UCLA · M.S. Computer Science · Sep 2026 – Jun 2028 (expected)

Intelligence,
built to work.

I’m Yeqiao Fu, a computer science master’s student at UCLA. I build and study AI systems: how they reason, use tools, and work reliably beyond a single model.

01 / Agent systems02 / Reasoning & evaluation03 / Distributed inference
Research ↔ Engineering

01 — Projects

Work

08 projects Latest work first

Research, engineering, and open-source work — with contributions, methods, and project notes.

01 / SystemsJan 2026 – May 2026

ParaMind

Distributed LLM inference prototype

Built layer-sharded inference and resource-aware placement; integrated a two-node text-generation prototype.

Role
Capstone Student
Context
HKU · Capstone project
02 / ResearchNov 2025 – Feb 2026

TABULA-R²

Table reasoning through executable operations

Built an execution-based evaluation pipeline over 129 real-world tables, with structured plans and multi-turn feedback.

Role
Remote Research Assistant
Context
JHU · Supervised by Prof. Philipp Koehn
03 / Open SourceJul 2025 – Aug 2025

Arkitect Cookbooks

Reference materials for people and coding agents

Authored developer and coding-agent references, both merged into the official Arkitect repository.

Role
AI Agent Engineering Intern
Context
ByteDance · Volcano Engine, Ark Group
04 / EngineeringJul 2025 – Aug 2025

Tool-Calling Evaluation

Inspecting tool paths, execution, and answers

Built a framework for local, HTTP API, and MCP tools, separating tool-path compliance from answer correctness.

Role
AI Agent Engineering Intern
Context
ByteDance · Volcano Engine, Ark Group
05 / EngineeringJul 2025 – Aug 2025

On-call Assistant

Retrieval support with explicit human takeover

Built incident retrieval and human-reviewed reply workflows for an LLM assistant piloted within a 30+ member team.

Role
AI Agent Engineering Intern
Context
ByteDance · Volcano Engine, Ark Group
06 / ResearchJul 2024 – Nov 2024

WorldMap

Reusable website graphs & trajectory generation

Built reusable interaction graphs and a WebArena pipeline producing annotated candidate training trajectories.

Role
Research Assistant
Context
HKU XLANG Lab · Supervised by Prof. Tao Yu
07 / ResearchMar 2024 – Jun 2024

Spider2-V

Executable multimodal workflow evaluation

Built executable workflow tasks and success checks; contributed validation, observation ablations, and paper material.

Role
Research Assistant
Context
HKU XLANG Lab · Supervised by Prof. Tao Yu
NeurIPS 2024
08 / ResearchJun 2023 – Sep 2023

NLP for Scam-Related Text Analysis

Exploratory language representations

Applied NLP preprocessing and explored word-vector representations for scam-related text.

Role
Independent Researcher
Context
Cambridge OSRP · Supervised by Kieren Lovell

02 — Background

Experience & Education

View CV ↗

Research, industry, and education, ordered by start date from newest to oldest.

  1. Current

    Education

    University of California, Los Angeles

    M.S. in Computer Science

    Expected completion: Jun 2028.

  2. Research & Industry

    Capstone Student · HKU · Capstone project

    Built layer-sharded LLM inference and resource-aware placement, integrating a two-node text-generation prototype.

  3. Research & Industry

    Remote Research Assistant · JHU · Supervised by Prof. Philipp Koehn

    Built TABULA-R² and documented its evaluations in an 86-page technical report.

  4. Research & Industry

    Agent Tooling & Evaluation

    AI Agent Engineering Intern · ByteDance · Volcano Engine, Ark Group

    Built tool evaluation and on-call assistant workflows, and authored the publicly merged Arkitect Cookbooks.

  5. Education

    University of California, Davis

    Exchange · Computer Science

    GPA: 3.8/4.0.

  6. Research & Industry

    Research Assistant · HKU XLANG Lab · Supervised by Prof. Tao Yu

    Built reusable website interaction graphs and a graph-guided trajectory-generation pipeline on WebArena.

  7. Research & Industry

    Research Assistant · HKU XLANG Lab · Supervised by Prof. Tao Yu

    Contributed benchmark tasks, validation, and experiments to Spider2-V.

  8. Research & Industry

    Independent Researcher · Cambridge OSRP · Supervised by Kieren Lovell

    Explored text preprocessing and word-vector representations for scam-related text.

  9. Education

    The University of Hong Kong

    BEng in Computer Science · Minor in Finance

    Departmental graduation rank: 10/133.

Milestones are ordered chronologically; spacing does not represent duration.

03 — Goals

Direction

Graduate school and beyond

I want to become an AI-native researcher and builder who understands intelligent systems deeply and uses AI as an extension of cognition.

During graduate school, I also want to understand how information and capital flow through the world, while systematically improving both my intellectual and physical performance.

01 / Foundations

Understand intelligent systems deeply.

02 / Cognition

Use AI to extend thinking and building.

03 / World models

Understand information and capital flows.

04 / Performance

Improve intellectual and physical performance.

04 — Get in touch

Contact

yqfu5166@ucla.edu
Project notes / 01 of 08

Systems Jan 2026 – May 2026

ParaMind

Distributed LLM inference prototype

Role
Capstone Student
Context
HKU · Capstone project

Overview

An undergraduate capstone exploring how to split a language model across multiple machines, place its layers within resource budgets, and connect execution across nodes. I implemented the inference and application components; my project partner developed the P2P transport subsystem and contributed architectural ideas.

My contribution

  • Model-aware sharding. Adapted loading and forward execution for Qwen and Llama, including embedding/output boundaries, shared weights where applicable, and positional information. Corrected a Llama rotary-position assumption against the implementation we were using.
  • Resource-aware placement. Used timed matrix multiplication as a compute proxy alongside available memory budgets. Developed greedy contiguous-layer placement and local replanning for node joins and departures.
  • Interface-based integration. Developed inference against mock connections, then integrated the partner-built transport interfaces carrying tokens, hidden states, logits, and route/request information.
  • Separate desktop output. Built an Electron/FastAPI prototype with persistent SQLite jobs and streamed local generation. This GUI was not connected to the real P2P inference path.
Fig. 01Distributed runtime & separate desktop prototype
Distributed runtime & separate desktop prototype. Simplified view of the demonstrated runtime. Layer boundaries are illustrative; the desktop application is a separate output.

Simplified view of the demonstrated runtime. Layer boundaries are illustrative; the desktop application is a separate output.

Open full-size diagram

Design note

Choosing a sufficient planner

I initially built dynamic-programming planners, then simplified to greedy placement when the prototype did not require the more elaborate optimizer. This was a complexity trade-off, not a measured speedup or a claim of optimal placement. Looking back, balancing stages also does not by itself establish lower latency for a sequential autoregressive request.

Results & scope

Demonstrated two-node loading and text generation with both Qwen and Llama on the same local network within HKU. These were executable proofs of concept, not comprehensive performance or correctness evaluations. Revised placement and staged reconfiguration did not establish reliable live migration.

Communication analysis, not a 72B run

A separate report calculation used a representative Qwen2.5-72B configuration to estimate activation traffic: about 32 MiB per boundary during a 2,048-token prefill, versus 16 KiB per decode token, under batch-one FP16 assumptions. This was an idealized communication budget—not a 72B execution result, a campus-network measurement, or proof of sufficient device memory.

Tools & methods

PythonPyTorchTransformersModel shardingResource-aware placementFastAPISQLiteElectron

Resources

Public development code; not yet cleaned into a polished release. The repository contains collaborative work.

Your place in the work list is saved.

Project notes / 02 of 08

Research Nov 2025 – Feb 2026

TABULA-R²

Table reasoning through executable operations

Role
Remote Research Assistant
Context
JHU · Supervised by Prof. Philipp Koehn

Overview

An independently organized evaluation and tooling project for context-limited local LLMs. Under Prof. Philipp Koehn’s broad research direction, I chose multi-table and distractor-table tasks and developed the data preparation, execution framework, experiments, analysis, and technical report.

My contribution

  • Data and task preparation. Selected, cleaned, and organized 129 Our World in Data tables totaling 70,528 rows. Constructed single-table, multi-table, and distractor tasks through LLM-assisted generation and different reference-answer checks: SQL, direct review, or an LLM judge.
  • Executable reasoning loop. Built a constrained table-operation language with PLAN/END parsing, syntax checks, and a pandas executor. Returned results and errors over multiple turns so models could attempt corrections.
  • Comparative evaluation. Compared models, few-shot and reasoning examples, and table-observation settings. Inspected execution errors and interaction traces alongside final answers; released code and data and documented the study in a technical report.
Fig. 02Task preparation & execution-based evaluation
Task preparation & execution-based evaluation. Reference-answer construction and final-answer scoring are separate. SQL checks were not rerun for each evaluated model answer.

Reference-answer construction and final-answer scoring are separate. SQL checks were not rerun for each evaluated model answer.

Open full-size diagram

Design note

Tools must also be usable by the model

Defining an action space was not enough: small local models often produced invalid outputs. Few-shot examples, explicit PLAN/END boundaries, syntax checking, and error feedback made the interaction more usable, while leaving meaningful failure modes to study.

Results & scope

The report recorded weaker performance on multi-table tasks and cases where extra examples or table context did not improve outcomes. These observations apply to this evaluation setup: formatting, parsing, execution, and grading all influenced the measured result. Feedback supported recovery but did not guarantee compliant or correct output.

Interpreting open-ended questions

Proxy-inference items reasoned from observed indicators to concepts not directly measured in the tables. Their answers depended on assumptions; agreement with a stored reference was not evidence that the inferred relationship was true. Likewise, a SQL NULL result was a question-specific check, not a general proof of unanswerability.

Tools & methods

PythonpandasPLAN/END DSLSQL reference checksLocal LLM inferenceLLM-assisted evaluation

Resources

Your place in the work list is saved.

Project notes / 03 of 08

Open Source Jul 2025 – Aug 2025

Arkitect Cookbooks

Reference materials for people and coding agents

Role
AI Agent Engineering Intern
Context
ByteDance · Volcano Engine, Ark Group

Overview

Two complementary references for the same agent framework: a Jupyter Notebook for developers and a concise Python reference for coding agents. Developed during my internship in the Data Department of ByteDance’s Volcano Engine Ark Group.

My contribution

  • Human-oriented examples. Wrote a tutorial notebook demonstrating context/state management, tool registration, streaming, and lifecycle hooks.
  • Agent-oriented reference. Prepared a concise Python counterpart that coding agents could consult when using the same interfaces.
  • Public contribution. Submitted both formats through PR #211, which was merged into the official volcengine/ai-app-lab repository.

Design note

One interface, two audiences

The two formats explain the same framework, but serve different readers: a notebook supports human exploration and explanation, while the concise Python reference exposes usable patterns for coding agents.

Results & scope

Both reference formats were publicly merged through PR #211. This is a distinct open-source contribution; the merge does not imply that the separate tool-evaluation framework or on-call assistant was publicly released.

Tools & methods

PythonJupyterArkitectContext managementTool registrationStreamingLifecycle hooks

Resources

Your place in the work list is saved.

Project notes / 04 of 08

Engineering Jul 2025 – Aug 2025

Tool-Calling Evaluation

Inspecting tool paths, execution, and answers

Role
AI Agent Engineering Intern
Context
ByteDance · Volcano Engine, Ark Group

Overview

An internally trialed framework for inspecting whether an agent called the right tools, whether those tools executed, and whether the answer used their output correctly. Designed to make evaluation decisions inspectable rather than reduce all behavior to one unexplained pass/fail.

My contribution

  • Unified tool registration. Integrated and ran local functions, HTTP APIs, and MCP tools, including search tools, through a common evaluation setup.
  • Configurable validation. Combined exact-match checks, regular expressions, and LLM-as-judge criteria. Task validation modules could specify expected tool sequences as well as answer checks.
  • Execution visibility. Retained structured outputs and human-readable execution logs, with analysis at batch, task-group, and individual-task levels.

Design note

A correct answer can follow a different path

A model could reach a correct answer through an alternative route and pass an answer check while failing a strict expected-sequence check. An LLM judge could also accept alternatives when the configured criteria allowed them. Keeping these judgments separate made the result easier to interpret.

Results & scope

All three tool categories were integrated and exercised, and I submitted a pull request. The framework remained in limited internal trial when the internship ended; a public release is not confirmed.

Tools & methods

PythonHTTP APIsMCPExact matchRegexLLM-as-judgeStructured logging

Resources

Internal internship work. No public code or internal execution records are linked.

Your place in the work list is saved.

Project notes / 05 of 08

Engineering Jul 2025 – Aug 2025

On-call Assistant

Retrieval support with explicit human takeover

Role
AI Agent Engineering Intern
Context
ByteDance · Volcano Engine, Ark Group

Overview

An LLM-assisted support workflow addressing repeated incident questions, unclear ownership, and dependence on a few specialists. My work connected reusable incident knowledge, relevant context, explicit human takeover, and usage feedback.

My contribution

  • Knowledge retrieval. Analyzed 1,000+ incident records and worked on an L1 knowledge base of resolved conversations and solutions, connected with the company’s L2 knowledge base through VikingDB retrieval.
  • Incident context. Given an issue identifier, fetched associated Grafana dashboard links for responders and model context. Used the existing secure TOS environment to store and retrieve chat images; I did not introduce a new security mechanism.
  • Human control. Implemented an explicit takeover control in the incident group chat. After takeover, the LLM suggested replies that the responder could discard, edit, or approve and send.
  • Feedback signals. Logged copy and one-click-send actions through Feishu to observe whether suggestions were used, without treating use as proof of factual correctness or incident resolution.
Fig. 03Incident context & human control
Incident context & human control. Conceptual, sanitized workflow based on my implementation. No internal data, dashboard addresses, or infrastructure details are reproduced.

Conceptual, sanitized workflow based on my implementation. No internal data, dashboard addresses, or infrastructure details are reproduced.

Open full-size diagram

Design note

Takeover is an action, not online presence

The responder explicitly clicked a takeover control; merely being online did not switch the assistant’s role. Relinquishing the incident or a configured inactivity interval could return handling to the LLM. The trial interval was configurable, so no fixed production timeout is claimed.

Results & scope

The assistant was piloted within a 30+ member team and received generally positive feedback. Answer accuracy and coordination with responders remained weaknesses. The pilot did not establish a measured reduction in resolution time or staffing needs.

Tools & methods

VikingDBRetrievalGrafana contextTOSFeishuHuman-in-the-loop workflows

Resources

Internal team pilot. This summary describes contributions and scope without publishing incident records, messages, or company documentation.

Your place in the work list is saved.

Project notes / 06 of 08

Research Jul 2024 – Nov 2024

WorldMap

Reusable website graphs & trajectory generation

Role
Research Assistant
Context
HKU XLANG Lab · Supervised by Prof. Tao Yu

Overview

A representation and data-generation project asking how an agent could understand where a component is, what interacting with it changes, and how to reach a desired state. I built reusable site interaction graphs and used them to generate annotated computer-use trajectories on WebArena.

My contribution

  • Interaction graph construction. Started from raw DOM and XPath-indexed components, merged similar elements, removed irrelevant content, and explored interaction types with a custom action space. Recorded within-page state changes and connected pages through navigation transitions.
  • Graph-guided generation. Sampled starting components, generated multiple action sequences with an LLM, annotated step intentions, and summarized or named the resulting tasks. Applied model-based deduplication to repeated trajectories.
  • Bounded planning context. Implemented segmented graph presentation with a configurable horizon. After traversing a segment, generation continued from the selected endpoint using the already constructed graph.
  • Assessment and replay. Recorded action types, inputs, XPath, accessibility roles/names, URLs, and intentions. Built a separate GPT-4o mini evaluator for completeness, complexity, conciseness, concreteness, and diversity.
Fig. 04Reusable graph → annotated trajectories
Reusable graph → annotated trajectories. Graph construction is an upfront step reused across generation runs. Segmented observation does not rebuild the graph.

Graph construction is an upfront step reused across generation runs. Segmented observation does not rebuild the graph.

Open full-size diagram

Design note

An expensive representation can still be reusable

Building the graph was costly, but it was constructed before trajectory generation and reused in the limited test environments. This distinguished the cost of discovering interaction structure from the cost of consulting it repeatedly. Agent-editable graphs remained an unimplemented idea.

Results & scope

The graph-guided pipeline ran on WebArena and produced saved annotated trajectories, including cross-page sequences. Recorded examples supported grounding and replay in the tested states. The work did not establish downstream training gains or reliable replay across arbitrary websites and changing page states.

A concrete recorded sequence

One saved forum example followed Reply → enter a response → Post → return to the discussion. Each action carried its locator and intention, and the final stop entry summarized the task. This illustrates the recorded format; it is not a newly executed live demo.

Tools & methods

HTML / DOMXPathWebArenaBrowser automationInteraction graphsLLM-generated trajectoriesGPT-4o mini evaluation

Resources

Research prototype. Raw trajectories are not published here; examples are described without copying local URLs or unreviewed content.

Your place in the work list is saved.

Project notes / 07 of 08

Research Mar 2024 – Jun 2024

Spider2-V

Executable multimodal workflow evaluation

Role
Research Assistant
Context
HKU XLANG Lab · Supervised by Prof. Tao Yu

Overview

A multimodal-agent benchmark for professional data science and engineering workflows, published at NeurIPS 2024, Datasets and Benchmarks Track. My work connected problem scoping, executable task construction, selected experiments, and communication of results.

My contribution

  • Research scoping. Explored Airflow and Snowflake as candidate applications, considering practical relevance, GUI interaction, and whether setup and evaluation could be made executable.
  • Task construction. Built abstract and verbose instructions, setup/reset scripts, and outcome checks for all Airflow tasks, LibreOffice Calc tasks, and selected Google Search tasks. Ran complete Snowflake workflows and inspected trajectories and success-check scripts.
  • Observation-space experiments. Ran and monitored selected screenshot, accessibility-tree, and Set-of-Mark ablations, summarized results, and helped interpret them. I did not run the final configuration combining all enhancements.
  • Research communication. Prepared benchmark visualizations, including a task-area ring chart, and assisted with a small part of the manuscript.
Fig. 05Making task completion checkable
Making task completion checkable. Illustration of my task-building workflow, not the full benchmark architecture. Checks target the requested result rather than incidental run-to-run differences.

Illustration of my task-building workflow, not the full benchmark architecture. Checks target the requested result rather than incidental run-to-run differences.

Open full-size diagram

Design note

Check the outcome, not incidental text

For task logs, merely having a file or some text was insufficient, but byte-for-byte matching was too strict because timestamps and other run-dependent details changed. I checked information that indicated the requested completion.

Results & scope

On the evaluated subset, unaligned screenshot-plus-tree input underperformed tree-only input, while the SoM configuration outperformed the unaligned combination. This compares the tested configurations; it does not establish alignment as the sole cause. Snowflake validation covered the complete task workflows, not every possible incorrect outcome.

Publication · NeurIPS 2024

Spider2-V: How Far Are Multimodal Agents From Automating Data Science and Engineering Workflows?

R. Cao, F. Lei, H. Wu, J. Chen, Y. Fu, et al.

Datasets and Benchmarks Track

Tools & methods

Benchmark constructionEnvironment setup / resetOutcome checksAirflowLibreOffice CalcSnowflakeObservation ablations

Resources

Your place in the work list is saved.

Project notes / 08 of 08

Research Jun 2023 – Sep 2023

NLP for Scam-Related Text Analysis

Exploratory language representations

Role
Independent Researcher
Context
Cambridge OSRP · Supervised by Kieren Lovell

Overview

An early exploratory research project at Cambridge OSRP, supervised by Kieren Lovell, applying text preprocessing and word-vector representations to scam-related text.

My contribution

  • Text preparation. Used NLTK tokenization, part-of-speech tagging, and lemmatization to prepare the text for exploratory analysis.
  • Representation exploration. Investigated word-vector representations of the preprocessed text.

Results & scope

The documented work is exploratory preprocessing and representation analysis. No validated scam-detection accuracy, deployed incident-response system, or public software release is claimed.

Tools & methods

PythonNLTKTokenizationPOS taggingLemmatizationWord vectors

Resources

Early research experience. No public artifact is linked.

Your place in the work list is saved.