UCLA · M.S. Computer Science · Sep 2026 – Jun 2028 (expected)

Intelligence,
built to work.

I’m Yeqiao Fu, a computer science master’s student at UCLA. I build and test AI systems that work with data, use software tools, and run across multiple computers.

01 / Agent systems02 / Reasoning & evaluation03 / Distributed inference
Research ↔ Engineering

01 — Projects

Work

Projects Latest work first

What I built, why it matters, and what each project achieved.

SystemsJan 2026 – May 2026

ParaMind

Exploring larger models on everyday computers

Explored how everyday computers could jointly host a larger model without a central cloud provider; built a two-computer text-generation prototype.

Role
Capstone Student
Context
HKU · Capstone project
ResearchAug 2025 – Nov 2025

TABULA-R²

Reasoning across large tables and question types

Built an evaluation framework using 129 real-world tables to test different question types, including questions spanning multiple tables and tests with irrelevant data.

Role
Remote Research Assistant
Context
JHU · Supervised by Prof. Philipp Koehn
Open SourceJul 2025 – Aug 2025

Arkitect Cookbooks

Practical guides for developers and AI coding assistants

Created a hands-on tutorial and a concise code reference, both accepted into the official Arkitect repository.

Role
AI Agent Engineering Intern
Context
ByteDance · Volcano Engine, Ark Group
EngineeringJul 2025 – Aug 2025

Tool-Calling Evaluation

Checking whether AI uses tools correctly

Built a test framework that checks which tools an AI calls, whether they run, and whether the final answer is correct.

Role
AI Agent Engineering Intern
Context
ByteDance · Volcano Engine, Ark Group
EngineeringJul 2025 – Aug 2025

On-call Assistant

Sharing incident knowledge beyond individual experts

Helped turn past solutions into shared, searchable knowledge and built human-reviewed AI replies; piloted within a 30+ member team.

Role
AI Agent Engineering Intern
Context
ByteDance · Volcano Engine, Ark Group
ResearchJul 2024 – Nov 2024

WorldMap

Mapping websites to create step-by-step browser tasks

Built reusable maps of website interactions, then used them to generate browser tasks with an explanation of each action.

Role
Research Assistant
Context
HKU XLANG Lab · Supervised by Prof. Tao Yu
ResearchMar 2024 – Jun 2024

Spider2-V

Testing whether AI can complete real software tasks

Built tasks and automatic success checks for a benchmark of AI assistants using data-science software.

Role
Research Assistant
Context
HKU XLANG Lab · Supervised by Prof. Tao Yu
NeurIPS 2024
ResearchJun 2023 – Sep 2023

NLP for Scam-Related Text Analysis

Preparing scam-related text for language analysis

Prepared scam-related text for analysis and explored ways to represent words numerically.

Role
Independent Researcher
Context
Cambridge OSRP · Supervised by Kieren Lovell

02 — Background

Experience & Education

View CV ↗

Research, industry, and education, ordered by start date from newest to oldest.

  1. Current

    Education

    University of California, Los Angeles

    M.S. in Computer Science

    Expected completion: Jun 2028.

  2. Research & Industry

    Capstone Student · HKU · Capstone project

    Built the model-running software and helped connect two computers so they could generate text together.

  3. Research & Industry

    Remote Research Assistant · JHU · Supervised by Prof. Philipp Koehn

    Built tests of how AI answers questions using real-world tables, released the code and data, and wrote a technical report.

  4. Research & Industry

    Agent Tooling & Evaluation

    AI Agent Engineering Intern · ByteDance · Volcano Engine, Ark Group

    Built AI tool tests and incident-support workflows, and published practical guides to the Arkitect framework.

  5. Education

    University of California, Davis

    Exchange · Computer Science

    GPA: 3.8/4.0.

  6. Research & Industry

    Research Assistant · HKU XLANG Lab · Supervised by Prof. Tao Yu

    Built reusable maps of website interactions, then used them to generate browser tasks with an explanation of each action.

  7. Research & Industry

    Research Assistant · HKU XLANG Lab · Supervised by Prof. Tao Yu

    Built tasks and automatic success checks for a benchmark of AI assistants using data-science software.

  8. Research & Industry

    Independent Researcher · Cambridge OSRP · Supervised by Kieren Lovell

    Prepared scam-related text for analysis and explored ways to represent words numerically.

  9. Education

    The University of Hong Kong

    BEng in Computer Science · Minor in Finance

    Departmental graduation rank: 10/133.

Milestones are ordered chronologically; spacing does not represent duration.

03 — Goals

Direction

Graduate school and beyond

I want to become an AI-native researcher and builder who understands intelligent systems deeply and uses AI as an extension of cognition.

During graduate school, I also want to understand how information and capital flow through the world, while systematically improving both my intellectual and physical performance.

01 / Foundations

Understand intelligent systems deeply.

02 / Cognition

Use AI to extend thinking and building.

03 / World models

Understand information and capital flows.

04 / Performance

Improve intellectual and physical performance.

04 — Get in touch

Contact

yqfu5166@ucla.edu
Project notes

Systems Jan 2026 – May 2026

ParaMind

Exploring larger models on everyday computers

Role
Capstone Student
Context
HKU · Capstone project

The goal

Explore whether everyday computers connected over ordinary networks can pool their memory and computing power to collectively host and run a language model too large for any single device. The goal is to make larger models more accessible through a decentralized alternative to cloud services, while reducing the need to send sensitive data to a central provider.

My contribution

  • Split the model into working parts. Adapted Qwen and Llama so different computers could load and run different sections of the model, handling the inputs, outputs, and position information each section needed.
  • Decide what runs where. Built a planner using available memory and an estimate of computing speed, with ways to revise assignments when computers joined or left. Simplified an early, more complex design to fit the prototype’s needs.
  • Connect the pieces. Tested model execution with simulated connections, then integrated the networking system built by my project partner. Also built a separate desktop app with saved jobs and live text output; it was not connected to the multi-computer system.
One model, two computers
Computer A runs the first part of a model and passes its results to computer B, which runs the remaining part. The separate desktop app is not connected to this system.

Each computer runs part of the model. The desktop app below is a separate, single-computer prototype.

Open full-size diagram

Outcome

Delivered a working prototype that generated text across two computers on the same local network at HKU, using both Qwen and Llama. A separate desktop prototype supported local generation and saved jobs. These demonstrated the implemented components, rather than large-scale performance or privacy guarantees.

What I learned

Understanding models by building the systems that run them.

Splitting a model across computers gave me a deeper understanding of its structure and how it generates text, beyond simply calling an inference API. I learned how model components, intermediate results, device limits, and communication fit together—and how to choose a system design that serves the prototype’s actual needs.

Tools

PythonPyTorchTransformersFastAPISQLiteElectron

Resources

Public development code from our capstone team; not yet cleaned into a polished release.

Your place in the work list is saved.

Project notes

Research Aug 2025 – Nov 2025

TABULA-R²

TL;DRI built an evaluation framework using 129 real-world tables to study how small, locally run language models answer questions through data-processing tools. Models struggled with both cross-table reasoning and following the required action format, showing that tool-use reliability mattered alongside reasoning ability.

Role
Remote Research Assistant
Context
JHU · Supervised by Prof. Philipp Koehn

The goal

Investigate whether small, locally run language models can answer different types of questions using large, real-world tables—including questions that require combining information across multiple tables. The goal is to understand how data-processing tools can help models work beyond what fits in a single prompt, identify relevant information, and recover from mistakes.

My contribution

  • Prepare the questions and data. Selected and cleaned 129 Our World in Data tables. Created different question types with AI assistance, including questions spanning several tables and tests with irrelevant tables. Checked reference answers using database queries, manual review, or another language model, depending on the question.
  • Let models work with the data. Built a small set of table operations and checks for the model’s instructions. Returned results and errors so models could try another step before submitting an answer.
  • Compare models and understand mistakes. Compared models, example prompts, and how much table data they could see. Examined their working steps and final answers, released the code and data, and documented the study in a technical report.
From table questions to model tests
Prepare table questions and check reference answers, then let a model run data operations, receive feedback, and submit an answer for scoring.

Questions and reference answers are prepared first. During testing, models receive feedback on their data operations before their final answers are scored.

Open full-size diagram

Outcome

Released an evaluation framework and a dataset built from 129 real-world tables, with experiments and analysis documented in a technical report. In this setup, models struggled more with multi-table questions, and additional examples or visible data did not consistently help. Invalid tool actions were another source of task failure; final-answer scores also depended on AI-based grading.

What I learned

Understanding what an evaluation actually measures.

Smaller models often generated invalid actions, so a failed task did not necessarily mean they could not reason about the data. I learned to distinguish reasoning difficulties from problems following the tool interface, and to examine the steps behind a final score. Organizing the project also gave me practice connecting data preparation, experiments, and analysis into a coherent study.

Tools

PythonpandasSQLLocal language modelsAI-assisted grading

Resources

Original study report and short summary; not peer-reviewed publications. Findings are specific to the tested setup, and final-answer scores depend on LLM-based grading.

Your place in the work list is saved.

Project notes

Open Source Jul 2025 – Aug 2025

Arkitect Cookbooks

Practical guides for developers and AI coding assistants

Role
AI Agent Engineering Intern
Context
ByteDance · Volcano Engine, Ark Group

The goal

Make it easier for developers and AI coding assistants to build applications with Arkitect. The goal is to bridge the gap between knowing what the framework offers and knowing how to use it, through practical examples tailored to each audience.

My contribution

  • Teach through runnable examples. Wrote a Jupyter notebook demonstrating conversation context, tool use, live response output, and code that responds to events during execution.
  • Make the same ideas easy to reuse. Created a concise Python reference for coding assistants using the same framework. Submitted both guides through the publicly merged PR #211.

Outcome

Delivered a tutorial notebook for developers and a concise Python reference for AI coding assistants. Both were merged into the official Arkitect repository through PR #211 and are publicly available.

What I learned

Explaining the same system to different audiences.

Developers benefit from explanations and examples they can run and explore; AI coding assistants need concise patterns they can reuse. Writing both versions taught me to adapt the structure and level of detail to the reader, while keeping the underlying technical meaning consistent.

Tools

PythonJupyterArkitect

Resources

Your place in the work list is saved.

Project notes

Engineering Jul 2025 – Aug 2025

Tool-Calling Evaluation

Checking whether AI uses tools correctly

Role
AI Agent Engineering Intern
Context
ByteDance · Volcano Engine, Ark Group

The goal

Make it possible to judge whether an AI assistant uses tools reliably—not just whether its final answer sounds convincing. The goal is to distinguish mistakes in choosing tools, running them, and using their results, so developers can understand what needs fixing.

My contribution

  • Test different kinds of tools. Built a common way to register and test local functions, web APIs, and tools using MCP, a standard way for AI applications to access tools. Connected and ran all three types.
  • Make the checks configurable. Combined exact-answer checks, text patterns, and AI-based grading. Tests could separately check the expected order of tool calls, their execution, and the final answer.
  • Show where things went wrong. Saved structured results and readable execution logs, with summaries for individual tasks and groups of tests. Kept following an expected tool sequence separate from reaching a correct answer.

Outcome

Delivered and tested a framework covering local functions, web APIs, and MCP tools, with separate checks for tool use, execution, and final answers. The framework was submitted for review and entered limited internal testing during my internship.

What I learned

Making evaluation results useful for debugging.

A correct answer can come from an unexpected sequence of tools, while following the expected sequence does not guarantee correctness. I learned to check tool choices, execution, and answers separately, and preserve clear records so developers can understand a failure rather than just see a score.

Tools

PythonHTTP APIsMCPAI-assisted gradingExecution logs

Resources

Internal internship project; code and execution records are not public.

Your place in the work list is saved.

Project notes

Engineering Jul 2025 – Aug 2025

On-call Assistant

Sharing incident knowledge beyond individual experts

Role
AI Agent Engineering Intern
Context
ByteDance · Volcano Engine, Ark Group

The goal

Turn experience from past incidents into shared knowledge, so solving a problem does not always depend on finding the one person who knows that system. The goal is to help engineers work across technical boundaries, reuse previous solutions, and reduce repetitive support work, with a clear handoff between AI assistance and human control.

My contribution

  • Find useful knowledge. Analyzed more than 1,000 incident records and helped turn resolved conversations and solutions into a searchable knowledge base, connected to the company’s existing documentation.
  • Bring the information together. Added tools to retrieve monitoring-dashboard links from an issue identifier and make chat images available as context, using the company’s existing storage system.
  • Let engineers take control. Built an explicit takeover button that switched the assistant from answering to suggesting replies for an engineer to edit, discard, or send. Logged copy and send actions to understand whether suggestions were used.
Find context, then support the responder
The assistant retrieves incident information and initially responds. An engineer can take over, then review, edit, or discard suggested replies before sending.

An engineer can take over the conversation and review AI suggestions. The workflow uses no real incident data in this illustration.

Open full-size diagram

Outcome

Delivered a workflow connecting past incident knowledge, relevant context, and human-reviewed AI replies. It was piloted within a 30+ member team and received generally positive feedback, while answer accuracy and coordination with engineers remained areas for improvement.

What I learned

Designing assistance around how people actually work.

Making past solutions searchable was only part of the challenge. Engineers also needed clear responsibility, relevant context, and control over the assistant’s replies. I learned to connect these needs into one workflow, and to distinguish whether a suggestion was used from whether it was correct or helped resolve the incident.

Tools

PythonVikingDBGrafanaFeishuTOS

Resources

Internal team pilot; incident records and company documentation are not public.

Your place in the work list is saved.

Project notes

Research Jul 2024 – Nov 2024

WorldMap

Mapping websites to create step-by-step browser tasks

Role
Research Assistant
Context
HKU XLANG Lab · Supervised by Prof. Tao Yu

The goal

Explore whether a reusable map of a website can help AI understand what actions are possible, what they change, and how to reach a goal—rather than rediscovering the website for every task. A related goal is to use that map to generate step-by-step training examples without manually writing every task and action.

My contribution

  • Map how a website works. Cleaned page elements, grouped similar controls, and explored how actions changed a page or opened another one. Connected these interactions into a reusable website map.
  • Generate tasks from the map. Built a pipeline on WebArena that used a language model to generate action sequences, explain each step’s purpose, and summarize the task. Saved each action with its page location so the sequence could be followed again.
  • Reuse the map and review the examples. Reused the map across runs and showed the model a few steps at a time. Filtered repeated sequences and built a separate AI-based review of their completeness, complexity, brevity, specificity, and variety.
Map the website, then generate tasks
Explore website controls to build a reusable map, then use small parts of that map to generate and review step-by-step browser tasks.

The map is built first and reused. Generated tasks record both the actions and what each action is intended to achieve.

Open full-size diagram

Outcome

Built a working pipeline on WebArena that reused website maps to generate step-by-step browser tasks with action explanations. Saved examples included tasks spanning multiple pages and supported replay in the tested page states. Their effect on subsequent AI training was not evaluated.

What I learned

Learning to move an open-ended research problem forward.

This was an early experience of independently working through a research problem under supervision: identifying an obstacle, finding or building a method, testing it, and using the result to decide what to tackle next. The most valuable lesson was learning to repeat that cycle until separate ideas and tools became a working system.

Tools

PythonHTML / DOMXPathWebArenaBrowser automationLanguage models

Resources

Research prototype; raw browser-action records are not published here.

Your place in the work list is saved.

Project notes

Research Mar 2024 – Jun 2024

Spider2-V

Testing whether AI can complete real software tasks

Role
Research Assistant
Context
HKU XLANG Lab · Supervised by Prof. Tao Yu

The goal

Measure how far AI assistants are from independently completing realistic data-science and engineering tasks in professional software. The goal is to test actual task completion, not just the ability to describe a solution, and identify where these assistants still struggle.

My contribution

  • Help define practical research tasks. Joined discussions about data-science and engineering workflows, exploring Airflow and Snowflake as applications where tasks could be set up, attempted, and checked.
  • Turn workflows into repeatable tests. Wrote instructions, setup and reset scripts, and success checks for all Airflow tasks, LibreOffice Calc tasks, and selected Google Search tasks. Ran through complete Snowflake workflows and reviewed their action records and success checks.
  • Run experiments and explain the results. Ran selected comparisons using screenshots, structured descriptions of page controls, and numbered visual markers. Summarized results, prepared benchmark figures, and contributed to a small part of the paper.
Give a task, run it, check the result
Prepare a task and its software environment, let the AI try it, and check the actual result rather than incidental changes such as timestamps.

Each test starts from a prepared environment and checks whether the requested outcome was achieved.

Open full-size diagram

Outcome

Contributed executable tasks, success checks, workflow validation, selected experiments, and figures to Spider2-V. The team’s benchmark was published in the NeurIPS 2024 Datasets and Benchmarks Track, with the paper and project available publicly.

What I learned

Learning the full cycle of a research project.

I gained experience across defining the problem, designing experiments, preparing and collecting data, interpreting results, running follow-up experiments, and contributing to the paper. Seeing these stages connect taught me that research is not a fixed sequence: findings can change the next experiment and reshape how the work is explained.

Publication · NeurIPS 2024

Spider2-V: How Far Are Multimodal Agents From Automating Data Science and Engineering Workflows?

R. Cao, F. Lei, H. Wu, J. Chen, Y. Fu, et al.

Datasets and Benchmarks Track

Tools

PythonAirflowLibreOffice CalcSnowflakeAutomated testing

Resources

Your place in the work list is saved.

Project notes

Research Jun 2023 – Sep 2023

NLP for Scam-Related Text Analysis

Preparing scam-related text for language analysis

Role
Independent Researcher
Context
Cambridge OSRP · Supervised by Kieren Lovell

The goal

Explore how natural language processing can support the analysis of scam-related text. The goal is to turn raw text into a form that software can compare and analyze, laying groundwork for studying patterns in such messages.

My contribution

  • Prepare the text. Used Python’s NLTK library to split text into words, identify grammatical roles, and reduce related word forms to a common base.
  • Explore word representations. Investigated word vectors: numerical representations that allow software to compare patterns between words.

Outcome

Completed an exploratory workflow for preparing scam-related text and examining numerical representations of words. The work established a basis for further text analysis rather than a validated scam-detection system.

What I learned

Building a foundation for working with language data.

This project introduced me to preparing raw text and representing words numerically so software could analyze them. It gave me an initial understanding of the steps between collecting text and asking useful questions about it, providing a foundation for my later work with language models.

Tools

PythonNLTKText preprocessingWord vectors

Resources

Early research experience; no public code or report is linked.

Your place in the work list is saved.