中文

ByteDance

Backend engineering intern · Agent infrastructure

- Shanghai

On the agent infrastructure team, worked on the backends of the TVLA data-collection platform and the GymHub environment platform.

Keeping large-scale collection running reliably

With tens of thousands of collectors and hundreds of millions of records, I mapped capacity and stability bottlenecks across collection, upload, processing, and review, then pushed database scaling, read/write separation, index and cache work, fuller monitoring, and API migration. SLI on the established core path rose from 99% to 99.9%.

On the product side, I built task bundles, task recommendations, quota management, review rules, a trial flow for new collectors, user profiles, device-damage reports, and real-device evaluation, and set up dashboards for scale, quality, efficiency, and activity. That work carried the platform through the first half of the year: 12.21 million new collected records, and 87,700 hours of valid video in total.

I also reworked the IDL layering and project structure, filled in CI, unit tests, and scheduled jobs, and organized task-bundle sync and API services, so later changes and debugging cost less.

One interface for different environments

Built, from scratch, an environment platform for training, evaluation, annotation, and synthesis. Defined a shared create, delete, act, and observe interface for virtual machines, phones, and sandboxes, and implemented the runtime, project, quota, key, auth-integration, and audit modules.

Decoupled the underlying resources behind a provider interface, and used declarative scheduling plus a runtime state machine to manage each environment's lifecycle, so adding a new resource costs less.

Designed and implemented a warm cache pool. When the image cache hits, virtual-machine startup drops from 70-100 seconds to under 5 seconds.