已下线 校招

Research Engineer - LLM/VLM Inference Optimization (Seed Infra)

字节跳动

  • 西雅图
  • 研发

收录时间

已下线 · 历史岗位确认下线时间:。以下为收录时的信息,不代表当前仍可申请。查看历史岗位档案

岗位职责

About the Team The Seed Infrastructures team oversees the distributed training, reinforcement learning framework, high-performance inference, and heterogeneous hardware compilation technologies for AI foundation models. Responsibilities

  1. Design, develop, and optimize high-performance inference systems for large-scale LLMs and VLMs, covering inference engines, serving frameworks, and end-to-end deployment pipelines.
  2. Build state-of-the-art model inference engines through advanced performance optimization techniques such as compiler-level optimizations, parallel computing, graph fusion, efficient CUDA kernel development, low-precision computation, streaming inference, speculative decoding, and high-concurrency request optimization.
  3. Collaborate closely with other research teams to identify performance bottlenecks, conduct in-depth performance analysis, and optimize large models; contribute to the development of model toolchains and the broader technical ecosystem.

职位要求

Minimum Qualifications: - Bachelor's degree or above in Computer Science, Electrical Engineering, Software Engineering, or a related field. - Strong proficiency in C/C++ and Python; solid foundations in algorithms, data structures, and systems programming; familiarity with containerization and server-side debugging. - Hands-on experience with at least one mainstream machine learning framework (e.g., PyTorch, TensorFlow). - Experience deploying or optimizing LLM/VLM inference at production scale, with demonstrated impact on latency, throughput, or serving cost. - Familiarity with GPU architecture and experience optimizing compute-intensive operators (e.g., FlashAttention, GEMM, GEMV, Conv2D). Preferred Qualifications: - Experience with large-scale LLM serving infrastructure or equivalent production LLM deployment experience. - Experience in GPU programming (CUDA/OpenCL) and familiarity with frameworks such as TensorRT, Triton, or CUTLASS. - Experience in performance modeling, profiling, and optimization, or strong knowledge of CPU/GPU architectures. - Familiarity with model/data parallelism frameworks for distributed inference.

有些机会,只在官网短暂出现

真正值得关注的岗位,常常只在企业官网短暂开放,可能两三天后就下线,也未必会同步到综合招聘平台。没有持续关注,你甚至不会知道它曾经出现。职先机持续聚合并核验官网岗位,帮你抓住职场先机,快人一步。

微信小程序 / 当前岗位

在微信里继续看这个岗位

字节跳动Research Engineer - LLM/VLM Inference Optimization (Seed Infra)

扫码直达当前岗位收藏、浏览记录和下线提醒留在微信里
电脑端可直接微信扫码;手机端可保存小程序码后在微信中识别。
微信扫码在职先机小程序查看Research Engineer - LLM/VLM Inference Optimization (Seed Infra)正在准备岗位码
职先机微信小程序一岗一码 · 正式版直达保存小程序码

岗位提醒 · 01

收藏历史岗位,保留参考信息

当前岗位Research Engineer - LLM/VLM Inference Optimization (Seed Infra)字节跳动 · 西雅图

01

扫码进入当前岗位无需重新搜索,直接打开当前岗位详情。

02

收藏历史岗位此岗位已经下线,收藏仅用于保存历史参考。

03

继续查看相似机会原岗位变化时,可继续浏览相关在招岗位。

职先机会持续核验公开岗位;岗位状态与最终招聘结果仍以企业招聘官网为准。

微信小程序扫码打开当前岗位

扫码查看并收藏历史岗位,不代表当前仍可申请。

微信扫码在职先机小程序收藏Research Engineer - LLM/VLM Inference Optimization (Seed Infra)正在准备岗位码

微信小程序 · 当前岗位

请前往微信小程序提交内推申请

字节跳动Research Engineer - LLM/VLM Inference Optimization (Seed Infra)校招 · 西雅图

01进入当前岗位扫码或打开后直达本岗位内推申请页。

02提交 PDF 简历简历替换、处理进度和消息提醒在小程序内同步。

手机端将尝试直接打开微信;电脑端请使用微信扫码。

微信扫码进入Research Engineer - LLM/VLM Inference Optimization (Seed Infra)内推申请页正在准备岗位码
当前岗位专属码扫码后无需重新搜索岗位保存小程序码

内推提示

暂不支持该岗位内推

如需投递,建议前往官网进行投递。