已下线 校招

Tech Lead Cloud Site Reliability Engineer - DCS Cloud

字节跳动

  • 圣何塞
  • 研发

收录时间

已下线 · 历史岗位确认下线时间:。以下为收录时的信息,不代表当前仍可申请。查看历史岗位档案

岗位职责

Our Infrastructure Engineering team supports the company's fast growth by building and operating hyper-scale datacenters, managing the life cycle of server fleet, providing cloud solutions, and developing various infrastructure services and making sure they are scalable and are reliable. We have three subgroups for this role: - Cloud Host Delivery, Delivery & Standardization - Cloud Host Operation, Operation Efficiency & Reliability - Cloud Management & Security Responsibilities - What You'll Do - Design, build, scale, and operate ByteDance’s global infrastructure, including large-scale systems spanning public and private clouds. - Develop tools, automation frameworks, visualizations, and monitoring systems to streamline operations and drive optimization of global infrastructure. - Create, manage, and standardize cloud AMIs/images for use across multiple environments, ensuring strict alignment with the company's global compliance standards. - Thrive in a fast-paced environment, engaging in technical operations and on-call rotations to address incidents related to cloud, OS, network, performance, and reliability. - Drive improvements across the entire infrastructure lifecycle, from ideation and design through development, deployment, user support, and continuous refinement.

职位要求

Minimal Qualifications - Bachelor’s degree or above in Computer Science, Software Engineering, Information Security, or a related field. - 5+ years of experience in Linux operations, SRE, or DevOps - Proficient in at least one programming language such as Go, Python, or C++, with solid engineering capabilities in platform development, system tooling, and automation. - Strong computer science fundamentals, with deep understanding of Linux OS principles, computer networks, storage systems, GPU systems, and databases, along with systematic troubleshooting and root-cause analysis skills. - Familiar with core reliability practices, including monitoring and alerting, capacity management, change management, canary/gray releases, incident response, and postmortem processes. - Strong communication and collaboration skills, with the ability to proactively identify problems, drive cross-team execution, and demonstrate strong ownership and results-oriented mindset. Preferred Qualifications - Hands-on experience operating public cloud platforms, or deep familiarity with major cloud providers such as OCI, AWS, Azure, GCP, etc, including understanding of their underlying mechanisms. - Experience with large-scale cloud host delivery, image/AMI systems, resource scheduling, network adaptation, and virtualization technologies such as KVM/QEMU. - Familiar with containers and cloud-native ecosystems, including Docker, Kubernetes, and containerd, with a solid understanding of isolation mechanisms like cgroups and namespaces. - Experience maintaining GPU clusters, including drivers, CUDA, MIG, topology awareness, troubleshooting, stress testing, and GPU delivery pipelines. - Proven experience in reliability-focused initiatives such as failure drill systems, capacity governance, change governance, observability platforms, and resource cost optimization. - Open-source contributions, technical blogs, patents, or technical sharing experience are highly preferred. - Experience operating large-scale production environments is a strong plus.

有些机会,只在官网短暂出现

真正值得关注的岗位,常常只在企业官网短暂开放,可能两三天后就下线,也未必会同步到综合招聘平台。没有持续关注,你甚至不会知道它曾经出现。职先机持续聚合并核验官网岗位,帮你抓住职场先机,快人一步。

微信小程序 / 当前岗位

在微信里继续看这个岗位

字节跳动Tech Lead Cloud Site Reliability Engineer - DCS Cloud

扫码直达当前岗位收藏、浏览记录和下线提醒留在微信里
电脑端可直接微信扫码;手机端可保存小程序码后在微信中识别。
微信扫码在职先机小程序查看Tech Lead Cloud Site Reliability Engineer - DCS Cloud正在准备岗位码
职先机微信小程序一岗一码 · 正式版直达保存小程序码

岗位提醒 · 01

收藏历史岗位,保留参考信息

当前岗位Tech Lead Cloud Site Reliability Engineer - DCS Cloud字节跳动 · 圣何塞

01

扫码进入当前岗位无需重新搜索,直接打开当前岗位详情。

02

收藏历史岗位此岗位已经下线,收藏仅用于保存历史参考。

03

继续查看相似机会原岗位变化时,可继续浏览相关在招岗位。

职先机会持续核验公开岗位;岗位状态与最终招聘结果仍以企业招聘官网为准。

微信小程序扫码打开当前岗位

扫码查看并收藏历史岗位,不代表当前仍可申请。

微信扫码在职先机小程序收藏Tech Lead Cloud Site Reliability Engineer - DCS Cloud正在准备岗位码

微信小程序 · 当前岗位

请前往微信小程序提交内推申请

字节跳动Tech Lead Cloud Site Reliability Engineer - DCS Cloud校招 · 圣何塞

01进入当前岗位扫码或打开后直达本岗位内推申请页。

02提交 PDF 简历简历替换、处理进度和消息提醒在小程序内同步。

手机端将尝试直接打开微信;电脑端请使用微信扫码。

微信扫码进入Tech Lead Cloud Site Reliability Engineer - DCS Cloud内推申请页正在准备岗位码
当前岗位专属码扫码后无需重新搜索岗位保存小程序码

内推提示

暂不支持该岗位内推

如需投递,建议前往官网进行投递。