
Ask a program team who owns the memory tier and you usually get four partial answers and one req. That was survivable while CXL lived in a lab. The specification stopped assuming it does.
Who writes the tiering policy? Who brings up the device and injects the errors? Who owns the channel at 128 GT/s? Who patches memory device firmware in a running fleet?
Four questions, and on most programs four partial answers and one req. That was survivable while CXL lived in a lab.
One Bad Quarter, Four Owners
The validation leads we work with describe a version of this that keeps recurring, and you may recognize it. A workload slows down at the ninety-ninth percentile, not on average. The memory device reports no errors. Somebody suspects the tiering policy is promoting the wrong pages, somebody else suspects the link is retraining, and your firmware team mentions that a device on that node took an update last week.
Untangling that takes four different competences, and on your org chart they are probably spread across teams that do not share a standup.
Platform software is one job: device enumeration, hotplug, NUMA and memory-region setup, the driver path, and now firmware activation on a device holding live data. The BSP and embedded engineers we place point out that this is Linux kernel and platform work with a hardware failure mode attached, a very different animal from application firmware.
Bring-up and error injection is another, covering poison handling and the RAS behavior that only shows up when you break something deliberately. Teams often try to cover it from the pre-silicon verification side. The engineers doing the job tell us that coverage closure in simulation and bench debug of a device misbehaving under real traffic have almost nothing in common.
The electrical path is a third, from connector through retimer into the slot, and the SI engineers we place say proving it takes measurement, not simulation alone.
Last, the tiering policy itself. What gets promoted, what gets demoted, on what signal, and what your application actually notices. That used to be partly the memory vendor's problem. The architects we talk to say it is now theirs.
When Did CXL Stop Being a Lab Project?
Around the time the specification started assuming the feature would run in production. CXL 3.2 arrived on December 3, 2024 with a CXL Hot-Page Monitoring Unit for memory tiering, common event records, online firmware activation, Post Package Repair enhancements, and Trusted Security Protocol extensions including meta-bits storage for host-only coherent device memory regions. Monitoring, in-place repair, live update, attestation. The engineers we place read that list as a description of something expected to run in a fleet, not on a bench.
The electrical bar rose alongside it. PCI-SIG released the PCIe 7.0 specification on June 11, 2025 at 128 GT/s. Above the tier, HBM4 landed as JESD270-4 in April 2025 with a 2048-bit interface and up to 2 TB/s per stack. A four-tier hierarchy is an ordinary architecture now, and the hardest failures the validation leads we work with report are no longer inside a tier but in the movement between tiers.
What the Teams Who Got It Right Did First
The hiring managers we talk to who have been through a CXL program are consistent about the sequence, and it is not the one most reqs imply.
They hired the validation engineer before the architect. Their reasoning: a tiering policy written against device behavior nobody has characterized gets rewritten the moment the device misbehaves, so the person who can make it fail in a controlled way earns their seat first.
Platform software they treated as permanent rather than project-scoped, because online firmware activation and post package repair are lifecycle features and the work does not stop at launch.
Several brought signal integrity in as a focused engagement instead of a standing seat, sized to the channel design and the correlation work, then again at the next connector or retimer change.
And the ones who hired the architect first say that is the decision they would change.
Their screening advice is worth borrowing too. Rather than asking whether a candidate knows CXL, describe the failure above and ask what they would want to rule out first. An engineer who has shipped this will not answer in order; they will ask you who sets the tiering thresholds, and whether that firmware update last week lines up with when the latency changed. A candidate working from the specification alone walks you back through the specification. The gap shows inside two minutes, which is more than a tool list will tell you in twenty.
Before You Run This Search Yourself
Everyone building AI infrastructure is reaching for the same people this year, and the engineers who have taken a CXL device from bring-up into production are mostly employed and not looking.
What we can tell you about our own numbers: two in three of the engineers we put in front of a hiring manager get an offer, about three times the industry norm, and 97 percent of the offers get accepted. Read together, those describe a shortlist of two or three engineers who fit the work instead of ten who fit a keyword, and an offer likely to hold the date on your req.
Memory stopped being the quiet part of your platform a while ago, and bring-up is an expensive place to discover which of the four jobs went unstaffed. Name all four owners early and bring-up gets boring, which on a memory program is about the highest compliment available.
If you are building this out, tell us what you are staffing and where the memory path worries you, and we will tell you which of those four searches is realistic on your timeline.
FAQ
Frequently Asked Questions
What Skills Does a CXL Memory Tiering Program Require?
Four different ones, usually written as a single req. Platform software for enumeration, hotplug, NUMA setup, the driver path and online firmware activation. Post-silicon validation for bring-up, error injection and poison handling. Signal and power integrity for the channel at 128 GT/s through a connector and retimer. Systems architecture for the tiering policy itself. The validation leads we work with say the failures live between those four rather than inside any one of them.
Can a Design Verification Engineer Cover CXL Post-Silicon Validation?
Usually not, though teams try. The engineers doing the job tell us that coverage closure in simulation and bench debug of a device misbehaving under real traffic have almost nothing in common. Post-silicon work is error injection, poison handling and RAS behavior that only appears when something is broken deliberately, on real hardware. It is a distinct hire from a pre-silicon DV engineer, even though both sit under verification on an org chart.
Which CXL Seat Should You Hire First?
The hiring managers we talk to who have been through a CXL program say validation, ahead of the architect. Their reasoning is that a tiering policy written against device behavior nobody has characterized gets rewritten the moment the device misbehaves, so the person who can make it fail in a controlled way earns the seat first. The ones who hired the architect first name that as the decision they would change.
Is CXL Still a Specialist Skill?
Less so every quarter. CXL 3.2 added a hot-page monitoring unit for tiering, online firmware activation and post package repair, which are operational features rather than research ones, and memory tiering is moving toward a per-server attach. The pool that has taken a device from bring-up into production is still small and concentrated at a handful of silicon and platform companies, so demand is outrunning supply rather than the skill staying exotic.
Written by
Game 7 Staff
