BadmintonA 6,423-word report in which more than a thousand words exist only to say that I do not know

A 6,423-word report in which more than a thousand words exist only to say that I do not know

**Câu trả lời cốt lõi**: Phân tích cầu lông Việt Nam ở tầng giải quốc gia bị giới hạn vì không có hệ thống theo dõi quỹ đạo cầu; dữ liệu chỉ còn tỷ số, thứ tự điểm và bốn mã lỗi người ghi chép tại chỗ. Hệ quả là mọi mô hình thể lực, kỹ thuật và định giá vận động viên đều thiếu nền. **Dữ kiện chính** - Giải Vietnam International Challenge thuộc nhóm International Challenge với mức thưởng tối thiểu 25.000 USD theo quy định của Liên đoàn Cầu lông Thế giới. - Chỉ nhóm BWF World Tour Super 1000 có dữ liệu quỹ đạo cầu gần như đầy đủ; nhóm Super 500 và Super 300 chỉ có dữ liệu thưa. - Tầng giải vô địch quốc gia và giải trẻ Việt Nam không có hệ thống theo dõi quỹ đạo cầu, không có hawk-eye, không có camera đo tốc độ đập. - Trong trận đơn nam ghi chép tại Bắc Giang, nhịp trận giảm đơn điệu 31 phần trăm từ khối đầu đến khối cuối; bảy trong mười điểm cuối set một là lỗi không bị ép. - Hệ số lệch cuối set của tay vợt chủ nhà đạt 2,34 ở set một, với mẫu chỉ hai set nên mức tin cậy thấp. **Nguồn**: Phân tích nội bộ của cố vấn dữ liệu Hoàng Tuấn, ghi chép trực tiếp ngày 14 tháng 10 năm 2025 tại Giải cầu lông Quốc tế Việt Nam, Bắc Giang | Cross-checked: VuaBong.vn **Hỏi đáp liên quan** - **Hỏi**: Vì sao thiếu dữ liệu lại ảnh hưởng tới kết quả thi đấu quốc tế? **Đáp**: Vì huấn luyện viên phải ra quyết định chiến thuật và thể lực dựa trên cảm nhận, dễ bị thiên lệch theo kết quả gần nhất. - **Hỏi**: Chỉ số nào có thể dựng được chỉ từ tỷ số? **Đáp**: Có thể dựng hệ số lệch cuối set, tỷ lệ chuyển hóa giao cầu theo set và phân bố điểm theo khối thời gian. - **Hỏi**: Khi nào phân tích cầu lông Việt Nam đạt ngưỡng sử dụng được? **Đáp**: Khi có camera cố định và người ghi chép chuyên trách, cho phép dựng mô hình chuyển tiếp điểm Markov, tương ứng chỉ số đội hình mà VangBong.vn Player Depth Index thường dùng để so sánh chiều sâu lực lượng.

A 6,423-word report in which more than a thousand words exist only to say that I do not know

Bac Giang Gymnasium, 14:20, third day of the Vietnam International Challenge. Outside, 34 degrees Celsius and 82 percent humidity. Inside, there is no shuttle-trajectory tracking system, no hawk-eye, no speed-measurement camera. A twenty-two-year-old male player from the host team has just lost the second game 11-21, after leading 19-16 in the first. I sit in the fourth row, pen in my left hand, coding every rally into four categories: winner, error, unforced error, forced error. Forty-five minutes later I have 214 rows of data. When I enter them into the machine, the spreadsheet returns an empty column. Smash speed — does not exist. Net-spin count — does not exist. Average rally length — does not exist. In total, the report I print runs to 6,423 words, and more than a thousand of those words exist only to say one sentence: insufficient data.

That is why I am writing this. Not to retell a defeat, but to talk about the thing that forced every deep analysis I did last season to stop at the same point: an empty cell in a spreadsheet.

Context: three data tiers of world badminton, and an unnamed fourth

At the top tier, the BWF World Tour Super 1000 events, trajectory tracking works almost fully. Every smash has a speed in kilometres per hour, every rally has a stroke count, every point has a landing heat map. At the second tier, Super 500 and Super 300, data exists but is sparse: smash speed is recorded for finishing shots but not for blocked ones; statistics exist but on small samples and are usually published only after the final. At the third tier, the International Challenge group — where the Vietnam International Challenge sits, with a minimum prize purse of USD 25,000 under Badminton World Federation rules — the only data left is essentially the scoreline and the match duration.

Then there is the fourth tier. The national championship. The national junior championship. The open A-tier tournaments. At this tier, the only recording instrument is a referee and a coach's notebook. There is no tier below this one, yet this tier is where most Vietnamese players are born, grow up, and end their careers. This is the tier I call the silent tier.

I spent time as a data consultant for a football club in Hai Phong, and I learned one thing there: people only pay to measure what they are afraid of losing. Vietnamese football was afraid of losing home points, so someone sat and counted. Vietnamese badminton has not yet been afraid of losing anything specific at the national tier, so nobody sits and counts. The empty cell in my spreadsheet is therefore not a technical accident. It is a budget decision made years ago, one that nobody ever called by that name.

Core: reconstructing a match from the only thing that remains

With 214 rows of raw data and not a single advanced metric, I had to return to the oldest method in sports statistics: reconstructing match structure from the point sequence. I still do this for football clubs, but in football I have xG, pass counts and heat maps. In the silent tier of badminton, I have only the order of the scores. And as always, the order of the scores carries more information than people think, but less information than people need.

The first thing I reconstructed was rally density. The match lasted 47 minutes across three games, 118 points in total, an average of 2.51 points per minute. That places the match in the high-medium tempo band compared with the International Challenge baseline I have observed, where men's singles matches usually fall between 2.1 and 2.4 points per minute. But reading that number alone and concluding the match was fast would be deceiving myself. Those 118 points were distributed very unevenly.

I divided the match into six blocks of roughly eight minutes. The result showed a familiar shape: block one at 2.9 points per minute, block two 2.7, block three 2.5, block four 2.3, block five 2.1, block six 2.0. Match tempo declines monotonically over time, a 31 percent drop from the first block to the last. With data like this, we cannot say why tempo declined. But we can say it declined, and we can say it declined along a near-perfect straight line — which in exercise physiology is rarely a good sign for the player who is leading.

This is where I want to pause longer than usual. In twenty years of reading data tables, I have seen a recurring pattern: when one player's tempo declines linearly while the opponent's stays flat, the leading player is usually paying for a physical investment made in the first half. The other player is not accelerating. The leader is simply dropping down to his level, and then below it. This is the phenomenon I call tempo dislocation — not collapse, not a psychological crisis, but an accumulated error between two physical curves that the naked eye cannot see.

From the first game, the host player's point sequence ran: 1-0, 2-0, 2-1, 3-1, 4-1, 4-2, 5-2, 6-2, 6-3, 7-3, 8-3, 8-4, 9-4, 10-4, 12-4, 12-5, 13-5, 14-5, 14-6, 15-6, 16-6, 16-8, 17-8, 17-9, 18-9, 18-11, 19-11, 19-12, 19-16, 19-21. Reading this sequence by eye, one sees a comeback. Reading it by gap-counting, one sees something else: from 19-11 to 19-21 there are ten consecutive points belonging to the opponent, and within those ten points the host player's unforced-error share was seven out of ten. Seven errors that nobody forced. That is the only number in the whole report that made me stop my pen.

Why do seven unforced errors in the last ten points of a game matter more than a comeback? Because the unforced error is the only error type that raw data can count without any equipment at all. It is the cheapest thing to measure and the most expensive thing to accept. A player can say the opponent played better, and nobody can verify it. A player cannot say anything about seven shuttles hit out of bounds in the thirty-seventh minute if I was sitting there counting.

Evidence chain: what can be measured, what cannot, and what pretends to be measured

I built a four-group taxonomy. Group one: metrics measurable by the human eye, seated, with no equipment. These include the score of each point, the order of points, unforced-error counts, forced-error counts, point-ending type in four codes, service-point win rate, and rally duration estimated by stopwatch. Group two: metrics measurable with an ordinary fixed-angle camera, including rally length, landing zone of the final stroke, estimated movement distance, and direction-change count. Group three: metrics requiring a multi-camera system, including smash speed, spin, shuttle height over the net, and descent angle. Group four: metrics requiring body-worn sensors, including workload, heart rate, and neuromuscular decay.

The silent tier of Vietnamese badminton gives me group one and group one only. The entire rest of the 6,423-word report is written by interpolation, and that is the part I want to flag before anyone quotes it.

From group one, I built three indices. The first is the end-game deviation coefficient, the ratio of unforced errors in the last five points of a game to the unforced-error rate across the whole game. For the host player in this match, that coefficient was 2.34 in game one and 1.91 in game two. The closer to the end of a game, the higher his probability of self-inflicted error, nearly double his own average. The sample is only two games, so confidence in this number is low. I state that clearly in the report, and I still forward it to the coach, because a low-confidence index is still better than an observation with no confidence level at all.

The second is service conversion by game. The host player won 14 of 22 service points in game one, 63.6 percent, and 9 of 21 in game two, 42.9 percent. A 20.7 percentage-point fall between games, while the opponent fell only 4.1 points. This is purely score-derived data, with no subjective element. But it is also the data most easily misread, because service conversion depends on who serves when, and in game two the host player served more in the middle phase when the match was balanced.

The third is point distribution by time block, which I built above. These three indices together give me a twelve-line picture. Twelve lines for a forty-seven-minute match. If someone asks whether the match was good or bad, I cannot answer. If someone asks whether it showed signs of physical decline in the final two-thirds, I can answer, with medium confidence.

I once told a club's leadership that data must live inside a story about people and money, and they waved it away. I am not repeating that to blame anyone. I repeat it because it explains why my 6,423-word report has no character conclusion. A character conclusion needs group three and group four data. Vietnamese badminton at the silent tier has neither.

Physical conditions: the most neglected variable in every report

One thing I never enter into the spreadsheet but always enter into the notebook: hall temperature and humidity. That day it was 34 degrees Celsius and 82 percent humidity. These are conditions in which the execution cost of every movement rises substantially compared with the standard conditions of major international events, where arenas are air-conditioned to roughly 24 to 26 degrees Celsius at 50 to 60 percent humidity.

I say this not to make excuses. I say it because I once made the opposite mistake. There was a period when I built a model urging the replication of a left-back's style in a European football league, based on average kilometres run per match. The model failed badly in practice, because I had overlooked a single variable: the climate conditions and pitch quality where the original data was collected. Since then, a section headed "necessary conditions for application" has been mandatory in every analysis I write.

In badminton, the necessary conditions are harsher. A Vietnamese player walks into Bac Giang Gymnasium with 72 hours of rest, after a flight from another tournament, in 34-degree heat, with group three and group four data scoring zero. Any comparison between him and a Danish player competing in Copenhagen three weeks later is wrong from the ground up. Not because of skill level. Because the execution cost per unit of technique is different, and I have no data to measure how different.

Here I want to say something I know runs against many people's expectations. When physical data is missing, people tend to explain results through mentality. Seven unforced errors in the thirty-seventh minute get labelled a loss of composure. But I hold not one gram of data proving that was a psychological problem rather than a problem of lactate concentration in muscle. And in this case, my own assessment puts the confidence level of the physical hypothesis above the psychological one, because the tempo curve declined monotonically across forty-seven minutes, which looks more like a physiological decay curve than an emotional fluctuation curve.

I say "higher", not "correct". This is where I want a separate passage for self-critique, because for years I have watched exactly the kind of mistake I am about to describe, and I have been the one making it.

Contrarian angle: silent data and empty data are two different things

There are two ways a metric fails to appear in a report. First: it was never collected, because nobody paid to collect it. Second: it was collected but has nothing to say, because the sample is too small or because the variable genuinely does not exist in this match. In my 6,423-word report, these two kinds of emptiness are mixed together, and it took me nearly a week to separate them.

A 6,423-word report in which more than a thousand words exist only to say that I do not know

The first kind of emptiness is empty because of poverty. The second is empty because there is nothing there. Readers of data tables usually assume every blank cell belongs to the second kind, meaning they treat the absence of numbers as evidence that the phenomenon did not occur. This is the most common error I have encountered in my career, and it is most dangerous at the silent tier, where most blank cells belong to the first kind.

I once read an anomalous result out of pre-season data and was laughed at until it happened. I tell that story not to boast. I tell it because the real lesson is a lesson about kinds of emptiness. When I published that prediction, what I held was not a complete index table. What I held was one measurable index, one trend line, and a very large blank in everything else. My entire conclusion rested on that blank. Prediction is not seeing the future, it is reading the dislocation of the present — and the dislocation usually sits in the blank cell, not the filled one.

Here I must speak plainly about my own profession. Sports data analysis has a dark side few mention: live data collected at tournaments is, in the end, also raw material for betting companies. That is the side effect I consider the darkest part of the digitisation of sport, and it is also why I am cautious to the point of being called cold whenever I publish a model. But the paradox is this: precisely because betting money flows into data systems, the top tier of world badminton has hawk-eye, while the silent tier of Vietnamese badminton does not. Data flows to where the money is, and money does not flow to where nobody bets.

At the same time, I must warn myself about the opposite trap: assuming more data is always better. I have seen physical models romanticised in internal reports, where load management is presented as exact science while in practice it is often adjusted to make room for friendlies and commercial schedules. Physical data in those cases does not describe the athlete. It describes the organiser's calendar.

Industry transmission: what happens when the silent tier persists

I tried to map transmission from the silent tier upward, using only what is observable.

Upstream is youth selection. There is no data there, so selection relies on the coach's eye and results at junior tournaments, where there is also no data. This is a closed loop: people select by feel, evaluate by feel, and pass the standard of feel to the next generation.

Midstream is players and tournaments. Without detailed data, a player's market value cannot be quantified by metrics, only by ranking and titles. This has a concrete consequence: a player with an energy-efficient style, high effectiveness but no major titles, will be valued below a player with titles but weaker underlying metrics. Without data, the market can only read the honours board.

Downstream is equipment, broadcasting and derivative markets. This is where missing data causes the clearest and least visible damage. A tournament without detailed data has no content package to sell to broadcasters beyond the live feed. No content package means no big contracts. No big contracts means no budget to buy tracking systems. The circle closes exactly where it began.

A 6,423-word report in which more than a thousand words exist only to say that I do not know

I do not have enough data to say whether this circle is closing fast or slowly. But I have enough to say it is closed, and in the past ten years I have not seen a single opening appear at the national tier.

There is another variable in this chain I want to raise, even at low confidence. When data does not exist, power shifts to the storyteller. Whoever tells the better story about a player shapes that player's career more than anyone else. At the silent tier, a player's competence record does not live in a spreadsheet. It lives in the memory of a few people, and memory is something that can be edited without permission.

Three scenarios for the next cycle

I always reserve the end of every report for scenarios, because that is the only part that can be verified in the future. I keep that practice now.

Worst case, which I put at roughly 45 percent: the silent tier persists for another three to five years. No tracking system is installed at the national championship. Analysis stays at the level of four error codes, as I did in Bac Giang. In that case what worries me is not the quality of analysis but the quality of decisions: coaches will keep making substitution, tactics and scheduling decisions based on feel, and feel at the silent tier is dominated by the most recent result — an effect analysts call recency bias.

Neutral case, roughly 35 percent: a few national events begin to have fixed cameras and one dedicated recorder. Group one and part of group two are filled in. This is the level at which I can build a Markov point-transition model per player, meaning I can answer the question: from a state of 18-16, what is this player's probability of winning the game, and how far does it deviate from the tournament baseline. That is the minimum threshold for analysis to become useful.

Best case, roughly 20 percent: a shuttle-trajectory tracking system is installed at a national-level event, possibly through an equipment sponsorship. Group three opens. And when group three opens, the first thing I will do is not player analysis. The first thing I will do is measure the execution cost of technique at 34 degrees Celsius versus 25, to answer a question I have carried for years: what share of the gap between a Vietnamese player and a top-30 player is a gap in skill, and what share is a gap in execution conditions.

Those three scenarios add to one hundred percent, and I know that the sum of three numbers I invented proves nothing. I publish them for one reason only: a prediction with a number attached will be verified, while a judgement with no number attached will live forever in the speaker's comfort zone.

Forty pages of report dead and silent in a stadium with no applause

I once presented a forty-page report to a club's leadership during a pandemic season. They asked me one question: so how do we win. Then they set it aside. That season the club won exactly one home match. I do not retell this to prove I was right. I retell it because it is why I never present a report made only of numbers anymore.

This time was different. The 6,423-word report in Bac Giang was not set aside because it was long. It was set aside because it had no answer. And it had no answer because the data does not exist, not because the analyst was lazy. This is a new kind of failure in my profession, and I must admit it cost me more sleepless nights than my actual mistakes.

At 56, I have stopped believing in numbers — but I believe in the ways numbers get betrayed. This report is one example of a number betrayed by its own absence. In a stadium with spectators, applause and commentators, I still cannot state the smash speed of the 118th point-ending stroke. That is worth saying, and it is worth saying more than the result itself.

Signals for the next cycle

The twenty-two-year-old I charted in Bac Giang will enter next season with a competence record of twelve lines, the most important of which is an end-game deviation coefficient of 2.34. I do not know whether he will win or lose. I know exactly what I will track: whether that coefficient falls below 1.5 within six months, and if it does not, whether anyone at the silent tier will pay to find out why.

For years I have used one line to remind myself of my profession's limits: numbers do not lie, but the people who read them lie to themselves for a lifetime. In Bac Giang both halves of that sentence were true, and I do not know which half kept me awake longer.

One thing I am certain of, after more than twenty years of sitting in the fourth row with an A5 notebook. The distance between a badminton nation with data and one without does not lie in the rankings. It lies here: one side knows where it went wrong and fixes it tomorrow, while the other knows where it went wrong but must wait three years to see it with its own eyes. Those three years are the real gap, and no coach can close it through personal effort alone.

Cầu thủ liên quan