A boxing robot on a Bracket Bot, driven from a web dashboard and posed by human moves seen through a camera.
by Chris Chen Sep 2026 to Sep 2026A hackathon project: a Bracket Bot that boxes. You can drive it from a web dashboard, and it copies human boxing moves seen by a laptop camera. Pose detection runs in the browser, and the Python server turns the moves into arm motion.
We were at the BracketBot workshop staring at a 16-DOF robot with two 7-joint arms and a meter-long telescoping mast on each side, and someone said "it kinda looks like a boxer". That was the whole origin. The arms were already the right reach (about 0.72 m), nearly the same as a human arm and the masts could drop the entire torso almost a full meter. If we could get it to mirror a person's arms in real time, we'd basically have a sparring partner.
Most teleoperation feels like remote control; we wanted embodiment. Instead of pushing joysticks and waiting for a reaction, your body becomes the interface. Raise your fist, the robot raises its fist. Duck, it ducks. Just a webcam, near-zero latency, and complete 1:1 movement.
BigFluffyRobot turns a BracketBot into a remotely operated boxing robot controlled entirely through body tracking:
The whole stack is Python on the backend and vanilla JS in the browser, with zero framework dependencies.
Kinematics from scratch. The robot's arms are not anthropomorphic: seven joints, none of which line up with a human shoulder-elbow-wrist. So there's no joint-to-joint mapping. We wrote a damped-least-squares IK solver in pure numpy that takes a target position in 3D space and finds joint angles to reach it. One FK call through the arm chain takes 0.14 ms, and a full IK solve for both arms takes about 12 ms, well within the 33 ms budget for 30 Hz.
Limb-for-limb retargeting. The first version just mapped the wrist position, scaled by arm length. The robot never fully extended; it topped out at about 70% of its reach. The fix was to stop thinking about endpoints and start thinking about limb segments. We take the direction of each human limb segment (upper arm, forearm) and lay them end-to-end using the robot's limb lengths. A target built from the robot's own geometry is inside its workspace by construction. This got us from 70% extension to 98%, and it made the robot actually look like it was copying the pose rather than just reaching for the same point.
Elbow tracking. Hand position alone doesn't pin down an arm's shape, since you can swing your elbow through a wide arc without moving your hand. We added the operator's elbow direction as a secondary IK objective, weighted below the hand and only active for the first few solver iterations. The hand converges to within a millimeter; the elbow follows within about 5 degrees on average. When the two objectives conflict, the hand wins, because the hand is what lands the punch.
Torso twist from shoulder width. To turn the base, we measure how much your shoulder line has foreshortened in the camera. A shoulder line rotating about the vertical foreshortens as cos(yaw), so acos(apparent_width / calibrated_width) recovers the twist angle from a plain 2D measurement, with no depth sensor and no depth estimate. It's accurate to within a thousandth of a radian up to 60 degrees. The sign (which way you're twisted) comes from which shoulder is nearer the camera. The base holds the heading with a proportional controller on its own odometry, so your shoulders and the robot always agree.
Pose detection in the browser. MediaPipe runs in the browser, not in Python. No pip install mediapipe, no OpenCV camera permissions, no CDN dependency at demo time (we vendor the model files). The browser sends landmarks over HTTP; retargeting and IK happen in Python, keeping kin.py as the single source of truth for the robot's geometry.
The reach problem. Our first retargeting approach anchored the arm at the shoulder link, which is actually the mast carriage, 17 cm away from the joint the arm pivots about. The reach envelope was measured in that wrong frame too. Result: a fully extended human arm produced a robot arm at barely half extension, and the last third of your arm travel did nothing. We didn't realize this for hours because the IK error readout was also wrong (we were clamping to a sphere instead of the real reachable envelope). Fixing the anchor point and measuring the actual non-spherical envelope in 128 directions solved both problems at once.
The base never turned. The dashboard appends a {"v":0,"w":0} idle drive command at the end of every POST batch, after the pose data. When each input wrote the base velocity from its own handler, the trailing zero command wiped the torso's turn command thirty times a second. Puppet mode worked perfectly for the arms; the base just never moved. The fix was centralizing all base control into one mixer function called from the 60 Hz control loop, with the keyboard winning ties.
Gripper chatter. Fingertip tracking near the open/close boundary made the gripper rattle. We added hysteresis (a latch with separate close and open thresholds) and dropped the hand model to 6 Hz instead of every frame. Fist detection doesn't need 30 fps, and the pose model (which positions the arms) does.
The countdown problem. You can't press Record and be in your stance at the same time. Without a countdown, every recorded move starts with three seconds of a hand travelling back from the keyboard. We added a configurable countdown (3s / 5s / 10s) with a big number overlay on the camera feed readable from three meters back.
The retargeting actually feels right. There's a moment when you forget you're controlling a robot and it just feels like the robot is you. That only happens when the pose copying is good enough that you stop noticing it, and getting there meant solving a bunch of problems nobody warns you about. The robot's arms have seven joints and none of them correspond to a human joint. We wrote an IK solver from scratch, figured out limb-for-limb mapping instead of endpoint mapping, added elbow shaping so the arm looks right and not just reaches right, and got the whole thing running at 30 fps in pure numpy with no IK library. A straight punch extends to 98% of the robot's physical reach. That number was 70% twelve hours earlier.
Torso twist from a 2D camera. We're proud of the shoulder-foreshortening trick. MediaPipe's depth estimates are noisy enough that the first version (which read shoulder depth directly) produced maybe 3 degrees of base yaw from a 45-degree twist. The arccos(width / calibrated_width) approach uses only 2D measurements, recovers the angle to a thousandth of a radian, and needs no depth sensor at all. It's the kind of solution that makes you wonder why you tried the complicated thing first.
Zero dependencies beyond numpy. No FastAPI, no uvicorn, no ROS, no IK library, no ML framework on the server side. The HTTP server is stdlib. The IK is 15 lines of damped least squares. Pose detection runs in the browser so there's no pip install mediapipe. The whole thing vendors into ~20 MB and runs fully offline. When venue WiFi went down during demos, we didn't notice.
Record and replay that actually works on stage. The countdown overlay, the big camera-filling numbers, the fact that the recorder runs on the 60 Hz control loop instead of the browser clock: all of that exists because we tried to demo it once without those things and the first three seconds of every recording were a hand reaching for the keyboard. Small UX decisions like that are the difference between a demo that works and a demo that works in front of people.
Writing IK from scratch instead of reaching for a library was the right call. We understood every failure mode, and when the solver did something weird at 2am we could actually diagnose it. The hardest part wasn't the math. It was getting the frames right: which link is the shoulder, where does the arm actually pivot, what coordinate system are the landmarks in. Every bug in the retargeting came down to an anchor point or a sign convention, not the solver itself.
We also learned that the most dramatic feature on this robot is the mast. A 1-metre vertical drop is far more legible on camera than a wheeled base backing up, and it doesn't risk tipping a self-balancing robot. Lead with the duck.