A drone programmed by GPT-6 Astra found the right person all by itself

Drone-Bench tests whether AI can write and debug code to control a drone via a software interface, rather than operating on board the drone itself.

0

During an experiment conducted by Andon Labs, OpenAI’s GPT-6 Astra model independently generated code for a drone which, in a controlled office environment, located a specified person using a photograph and followed them. According to the laboratory, in one of the trials the system outperformed the human baseline result at all stages of the Drone-Bench test; however, its typical reliability remains low.

Briefly about the main points

  • Astra has developed code to track and monitor a person using a drone.
  • The model achieved scores higher than the human baseline on five Drone-Bench tasks at least once.
  • On average, the entire sequence of tasks was completed with a probability of 2.8%.
  • The test checks that the code works via the drone’s interface, rather than via the AI inside the drone.
  • Andon Labs plans to make the scenarios more complex and assess the associated risks.

The model wrote the code for a flight in an office scenario

The researchers set out to Astra The task was to locate a specific person and follow them. The model was designed to convert data from the camera and available software tools into operational control: the drone moved around the office, identified the target using a reference image and kept it in its field of view.

This is not about autonomous intelligence built into a quadcopter. In Drone-Bench, an agent creates and corrects the code, whilst the drone executes it via a programming interface.

What exactly does Drone-Bench measure?

The benchmark breaks down the end-to-end task into five technical components. According to Andon Labs, Astra became the first model to exceed the human baseline score in each of these components in at least one run. This baseline is not a comparison with professional surveillance systems or industrial drones.

In the test, an agent can submit up to ten versions of a solution, receive feedback and refine the code. This rapid feedback makes the environment more conducive to debugging than an unprepared real-world setting.

  • Refurbishment: creation of a 3D model of the office and a map of obstacles.
  • Localisation: determining the drone’s position on the map that has been created.
  • Navigation: planning a route to the desired location.
  • Identification: searching for a person using a reference photograph.
  • Escort: keeping a person in sight whilst on the move.

A single success does not mean sustained autonomy

On average, an Astra run had only a 2.8% probability of completing the entire sequence from start to finish. The most challenging components remained environment reconstruction and reliable human detection. An error in the 3D map can give a false impression of obstacles and hinder safe route planning.

Therefore, the best result demonstrates the model’s potential, but not its readiness to operate reliably without human supervision. The test was carried out in an office using a low-cost drone; it would be incorrect to extrapolate the findings to real-world surveillance systems or military applications.

The laboratory offers to assess risks prior to the roll-out of systems

Andon Labs has created Drone-Bench to monitor at an early stage how language models acquire the ability to make physical observations. Researchers warn that AI access to physical devices may give rise to risks of misuse, and the limits of acceptable use should be discussed not only by laboratories.

The developers plan to make Drone-Bench more complex and closer to real-world scenarios.

 

WRITE A REPLY

enter your comment!
enter your name here