Target object reconnaissance method, electronic device, storage medium and program product
Through the multi-view coordinated reconnaissance method of high-altitude and ground acquisition devices, the use of time-space synchronization calibration and object recognition model, the problem of single drone reconnaissance perspective and information sharing in the game is solved, and efficient and accurate reconnaissance of target objects in the virtual world is achieved.
Patent Information
- Application Number
- CN202510639039.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2045-05-19
AI Technical Summary
In the prior art, the drone reconnaissance in the game has the problems of a single reconnaissance perspective, limited coverage, and it is difficult to achieve efficient sharing and real-time interaction of air-ground equipment information, resulting in low reconnaissance efficiency and difficult to meet the target reconnaissance needs in complex environments.
The multi-view coordinated reconnaissance method of high-altitude and ground acquisition devices is adopted, and the object recognition model is optimized by space-time synchronization calibration processing and object recognition model, combined with environmental perception data, to achieve efficient fusion of data from multiple devices and precise positioning of target objects.
It improves the efficiency and accuracy of target object reconnaissance in the virtual world, solves the limitations of a single perspective, enhances the ability to distinguish interfering objects, and ensures efficient identification and positioning of target objects.
Smart Images

Figure CN120154898B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and more specifically, to a target object detection method, electronic equipment, storage medium, and program product. Background Art
[0002] In gaming, scouting objects is crucial for strategic decision-making. While various techniques exist, they still face challenges such as low scouting efficiency, limiting players' ability to accurately control and quickly respond to the battlefield. Therefore, a more efficient scouting method is urgently needed to enhance the gaming experience. Summary of the Invention
[0003] The purpose of the embodiments of the present application is to provide a target object detection method, electronic device, storage medium and program product, so as to achieve the technical effect of efficiently detecting target objects in a virtual world.
[0004] A first aspect of an embodiment of the present application provides a target object detection method, the method comprising:
[0005] Acquire a first object visual data set collected by an aerial acquisition device on a preset operation path in a virtual world, and acquire a second object visual data set collected by a ground acquisition device on the operation-related objects on the preset operation path; wherein the operation-related objects include disappearable interference objects and disappearable target objects; each piece of first object visual data in the first object visual data set and each piece of second object visual data in the second object visual data set carry position information, and the position information is used to represent the position of the operation-related objects in the virtual world when they are not disappearing;
[0006] Performing spatiotemporal synchronization calibration processing on the first object visual data set and the second object visual data set to obtain a processed visual data set;
[0007] inputting the processed visual data set into a trained object recognition model to obtain an object recognition result, and if the object recognition result indicates that the target object is recognized, determining target visual data including the target object from the processed visual data set;
[0008] The position information of the target object is determined based on the position information carried by the target visual data.
[0009] In the above implementation process, the use of multiple perspectives of different acquisition devices to collect visual data can compensate for the limitations of a single reconnaissance perspective, thereby fully capturing the characteristics of the target object. The use of spatiotemporal synchronization calibration processing eliminates visual data errors caused by differences in the acquisition time and position of different acquisition devices, making subsequent recognition more accurate. The use of object recognition models to build unified data interfaces and data standards enables efficient fusion of data collected by multiple types of devices, improving the overall efficiency and accuracy of visual data processing. Furthermore, the object recognition model can accurately distinguish between targets and interfering objects, and combined with the target visual data position information, the target object can be accurately located, providing a reliable basis for virtual world reconnaissance and achieving efficient reconnaissance of target objects in the virtual world.
[0010] Furthermore, before obtaining the first object visual data set collected by the high-altitude collection device on the preset work path in the virtual world, the method further includes:
[0011] The environmental data of the area to be operated is input into the trained path generation model to obtain the preset operation path; the environmental data includes the topographic data of the area to be operated, the obstacle distribution data, and the morphological data of the target object.
[0012] In the above implementation process, the optimal operation path is generated based on environmental characteristics (terrain, obstacles, target shape), so that the collection device can efficiently cover the potential target area, avoid invalid data collection, and improve reconnaissance efficiency.
[0013] Furthermore, before obtaining the first object visual data set collected by the high-altitude collection device on the preset work path in the virtual world, the method further includes:
[0014] Acquire first environmental perception data collected by the high-altitude acquisition device at a first acquisition angle at a first position on the preset operation path, and acquire second environmental perception data collected by the ground acquisition device at a second acquisition angle at a second position on the preset operation path;
[0015] Based on the first environmental perception data and the second environmental perception data, target collaborative operation parameters are determined, and the high-altitude acquisition device is controlled to collect the first object visual data set according to the target collaborative operation parameters, and the ground acquisition device is controlled to collect the second object visual data set according to the target collaborative operation parameters; wherein the target collaborative operation parameters are used to indicate the first target position and the first target acquisition angle of the high-altitude acquisition device, and the second target position and the second target acquisition angle of the ground acquisition device.
[0016] In the above implementation process, the spatial layout and collection angle of the collection equipment are dynamically optimized through environmental perception data, and a multi-perspective collaborative observation network is constructed to effectively solve the blind spot problem of a single perspective and improve the efficiency of target discovery.
[0017] Furthermore, the high-altitude collection device is equipped with a first environment perception sensor; the ground collection device is equipped with a second environment perception sensor; the first environment perception data is collected by the first environment perception sensor; the second environment perception data is collected by the second environment perception sensor;
[0018] The determining of target collaborative operation parameters based on the first environmental perception data and the second environmental perception data includes:
[0019] The first environmental perception data and the second environmental perception data are input into a trained operation parameter inference model to obtain the target collaborative operation parameters.
[0020] In the above implementation process, a collaborative operation parameter reasoning model is introduced to realize the intelligent dynamic adjustment of acquisition parameters and enhance the adaptability and operation flexibility to different scenarios.
[0021] Furthermore, the target collaborative operation parameters include visual range parameters, relative position parameters between the high-altitude acquisition device and the ground acquisition device, and formation parameters; the operation parameter inference model includes a visual range processing sub-model, a relative position processing sub-model, and a formation processing sub-model; the trained operation parameter inference model is trained by the following steps:
[0022] Acquire a multi-task training data set; the multi-task training data set includes a visual range training data subset, a relative position training data subset, and a formation training data subset; the visual range training data subset includes a one-to-one correspondence between a visual range image and a visual range label; the relative position training data subset includes a one-to-one correspondence between acquisition device position data and a relative position label; the formation training data subset includes a one-to-one correspondence between the acquisition device position data and a formation label; the acquisition device position includes the position data of the high-altitude acquisition device and the position data of the ground acquisition device;
[0023] The visual range processing sub-model is supervisedly trained using the visual range training data subset, the relative position processing sub-model is supervisedly trained using the relative position training data subset, and the formation formation processing sub-model is supervisedly trained using the formation formation training data subset to obtain the trained operation parameter inference model. The trained operation parameter inference model is used to generate the target collaborative operation parameters including the visual range parameters, the relative position parameters, and the formation formation parameters.
[0024] In the above implementation process, a multi-task collaborative training strategy is adopted to enable the model to have multi-objective optimization capabilities, significantly improving the accuracy of parameter reasoning.
[0025] Furthermore, the method further comprises:
[0026] When the object recognition result indicates that the target object is not recognized, an updated multi-task training data set including historical reconnaissance data is obtained, and the trained operation parameter inference model is trained using the updated multi-task training data set to obtain the trained updated operation parameter inference model, and the step of obtaining a first object visual data set collected by the high-altitude acquisition device on the operation-related object on the preset operation path in the virtual world is returned to be executed until the high-altitude acquisition device and the ground acquisition device collect visual data including the target object according to the updated target collaborative operation parameters generated by the trained updated operation parameter inference model.
[0027] In the above implementation process, a closed-loop feedback mechanism is constructed to use historical reconnaissance data to continuously iterate and optimize the operation parameter inference model.
[0028] Furthermore, each piece of the first object visual data carries first acquisition time information, and each piece of the second object visual data carries second acquisition time information;
[0029] The performing spatiotemporal synchronization calibration processing on the first object visual data set and the second object visual data set to obtain a processed visual data set includes:
[0030] Based on the first acquisition time information and the second acquisition time information, sorting the plurality of first object visual data and the plurality of second object visual data according to acquisition time to obtain a time-calibrated data set;
[0031] For each first object visual data, determining target second object visual data having characteristic pixel correlation with the first object visual data from a plurality of second object visual data, and integrating the first object visual data with the target second object visual data in pixel space based on the characteristic pixel correlation to obtain a spatially fused data set;
[0032] The processed visual dataset is determined based on the spatiotemporal correlation between the temporally aligned dataset and the spatially fused dataset.
[0033] In the above implementation process, by constructing a spatiotemporally consistent dataset, the spatiotemporal continuity of target reconnaissance in dynamic virtual environments is significantly improved, and the problem of target feature fragmentation caused by asynchronous acquisition of multiple acquisition devices and perspective differences is effectively solved, enabling the object recognition model to obtain more complete spatiotemporal context information, thereby improving the accuracy of target detection.
[0034] According to a second aspect of an embodiment of the present application, an electronic device is provided, comprising:
[0035] processor;
[0036] a memory for storing processor-executable instructions;
[0037] Wherein, when the processor calls the executable instruction, any method described in the first aspect is implemented.
[0038] A third aspect of an embodiment of the present application provides a computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the steps of any of the methods described in the first aspect.
[0039] A fourth aspect of the embodiments of the present application provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements any method described in the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without creative work.
[0041] Figure 1 A flowchart of a target object detection method provided in an embodiment of the present application;
[0042] Figure 2This is a structural block diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0043] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application.
[0044] It should be noted that similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings. At the same time, in the description of this application, the terms "first", "second", etc. are only used to distinguish the description and should not be understood as indicating or implying relative importance.
[0045] In related technologies, drone reconnaissance in games typically relies on a single drone or a simple drone formation. This approach presents significant limitations when performing reconnaissance missions, such as a single reconnaissance perspective, which can result in targets being obscured or missed; and limited coverage, requiring frequent movement to cover larger areas, which reduces reconnaissance efficiency. Meanwhile, while ground-based reconnaissance equipment, such as robotic dogs, can confirm targets at close range and provide detailed information, their narrow ground field of view makes it difficult to grasp the overall environmental conditions. Furthermore, they are constrained by terrain and obstacles, resulting in insufficient flexibility. Furthermore, the coordinated control technology between aerial drones and ground-based equipment is relatively complex, involving multiple steps such as flight path planning, inter-device communication, and data fusion and analysis. The current lack of effective intelligent decision-making and control solutions hinders efficient information sharing and real-time interaction between airborne and ground-based equipment, making it difficult to meet the requirements for target reconnaissance and identification in complex environments or demanding missions. This severely restricts the effectiveness and efficiency of unmanned reconnaissance systems in real-world mission scenarios.
[0046] In response to any of the above-mentioned problems, the present invention provides a method for detecting a target object. Figure 1 , Figure 1 A flowchart of a target object detection method provided in an embodiment of the present application.
[0047] In this embodiment, the method includes:
[0048] Step S10: Acquire a first object visual data set collected by an aerial acquisition device on a preset operation path in the virtual world, and acquire a second object visual data set collected by a ground acquisition device on the operation-related objects on the preset operation path; wherein the operation-related objects include disappearable interference objects and disappearable target objects; each first object visual data set in the first object visual data set and each second object visual data set in the second object visual data set carry position information, and the position information is used to represent the position of the operation-related objects in the virtual world when they are not disappearing;
[0049] It should be noted that the application scenario of this embodiment includes a virtual game world.
[0050] In the virtual world, high-altitude data collection devices can correspond to the game engine's preset high-altitude cameras or flying units (such as drones), simulating a bird's-eye view and covering large scenes (such as towns and battlefields). Ground-based data collection devices can be robot dogs, ground cameras, player characters, or NPCs, providing a close-up perspective to capture details.
[0051] Preset operation paths can be planned by game AI or level designers, avoiding virtual obstacles (such as walls and rivers) based on navigation meshes and covering areas with high probability that target objects may appear.
[0052] Object visual data may be image data of an object, including pictures and videos of the object.
[0053] Task-related objects include destructive interference objects and destructive target objects. Destructive interference objects include dynamically generated special effects (explosions, smoke), temporary NPCs, or objects triggered by random events. Destructive target objects include mission-critical objects (such as fortresses, hidden treasure chests) or enemy units, which are only visible under specific conditions (such as time of day or player behavior) or can be destroyed and disappear at any time.
[0054] The position of work-related objects in the virtual world can be based on the coordinate system of the high-altitude acquisition device or the ground acquisition device, indirectly derived using the sensor pose and observation vector, or it can be directly determined by the three-dimensional geometric coordinates of the work-related objects in the global model coordinate system of the virtual scene to which they belong (such as the Unity world coordinate system or the Unreal scene coordinate system).
[0055] Step S20: performing spatiotemporal synchronization calibration processing on the first object visual data set and the second object visual data set to obtain a processed visual data set;
[0056] It should be noted that spatiotemporal synchronization calibration is necessary because, on the one hand, there may be a time axis misalignment between high-altitude and ground-based acquisition devices (i.e., the acquisition frame rates or triggering times of the high-altitude and ground-based acquisition devices may be out of sync, e.g., the high-altitude acquisition device acquires once every 2 seconds, while the ground-based acquisition device acquires one frame every 1 second). This results in a time offset for the same target in different acquisition devices. Ignoring this time offset and directly fusing the data collected by different acquisition devices may produce "ghost targets" (i.e., the same target is identified as multiple independent objects). On the other hand, the observation geometry between a high-altitude bird's-eye view and a ground-level view differs significantly. For example, the camera of a ground-based acquisition device may not be able to observe a target obscured by a building due to its viewing angle, while the high-altitude view of a high-altitude acquisition device may cause the target to be blurred due to its distance. Target positions that have not been spatially calibrated may be deviated (e.g., the target coordinates from the high-altitude view must be projected onto the ground to align with the ground data). On the other hand, the dynamic environment in the virtual world (such as sudden rainstorms and sandstorm effects) may temporarily reduce the visual data collection quality of a certain collection device (such as the camera of a ground collection device being blocked by rain, causing the image to be blurred), making it difficult to correctly integrate the visual data collected by different collection devices, thereby affecting subsequent target recognition.
[0057] As an example, time synchronization calibration is completed through timestamp alignment. Specifically, an acquisition timestamp or engine timestamp is attached to each frame of data. Then, the time axis of the high-altitude acquisition device is selected as the reference, and the timestamp of the visual data collected by the ground acquisition device is mapped to the reference time axis through linear interpolation to complete the time synchronization calibration.
[0058] Optionally, spatial synchronization calibration is completed through coordinate system conversion. Specifically: the screen coordinates of the high-altitude acquisition device are converted into world coordinates using the engine API, and then projected into the viewing cone of the ground acquisition device to achieve cross-viewing geometric alignment, thereby completing spatial synchronization calibration.
[0059] Optionally, spatial synchronous calibration is completed by using image registration (i.e., feature point matching). Specifically, feature points (such as corner points, edges, etc.) are extracted from the visual data of the first object and the visual data of the second object, and then the geometric transformation relationship between the two visual data is calculated through matching and optimization methods to complete the spatial synchronous calibration.
[0060] Step S30: inputting the processed visual data set into a trained object recognition model to obtain an object recognition result, and if the object recognition result indicates that the target object is recognized, determining target visual data including the target object from the processed visual data set;
[0061] Optionally, the virtual world refers to a virtual game world. The trained object recognition model uses a lightweight architecture (such as MobileNet) adapted to the real-time requirements of gaming devices. The model can integrate game metadata (such as an object label database) to improve recognition efficiency. The model can filter recognition results based on task requirements (such as "find and identify a red treasure chest") and extract spatial-temporal segments containing the target (for example, recording the frame number and location where the target appears).
[0062] It should be noted that the object recognition model can be a deep neural network model that can learn higher-level feature representations from images to achieve accurate object recognition. The object recognition results include recognition results for each object in the visual data.
[0063] In its implementation, deep learning and multi-task learning techniques are used to automatically detect, classify, and extract features of targets. These features can be identified based on their location, size, trajectory, and other attributes, and automatically categorized into different types of targets, such as people, vehicles, animals, equipment, or other specifically labeled objects. Furthermore, ground-based acquisition devices can conduct close-range confirmation reconnaissance based on preliminary identification results. This close-range confirmation by the ground-based acquisition device allows for more detailed observation and verification of target features, further enhancing the reliability and accuracy of target identification results.
[0064] Step S40: Determine the position information of the target object based on the position information carried by the target visual data.
[0065] It should be noted that the position information of the target object may refer to two-dimensional position information of the target object in the virtual world, or may refer to three-dimensional position information of the target object in the virtual world, and this embodiment does not impose any limitation on this.
[0066] In this embodiment, by utilizing the different viewing angles and acquisition positions of high-altitude and ground-based acquisition devices, information about the target object can be acquired from multiple dimensions, thereby increasing the richness and comprehensiveness of the data and facilitating improved reconnaissance accuracy. By performing spatiotemporal synchronization calibration on the visual data collected by different devices, errors caused by differences in acquisition time and spatial location can be eliminated, making subsequent object recognition and position determination more accurate. Analyzing the data set using the model's recognition capabilities can effectively distinguish between interfering objects and target objects, improving the accuracy and efficiency of identifying target objects. By determining the target object's position information based on the position information carried by the target's visual data, precise positioning of the target object in the virtual world is achieved.
[0067] Based on any of the above embodiments, before step S10, the method further includes:
[0068] The environmental data of the area to be operated is input into the trained path generation model to obtain the preset operation path; the environmental data includes the topographic data of the area to be operated, the obstacle distribution data, and the morphological data of the target object.
[0069] It is understandable that environmental data includes topographic data (topographic data is used to reflect the terrain undulations, distribution of mountains and rivers, and other information of the area to be operated, which helps to plan the flight or movement route of the collection device and avoid collision with terrain obstacles), obstacle distribution data (obstacle distribution data clarifies the position and shape of various obstacles in the area to be operated, so that path planning can avoid these obstacles and ensure the driving safety of the collection device), and target object morphological data (the target object morphological data includes the size, shape, appearance characteristics, etc. of the target object, which helps the path generation model to plan a path that is more conducive to collecting the target object's visual data according to the characteristics of the target object).
[0070] It should be noted that the trained path generation model performs computations and inferences based on input environmental data. The model may incorporate various algorithms and rules, such as map-based path search algorithms, path planning algorithms, and optimization algorithms that consider target object characteristics. By analyzing environmental data, the model generates one or more predefined operational paths suitable for both aerial and ground-based collection devices operating in the virtual world. These predefined operational paths indicate the flight paths and ground deployment plans for the collaborative operation of the aerial and ground collection devices. These operational paths must also avoid terrain obstacles, circumvent obstacles, and approach the target object, ensuring that the collection devices operate safely and efficiently and quickly collect visual data containing the target object. Furthermore, the path planning process must consider the flight performance, endurance, and payload limitations of the aerial collection device, as well as the travel capabilities and endurance of the ground collection device, to ensure the planned path is both feasible and optimal.
[0071] Optionally, the spatial data of the area to be operated, obstacle information, and distribution characteristics of potential target objects are input into a model equipped with a path planning algorithm to develop an efficient and reasonable operation path.
[0072] In this embodiment, by inputting the environmental data of the area to be operated into a path generation model that has been specifically trained to obtain a preset operation path, it can ensure that the acquisition device can quickly, efficiently and safely collect visual data of the target object, thereby improving the accuracy and efficiency of the entire reconnaissance process.
[0073] Based on any of the above embodiments, before step S10, the method further includes:
[0074] Acquire first environmental perception data collected by the high-altitude acquisition device at a first acquisition angle at a first position on the preset operation path, and acquire second environmental perception data collected by the ground acquisition device at a second acquisition angle at a second position on the preset operation path;
[0075] It should be noted that the environmental perception data is used to reflect the environmental conditions at the current location and viewing angle of the collection device.
[0076] The first collection angle can be the initial collection angle of the high-altitude collection device. The first environmental perception data can include various environmental information of the working area at the first position and the first collection angle, such as topography, obstacle distribution, partial features of objects, etc. The first environmental perception data provides high-altitude perspective data for subsequent analysis and determination of collaborative operation parameters.
[0077] Similarly, the second acquisition angle can also be the initial acquisition angle of the ground acquisition device. The second environmental perception data supplements the environmental information of the work area from a ground perspective, including ground details, the relationship between objects and the ground, and the specific characteristics of objects. The second environmental perception data complements the first environmental perception data, together forming a comprehensive environmental perception of the work area.
[0078] Based on the first environmental perception data and the second environmental perception data, target collaborative operation parameters are determined, and the high-altitude acquisition device is controlled to collect the first object visual data set according to the target collaborative operation parameters, and the ground acquisition device is controlled to collect the second object visual data set according to the target collaborative operation parameters; wherein the target collaborative operation parameters are used to indicate the first target position and the first target acquisition angle of the high-altitude acquisition device, and the second target position and the second target acquisition angle of the ground acquisition device.
[0079] It's important to note that the target collaborative operation parameters indicate the target positions and acquisition angles of the two acquisition devices, enabling more efficient and accurate data collection. They also ensure that the two acquisition devices can collaborate during data collection, acquiring comprehensive and accurate visual data of task-related objects from different angles and positions, thereby improving data quality and usability.
[0080] Exemplarily, based on the first environmental perception data and the second environmental perception data, analysis and processing are performed through algorithms or models, and environmental information from both high-altitude and ground perspectives is comprehensively considered to calculate parameters that enable the high-altitude collection device and the ground collection device to work together to achieve the best effect, including the target position and collection angle of the two collection devices. These parameters can ensure that the two collection devices can perform comprehensive and complete collection of objects from different angles during the subsequent data collection process, avoiding repeated collection or collection blind spots.
[0081] In this embodiment, by analyzing and processing the environmental perception data collected by the high-altitude collection device and the ground collection device at different or the same positions and angles, determining the optimized collaborative operation parameters, and controlling the collection device to perform subsequent collection work with the optimized collaborative operation parameters, it can be ensured that the collection device collects the visual data of objects efficiently and without blind spots, and constructs a virtual environment with rich details and high restoration.
[0082] Based on any of the above embodiments, the high-altitude collection device is equipped with a first environment perception sensor; the ground collection device is equipped with a second environment perception sensor; the first environment perception data is collected by the first environment perception sensor; and the second environment perception data is collected by the second environment perception sensor;
[0083] The determining of target collaborative operation parameters based on the first environmental perception data and the second environmental perception data includes:
[0084] The first environmental perception data and the second environmental perception data are input into a trained operation parameter inference model to obtain the target collaborative operation parameters.
[0085] In a specific implementation, both the aerial and ground-based acquisition devices are equipped with sensors such as lidar, cameras, infrared sensors, and ultrasonic sensors. Each device transmits its sensor data, environmental data, and location data in real time to the operation parameter inference model. The operation parameter inference model dynamically calculates the target collaborative operation parameters for the two acquisition devices through real-time processing and inference. Different sensors have different functional characteristics and can perceive the operating environment from multiple dimensions. Specifically, lidar can accurately measure the distance and shape of objects; cameras can capture visual images of objects, facilitating object recognition and feature analysis; infrared sensors can detect infrared radiation emitted by objects and are suitable for sensing objects at night or in low-light environments; and ultrasonic sensors can be used for detecting close-range obstacles and measuring distances. For example, the operation parameter inference model uses algorithms and machine learning techniques to mine patterns and information contained in real-time data. For example, it identifies the position and posture of the target object based on lidar data and camera images, evaluates the current operating status of the acquisition device based on environmental data and location data, and calculates the target collaborative operation parameters between the aerial and ground-based acquisition devices through data processing and inference.
[0086] In this embodiment, comprehensive and accurate perception data is collected by environmental perception sensors, and with the help of an operation parameter inference model, target collaborative operation parameters can be efficiently and accurately analyzed from a variety of data, thereby realizing scientific collaborative operation between high-altitude collection devices and ground collection devices.
[0087] Based on any of the above embodiments, the target collaborative operation parameters include visual range parameters, relative position parameters between the high-altitude acquisition device and the ground acquisition device, and formation parameters; the operation parameter inference model includes a visual range processing sub-model, a relative position processing sub-model, and a formation processing sub-model; the trained operation parameter inference model is trained by the following steps:
[0088] Acquire a multi-task training data set; the multi-task training data set includes a visual range training data subset, a relative position training data subset, and a formation training data subset; the visual range training data subset includes a one-to-one correspondence between a visual range image and a visual range label; the relative position training data subset includes a one-to-one correspondence between acquisition device position data and a relative position label; the formation training data subset includes a one-to-one correspondence between the acquisition device position data and a formation label; the acquisition device position includes the position data of the high-altitude acquisition device and the position data of the ground acquisition device;
[0089] The visual range processing sub-model is supervisedly trained using the visual range training data subset, the relative position processing sub-model is supervisedly trained using the relative position training data subset, and the formation formation processing sub-model is supervisedly trained using the formation formation training data subset to obtain the trained operation parameter inference model. The trained operation parameter inference model is used to generate the target collaborative operation parameters including the visual range parameters, the relative position parameters, and the formation formation parameters.
[0090] It should be noted that the visual range training data subset includes a one-to-one correspondence between visual range images and visual range labels. A visual range image can refer to an image captured by the acquisition device in different scenarios, reflecting the area it can perceive, while a visual range label specifies the actual visual range information corresponding to the image, such as the specific acquisition angle range and acquisition coverage. The visual range training data subset is used to train the model to identify the visual range corresponding to different images, so that reasonable visual range parameters can be subsequently generated.
[0091] The relative position training data subset includes a one-to-one correspondence between acquisition device location data and relative position labels. The acquisition device location data records the specific locations of the high-altitude and ground-based acquisition devices at different times, while the relative position labels represent the relative positional relationship between the two acquisition devices, such as the distance. Using this relative position training data subset, the model can learn how to infer the relative positions of the acquisition devices based on their location data.
[0092] This embodiment does not limit the number of high-altitude and ground-based collection devices; there can be one or more high-altitude and one or more ground-based collection devices. The formation training data subset includes a one-to-one correspondence between the position data of the collection devices and formation labels. The formation labels describe the shape and layout of a formation composed of multiple collection devices, such as a linear formation, a triangular formation, or other specific formations. By learning from the data in the formation training data subset, the model can generate an appropriate formation based on the position of each collection device.
[0093] It is understandable that the operation parameter inference model includes but is not limited to a visual range processing sub-model, a relative position processing sub-model, and a formation formation processing sub-model. Each sub-model is responsible for processing and learning a corresponding portion of the data in the data set to achieve the generation of different types of parameters. The actual carrier of each sub-model can be expressed as an algorithm module or a model unit. Among them, the visual range processing sub-model learns the relationship between the visual range image and the corresponding label, and masters how to extract key features from the input image data and predict reasonable visual range parameters. The relative position processing sub-model can calculate the relative position parameters between them based on the real-time position data of the acquisition device. The formation formation processing sub-model can calculate reasonable formation formation parameters based on the position data of the acquisition device to ensure that multiple acquisition devices can work together in the best formation.
[0094] It should be understood that supervised training refers to the use of supervised learning methods to train each sub-model. Supervised learning means that during the training process, the model has known input data (such as visual range images, acquisition device position data) and corresponding expected outputs (such as visual range labels, relative position labels, and formation labels). By continuously adjusting the model parameters, the model output is made as close to the expected output as possible. After supervised training of each sub-model, a trained operation parameter inference model can be obtained. This model can comprehensively process the input first and second environmental perception data to generate target collaborative operation parameters including visual range parameters, relative position parameters, and formation parameters, thereby guiding the collaborative operation of high-altitude and ground-based acquisition devices and improving the efficiency and quality of data collection.
[0095] In this embodiment, a lightweight, modular model design reduces the computational burden and latency of the inference process, enhancing real-time data processing and decision-making response capabilities. Furthermore, by constructing a multi-task training dataset containing subsets of visual range, relative position, and formation training data, and conducting supervised training on the three sub-models for visual range processing, relative position processing, and formation processing, a comprehensive operational parameter inference model capable of generating comprehensive target collaborative operational parameters is created. This model, when presented with environmental perception data transmitted by high-altitude and ground-based acquisition devices, can calculate appropriate visual range parameters (used to instruct the acquisition devices to clearly capture visual data of objects), relative position parameters (used to ensure that the acquisition devices maintain optimal relative positions to avoid blind spots and missed targets), and formation parameters (used to instruct efficient collaborative operations), thereby improving the comprehensiveness, accuracy, and collaborative efficiency of data acquisition by the acquisition devices in the virtual world.
[0096] Based on any of the above embodiments, the method further includes:
[0097] When the object recognition result indicates that the target object is not recognized, an updated multi-task training data set including historical reconnaissance data is obtained, and the trained operation parameter inference model is trained using the updated multi-task training data set to obtain the trained updated operation parameter inference model, and the step of obtaining a first object visual data set collected by the high-altitude acquisition device on the operation-related object on the preset operation path in the virtual world is returned to be executed until the high-altitude acquisition device and the ground acquisition device collect visual data including the target object according to the updated target collaborative operation parameters generated by the trained updated operation parameter inference model.
[0098] It should be noted that in the reconnaissance operation in the virtual world, the target object has been included in the preset operation path during the path planning stage. Under normal circumstances, the target object can be identified. However, if the target object cannot be identified, it may be because the target collaborative operation parameters currently set for the acquisition device are not reasonable. Based on the unreasonable parameters, it takes too long for the acquisition device to go to the location of the target object to collect data. In the process of the acquisition device moving along the preset operation path based on the unreasonable target collaborative operation parameters, the target object may disappear due to various factors. For example, the target object disappears due to the explosion special effects in the virtual world, or disappears from the field of view due to the switching of scene special effects. In view of this, it is necessary to optimize the operation parameter inference model so that the target collaborative operation parameters generated by the model are more in line with the actual operation needs, thereby reducing or eliminating the situation where the acquisition device fails to collect the visual data of the target object before it disappears.
[0099] Specifically, when the target object is not identified, an updated multi-task training dataset containing historical reconnaissance data is obtained. This historical reconnaissance data records past acquisitions by the acquisition device under different operating parameters, as well as information about the target object and environment in the corresponding scene. This data helps the model learn more appropriate parameter settings for various complex situations, particularly those where the target object is prone to disappearing. Next, the trained operation parameter inference model is retrained using the updated multi-task training dataset. During training, the model adjusts its internal parameters and structure based on the new data. For example, historical data is analyzed to identify parameters such as the optimal relative position, visual range, and formation of the acquisition devices in scenarios where the target object is likely to disappear, enabling the model to generate more appropriate target collaborative operation parameters. Once model training is complete and the updated operation parameter inference model is obtained, the process returns to the step of acquiring the first set of visual data collected by the high-altitude acquisition device on the task-related object along the preset operation path. At this point, the acquisition device navigates the virtual world based on the updated target collaborative operation parameters. These optimized parameters increase the likelihood that the acquisition device will capture visual data of the target object before it disappears. The above-mentioned identification, optimization, and collection processes are continuously repeated. If the target object is still not identified, the training data set and training model are updated again until the collection device successfully collects visual data containing the target object according to the updated target collaborative operation parameters, thus completing the reconnaissance mission.
[0100] In its implementation, the object recognition model provides real-time feedback on current reconnaissance results, continuously optimizing the subsequent deployment paths and reconnaissance angles of both ground-based and high-altitude acquisition devices. Leveraging real-time feedback and reinforcement learning, the path generation model and operational parameter inference model rapidly analyze and evaluate information gathered during the reconnaissance mission, dynamically adjusting the mission path, operational angle, and formation. Through continuous iterative optimization, it can adaptively address uncertainties and emergencies in complex environments, ensuring the continued and effective advancement of the mission while effectively reducing the need for human intervention and improving reconnaissance efficiency.
[0101] In this embodiment, by integrating historical reconnaissance data into the generation of an updated multi-task training data set, the operation parameter inference model is optimized, prompting it to generate more suitable target collaborative operation parameters, thereby increasing the probability of the acquisition device collecting its visual data before the target object disappears, reducing acquisition delays and other problems, and thus ensuring the completion of the reconnaissance mission.
[0102] Based on any of the above embodiments, each piece of the first object visual data carries first acquisition time information, and each piece of the second object visual data carries second acquisition time information;
[0103] The performing spatiotemporal synchronization calibration processing on the first object visual data set and the second object visual data set to obtain a processed visual data set includes:
[0104] Based on the first acquisition time information and the second acquisition time information, sorting the plurality of first object visual data and the plurality of second object visual data according to acquisition time to obtain a time-calibrated data set;
[0105] For each first object visual data, determining target second object visual data having characteristic pixel correlation with the first object visual data from a plurality of second object visual data, and integrating the first object visual data with the target second object visual data in pixel space based on the characteristic pixel correlation to obtain a spatially fused data set;
[0106] The processed visual dataset is determined based on the spatiotemporal correlation between the temporally aligned dataset and the spatially fused dataset.
[0107] It's important to note that the purpose of time-based sorting is to ensure that visual data from different acquisition devices are consistent in their temporal order, facilitating subsequent processing and analysis. By sorting by time, we can clearly understand the data collected at different times and avoid data processing errors caused by time sequence confusion.
[0108] Feature pixel correlation refers to the similarity of pixel features related to key scene elements in two or more visual data sets, such as the shape, color, and texture of an object. By searching for visual data of a second object with feature pixel correlation, we can ensure that we select visual data that spatially corresponds to and complements the first object's visual data.
[0109] Alternatively, techniques such as image stitching and feature fusion can be used to merge two related visual data sets at the pixel level. This allows the processed data to fully reflect information about the object and scene, achieving spatial fusion and overcoming the limitations of using only high-altitude or ground-based acquisition devices. Furthermore, data integration is performed based on the spatiotemporal connections between the data sets to produce a processed visual dataset.
[0110] In its implementation, high-precision image registration technology and data fusion algorithms are used to precisely align and fuse multi-view data across multiple devices and perspectives, capturing multi-angle images or video data from both aerial and ground-based acquisition devices. This process involves calibrating the multi-view data in both time and space, eliminating time differences and spatial position errors, and performing information compensation and enhancement, resulting in a more comprehensive and accurate processed visual dataset.
[0111] In this embodiment, by processing and calibrating the two dimensions of time and space, the spatiotemporal synchronization calibration of the first object visual data set and the second object visual data set is achieved, which improves the quality of the visual data and enhances the accuracy and reliability of the reconnaissance of target objects in the virtual world.
[0112] Based on the method described in any of the above embodiments, the present application also provides Figure 2 A schematic diagram of the structure of an electronic device is shown in FIG. Figure 2 At the hardware level, the electronic device includes a processor, an internal bus, a network interface, memory, and non-volatile storage, and may also include other hardware required for its services. The processor reads the corresponding computer program from the non-volatile storage into the memory and then runs it to implement the method described in any of the above embodiments.
[0113] Based on the method described in any of the above embodiments, the present application also provides a computer storage medium, which stores a computer program. When the computer program is executed by a processor, it can be used to execute the method described in any of the above embodiments.
[0114] Based on the method described in any of the above embodiments, the present application further provides a computer program product comprising one or more computer programs or instructions. The computer program or instructions may be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another. When the computer program is executed by a processor, the method described in any of the above embodiments is implemented.
[0115] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions and operations of the devices, methods and computer program products according to the multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of the code, and the module, program segment or a part of the code contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or can be implemented using a combination of dedicated hardware and computer instructions.
[0116] In addition, the functional modules in each embodiment of the present application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0117] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard drives, read-only memories (ROM), random access memories (RAM), magnetic disks or optical disks.
[0118] The foregoing is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included within the scope of protection of the present application. It should be noted that similar reference numerals and letters represent similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined or explained in subsequent figures.
[0119] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
[0120] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.
Claims
1. A target object detection method, characterized in that: The method comprises: Acquire first environmental perception data collected by a high-altitude acquisition device on a preset operation path in the virtual world, and acquire second environmental perception data collected by a ground acquisition device on the preset operation path; The first environmental perception data and the second environmental perception data are input into a trained operation parameter inference model to obtain target collaborative operation parameters; the operation parameter inference model includes a visual range processing sub-model, a relative position processing sub-model, and a formation formation processing sub-model; the target collaborative operation parameters include visual range parameters, relative position parameters between the high-altitude acquisition device and the ground acquisition device, and formation formation parameters; the trained operation parameter inference model is trained by the following steps: obtaining a multi-task training data set; the multi-task training data set includes a visual range training data subset, a relative position training data subset, and a formation formation training data subset; the visual range training data subset includes a one-to-one corresponding visual range image and a visual range label; the relative position training data subset includes One-to-one corresponding acquisition device position data and relative position labels; the formation formation training data subset includes one-to-one corresponding acquisition device position data and formation formation labels; the acquisition device position includes the position data of the high-altitude acquisition device and the position data of the ground acquisition device; the visual range processing sub-model is supervised trained using the visual range training data subset, the relative position training data subset is supervised trained on the relative position processing sub-model, and the formation formation training data subset is supervised trained on the formation formation processing sub-model to obtain the trained operation parameter inference model, and the trained operation parameter inference model is used to generate the target collaborative operation parameters including the visual range parameters, the relative position parameters, and the formation formation parameters; Acquire a first object visual data set collected by the high-altitude acquisition device on the preset operation path according to the target collaborative operation parameters, and acquire a second object visual data set collected by the ground acquisition device on the preset operation path according to the target collaborative operation parameters; wherein the operation-related objects include disappearable interference objects and disappearable target objects; each first object visual data set in the first object visual data set and each second object visual data set in the second object visual data set carry position information, and the position information is used to represent the position of the operation-related object in the virtual world when it has not disappeared; Performing spatiotemporal synchronization calibration processing on the first object visual data set and the second object visual data set to obtain a processed visual data set; inputting the processed visual data set into a trained object recognition model to obtain an object recognition result, and if the object recognition result indicates that the target object is recognized, determining target visual data including the target object from the processed visual data set; The position information of the target object is determined based on the position information carried by the target visual data.
2. The method according to claim 1, wherein Before acquiring the first object visual data set collected by the high-altitude collection device on the preset operation path according to the target collaborative operation parameters, the method further includes: The environmental data of the area to be operated is input into the trained path generation model to obtain the preset operation path; the environmental data includes the topographic data of the area to be operated, the obstacle distribution data, and the morphological data of the target object.
3. The method according to claim 1, wherein The method further comprises: When the object recognition result indicates that the target object is not recognized, an updated multi-task training data set including historical reconnaissance data is obtained, and the trained operation parameter inference model is trained using the updated multi-task training data set to obtain the trained updated operation parameter inference model, and the step of obtaining the first object visual data set collected by the high-altitude acquisition device on the operation-related objects according to the target collaborative operation parameters on the preset operation path is returned to execute until the high-altitude acquisition device and the ground acquisition device collect visual data including the target object according to the updated target collaborative operation parameters generated by the trained updated operation parameter inference model.
4. The method according to claim 1, wherein Each piece of the first object visual data carries first acquisition time information, and each piece of the second object visual data carries second acquisition time information; The performing spatiotemporal synchronization calibration processing on the first object visual data set and the second object visual data set to obtain a processed visual data set includes: Based on the first acquisition time information and the second acquisition time information, sorting the plurality of first object visual data and the plurality of second object visual data according to acquisition time to obtain a time-calibrated data set; For each first object visual data, determining target second object visual data having characteristic pixel correlation with the first object visual data from a plurality of second object visual data, and integrating the first object visual data with the target second object visual data in pixel space based on the characteristic pixel correlation to obtain a spatially fused data set; The processed visual dataset is determined based on the spatiotemporal correlation between the temporally aligned dataset and the spatially fused dataset.
5. An electronic device, characterized in that: The electronic device comprises: processor; a memory for storing processor-executable instructions; Wherein, when the processor calls the executable instruction, it implements the method described in any one of claims 1-4.
6. A computer-readable storage medium, characterized in that Computer instructions are stored thereon, and when the computer instructions are executed by a processor, the steps of the method according to any one of claims 1 to 4 are implemented.
7. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.
Citation Information
Patent Citations
Image processing method and device, and storage medium
CN108022301A
Machine vision-based detecting and processing of table game events
TW202449723A