Target object reconnaissance method, electronic equipment, storage medium and program product
By obtaining visual data of high-altitude and ground acquisition devices in the game world, performing time-spatial and spatial calibration and object recognition, the problem of low reconnaissance efficiency of target objects in the prior art is solved, and efficient and accurate reconnaissance of target objects in the virtual world is achieved.
Patent Information
- Application Number
- CN202510639039.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-05-19
AI Technical Summary
The existing technology has low reconnaissance efficiency of target objects in the game world, which limits players' precise control and rapid response capabilities of the overall battlefield.
By obtaining the visual data set on the preset job path of the high altitude and ground acquisition devices in the virtual world, and performing time-time synchronization calibration processing, input the trained object recognition model for identification, and determine the position information of the target object.
It realizes efficient reconnaissance of target objects in the virtual world. Through multi-view data acquisition and time-time synchronization calibration, the accuracy and efficiency of reconnaissance are improved, and the target objects can be accurately distinguished from interference objects, and precisely positioned the target objects.
Smart Images

Figure CN120154898A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology. Specifically, it relates to a method for detecting target objects, an electronic device, a storage medium, and a program product. Background Art
[0002] In the game world, detecting objects is crucial for strategic decision-making. Although there are already diverse means in related technologies, challenges such as low detection efficiency still exist, which limits the player's ability to accurately control the overall battlefield situation and respond quickly. Therefore, there is an urgent need for a more efficient detection method to improve the game experience. Summary of the Invention
[0003] The purpose of the embodiments of this application is to provide a method for detecting target objects, an electronic device, a storage medium, and a program product, so as to achieve the technical effect of efficiently detecting target objects in a virtual world.
[0004] The first aspect of the embodiments of this application provides a method for detecting target objects, and the method includes: Obtain a first set of object visual data collected by an aerial acquisition device for operation-related objects on a preset operation path in a virtual world, and obtain a second set of object visual data collected by a ground acquisition device for the operation-related objects on the preset operation path; wherein, the operation-related objects include disappearable interference objects and disappearable target objects; each piece of first object visual data in the first set of object visual data and each piece of second object visual data in the second set of object visual data carry position information, and the position information is used to represent the position of the operation-related objects in the virtual world when they are in an un-disappeared state; Perform spatio-temporal synchronization calibration processing on the first set of object visual data and the second set of object visual data to obtain a processed visual data set; Input the processed visual data set into a trained object recognition model to obtain an object recognition result, and when the object recognition result indicates that the target object is recognized, determine target visual data including the target object from the processed visual data set; Determine the position information of the target object based on the position information carried by the target visual data.
[0005] In the above implementation process, multi-perspective visual data is collected using different acquisition devices, which can make up for the limitations of a single reconnaissance perspective, thereby comprehensively capturing the characteristics of target objects; spatio-temporal synchronization calibration processing is used to eliminate visual data errors caused by differences in acquisition time and position of different acquisition devices, making subsequent recognition more accurate; an object recognition model is used to construct a unified data interface and data standard to achieve efficient fusion of data collected by multiple types of devices, improving the overall visual data processing efficiency and accuracy. In addition, based on the object recognition model, the target and interference objects can be accurately distinguished, and combined with the position information of the target visual data, the target object can be accurately located, providing a reliable basis for reconnaissance in the virtual world and realizing efficient reconnaissance of the target object in the virtual world.
[0006] Further, before obtaining the first set of object visual data collected for operation-related objects on the preset operation path of the high-altitude acquisition device in the virtual world, it further includes: Inputting the environmental data of the area to be operated into the trained path generation model to obtain the preset operation path; the environmental data includes the topographic and geomorphic data, obstacle distribution data, and morphological data of the target object in the area to be operated.
[0007] In the above implementation process, the optimal operation path is generated based on environmental characteristics (terrain, obstacles, target morphology), enabling the acquisition device to efficiently cover potential target areas, avoiding the collection of invalid data, and improving reconnaissance efficiency.
[0008] Further, before obtaining the first set of object visual data collected for operation-related objects on the preset operation path of the high-altitude acquisition device in the virtual world, it further includes: Obtaining the first environmental perception data collected by the high-altitude acquisition device at the first position on the preset operation path at the first acquisition angle, and obtaining the second environmental perception data collected by the ground acquisition device at the second position on the preset operation path at the second acquisition angle; Determining target collaborative operation parameters based on the first environmental perception data and the second environmental perception data, and controlling the high-altitude acquisition device to collect the first set of object visual data according to the target collaborative operation parameters, and controlling the ground acquisition device to collect the second set of object visual data according to the target collaborative operation parameters; wherein, the target collaborative operation parameters are used to indicate the first target position and the first target acquisition angle of the high-altitude acquisition device, and the second target position and the second target acquisition angle of the ground acquisition device.
[0009] In the above implementation process, the spatial layout and acquisition angle of the acquisition device are dynamically optimized through environmental perception data, and a multi-perspective collaborative observation network is constructed, effectively solving the problem of blind spots in a single perspective and improving the target discovery efficiency.
[0010] Furthermore, the high-altitude acquisition device is equipped with a first environmental perception sensor; the ground acquisition device is equipped with a second environmental perception sensor; the first environmental perception data is collected by the first environmental perception sensor; the second environmental perception data is collected by the second environmental perception sensor; Determining the target collaborative operation parameters based on the first environmental perception data and the second environmental perception data includes: Inputting the first environmental perception data and the second environmental perception data into a trained operation parameter inference model to obtain the target collaborative operation parameters.
[0011] In the above implementation process, a collaborative operation parameter inference model is introduced to realize the intelligent dynamic adjustment of acquisition parameters, enhancing the adaptability to different scenarios and the flexibility of operations.
[0012] Furthermore, the target collaborative operation parameters include visual range parameters, relative position parameters between the high-altitude acquisition device and the ground acquisition device, and formation parameters; the operation parameter inference model includes a visual range processing sub-model, a relative position processing sub-model, and a formation processing sub-model; the trained operation parameter inference model is trained through the following steps: Obtain a multi-task training data set; the multi-task training data set includes a visual range training data subset, a relative position training data subset, and a formation training data subset; the visual range training data subset includes corresponding visual range images and visual range labels; the relative position training data subset includes corresponding acquisition device position data and relative position labels; the formation training data subset includes corresponding acquisition device position data and formation labels; the acquisition device position includes the position data of the high-altitude acquisition device and the position data of the ground acquisition device; Use the visual range training data subset to perform supervised training on the visual range processing sub-model, use the relative position training data subset to perform supervised training on the relative position processing sub-model, and use the formation training data subset to perform supervised training on the formation processing sub-model to obtain the trained operation parameter inference model, and the trained operation parameter inference model is used to generate the target collaborative operation parameters including the visual range parameters, the relative position parameters, and the formation parameters.
[0013] In the above implementation process, a multi-task collaborative training strategy is adopted to enable the model to have the ability of multi-objective optimization, significantly improving the accuracy of parameter inference.
[0014] Furthermore, the method further includes: In the case that the object recognition result indicates that the target object is not recognized, obtain an updated multi-task training data set including historical reconnaissance data, and use the updated multi-task training data set to train the trained job parameter inference model to obtain a trained updated job parameter inference model. Return to execute the step of acquiring the first object visual data set collected by the aerial acquisition device for job-related objects on the preset job path in the virtual world until the aerial acquisition device and the ground acquisition device collect visual data including the target object according to the updated target collaborative job parameters generated by the trained updated job parameter inference model.
[0015] In the above implementation process, a closed-loop feedback mechanism is constructed to continuously iteratively optimize the job parameter inference model using historical reconnaissance data.
[0016] Further, each of the first object visual data carries first acquisition time information, and each of the second object visual data carries second acquisition time information; The step of performing spatio-temporal synchronization calibration processing on the first object visual data set and the second object visual data set to obtain a processed visual data set includes: Based on the first acquisition time information and the second acquisition time information, sort the multiple first object visual data and the multiple second object visual data according to the acquisition time to obtain a time-calibrated data set; For each of the first object visual data, determine a target second object visual data that has feature pixel correlation with the first object visual data from the multiple second object visual data, and integrate the first object visual data and the target second object visual data in the pixel space based on the feature pixel correlation to obtain a spatially fused data set; Based on the spatio-temporal correlation between the time-calibrated data set and the spatially fused data set, determine the processed visual data set.
[0017] In the above implementation process, by constructing a data set with spatio-temporal consistency, the spatio-temporal continuity of target reconnaissance in a dynamic virtual environment is significantly improved, effectively solving the problem of fragmented target features caused by asynchronous acquisition of multiple acquisition devices and perspective differences, enabling the object recognition model to obtain more complete spatio-temporal context information, and thus improving the accuracy of target detection.
[0018] A second aspect of the embodiments of the present application provides an electronic device, where the electronic device includes: A processor; A memory for storing instructions executable by the processor; Wherein, when the processor calls the executable instructions, the method described in any one of the first aspect is implemented.
[0019] In the third aspect of the embodiments of the present application, a computer-readable storage medium is provided, on which computer instructions are stored, and when the computer instructions are executed by a processor, the steps of any of the methods in the first aspect are implemented.
[0020] In the fourth aspect of the embodiments of the present application, a computer program product is provided, the computer program product includes a computer program, and when the computer program is executed by a processor, the method of any of the methods in the first aspect is implemented. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required to be used in the embodiments of the present application will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0022] Figure 1 It is a schematic flowchart of a method for detecting a target object provided by an embodiment of the present application; Figure 2 It is a structural block diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] The technical solutions in the embodiments of the present application will be described below with reference to the drawings in the embodiments of the present application.
[0024] It should be noted that: similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of the present application, the terms "first", "second", etc. are only used for distinguishing descriptions, and cannot be understood as indicating or implying relative importance.
[0025] In the related art, the drone reconnaissance in games usually relies on a single drone or a simple drone formation. This method has obvious limitations when performing reconnaissance tasks. For example, the reconnaissance perspective is single, resulting in the possibility that the target may be blocked or missed; the coverage area is limited, and frequent movement is required to cover a large area, reducing the reconnaissance efficiency. At the same time, although ground reconnaissance equipment such as a robotic dog can perform close-range target confirmation and provide detailed information, the ground vision is narrow, it is difficult to grasp the overall environmental conditions, and it is limited by terrain and obstacles, lacking flexibility. In addition, the cooperative control technology of aerial drones and ground equipment is relatively complex, involving multiple links such as flight path planning, communication between devices, and data fusion analysis. Currently, there is a lack of an effective intelligent decision-making control scheme, so it is difficult to achieve efficient sharing and real-time interaction of information between air and ground devices, and it is difficult to meet the requirements of target reconnaissance and identification in complex environments or high-demand tasks, seriously restricting the application effect and efficiency improvement of the unmanned reconnaissance system in actual mission scenarios.
[0026] In view of any of the above problems, an embodiment of the present application provides a method for reconnoitering a target object, referring to Figure 1 , Figure 1 which is a schematic flowchart of a method for reconnoitering a target object provided by an embodiment of the present application.
[0027] In this embodiment, the method includes: Step S10: Obtain a first set of object visual data collected by a high-altitude acquisition device for operation-related objects on a preset operation path in a virtual world, and obtain a second set of object visual data collected by a ground acquisition device for the operation-related objects on the preset operation path; wherein, the operation-related objects include disappearable interference objects and disappearable target objects; each first object visual data in the first set of object visual data and each second object visual data in the second set of object visual data carry position information, and the position information is used to represent the position of the operation-related objects in the virtual world in an un-disappeared state; It should be noted that the application scenario of this embodiment includes a virtual game world.
[0028] The high-altitude acquisition device in the virtual world can correspond to a high-altitude camera or a flying unit (such as a drone) preset by the game engine, which is used to simulate an overlooking perspective and cover a large range of scenes (such as a town, a battlefield). The ground acquisition device can be a robotic dog, a ground camera, a player character, or an NPC, which is used to provide a near-ground perspective and capture details.
[0029] The preset operation path can be planned by the game AI or the level designer, avoiding virtual obstacles (such as walls, rivers) based on the navigation mesh, and covering high-probability areas where the target object may appear.
[0030] The object visual data can be the image data of the object, including pictures and videos of the object.
[0031] The operation-related objects include disappearable interference objects and disappearable target objects. Among them, the disappearable interference objects can refer to, for example, dynamically generated special effects (explosions, smoke), temporary NPCs, or objects triggered by random events, etc. The disappearable target objects can refer to mission-critical objects (such as fortresses, hidden treasure chests) or enemy units, etc., which are only visible under specific conditions (such as time, player behavior) or will be blown up and disappear at any time.
[0032] The positions of the operation-related objects in the virtual world can be indirectly derived using the coordinate system of the high-altitude acquisition device or the ground acquisition device, based on the sensor pose and the observation vector, or directly determined by the three-dimensional geometric coordinates of the operation-related objects in the global model coordinate system of the virtual scene to which they belong (such as the Unity world coordinate system or the Unreal scene coordinate system).
[0033] Step S20: Perform spatio-temporal synchronization calibration processing on the first object visual data set and the second object visual data set to obtain a processed visual data set; It should be noted that the reason for performing spatio-temporal synchronization calibration processing is as follows: On the one hand, the high-altitude acquisition device and the ground acquisition device may have a time-axis misalignment (that is, the acquisition frame rates or trigger moments of the high-altitude acquisition device and the ground acquisition device may not be synchronized. For example, the high-altitude acquisition device acquires once every 2 seconds, and the ground acquisition device acquires one frame per second), which may lead to a time offset of the same target in different acquisition devices. Ignoring the time offset and directly fusing the data acquired by different acquisition devices may result in "ghost targets" (that is, the same target is recognized as multiple independent objects). On the other hand, the observation geometry of the high-altitude overlooking view and the ground-level view is significantly different. For example, the camera of the ground acquisition device may not be able to observe a target blocked by a building due to the viewing angle limitation, while the high-altitude view of the high-altitude acquisition device may make the target blurred due to the excessive distance. The position of the target without spatial calibration may deviate (for example, the target coordinates in the high-altitude view need to be projected onto the ground to align with the ground data). On the third hand, the dynamic environment in the virtual world (such as a sudden rainstorm, sandstorm special effects) may temporarily reduce the visual data acquisition quality of a certain acquisition device (such as the camera of the ground acquisition device being blocked by rain, resulting in a blurred picture), making it difficult to correctly fuse the visual data acquired by different acquisition devices, thus affecting subsequent target recognition.
[0034] As an example, time synchronization calibration is completed through timestamp alignment. Specifically: an acquisition timestamp or an engine timestamp is attached to each frame of data, and then, the time axis of the aerial acquisition device is selected as the reference, and the timestamps of the visual data acquired by the ground acquisition device are mapped to the reference time axis through linear interpolation, thereby completing the time synchronization calibration.
[0035] Optionally, spatial synchronization calibration is completed through coordinate system transformation. Specifically: the screen coordinates of the aerial acquisition device are converted into world coordinates using the engine API, and then projected into the frustum of the ground acquisition device to achieve cross-view geometric alignment, thereby completing the spatial synchronization calibration.
[0036] Optionally, image registration (i.e., feature point matching) is used to complete the spatial synchronization calibration. Specifically, feature points (such as corner points, edges, etc.) are extracted from the first object visual data and the second object visual data, and then, the geometric transformation relationship between the two visual data is calculated through matching and optimization methods, thereby completing the spatial synchronization calibration.
[0037] Step S30: Input the processed visual data set into the trained object recognition model to obtain the object recognition result, and when the object recognition result indicates that the target object is recognized, determine the target visual data including the target object from the processed visual data set; Optionally, the virtual world refers to a virtual game world, and the trained object recognition model adopts a lightweight architecture (such as MobileNet) to adapt to the real-time requirements of game devices. The model can combine game metadata (such as an object label database) to improve the recognition efficiency. The model can filter the recognition results according to the task requirements (such as "finding and recognizing a red treasure chest") and extract the spatio-temporal segments including the target (such as recording the frame numbers and positions where the target appears).
[0038] It should be noted that the object recognition model can be a deep neural network model, which can learn higher-level feature representations from images, thereby achieving accurate object recognition. The object recognition result includes the recognition results for each object in the visual data.
[0039] In specific implementation, the technologies of deep learning and multi-task learning are used to automatically detect, classify, and extract features of the target, and the attribute features such as the position, size, and motion trajectory of the target can be recognized, and different types of targets can be automatically classified, such as people, vehicles, animals, devices, or other objects with specific marks. At the same time, the ground acquisition device can conduct close-range confirmation and investigation based on the preliminary recognition result. Through the close-range confirmation of the ground acquisition device, the target features can be observed and verified more carefully, further enhancing the reliability and accuracy of the target recognition result.
[0040] Step S40: Determine the position information of the target object based on the position information carried by the target visual data.
[0041] It should be noted that the position information of the target object can refer to the two-dimensional position information of the target object in the virtual world, or the three-dimensional position information of the target object in the virtual world. This embodiment does not limit this.
[0042] In this embodiment, by using the different perspectives and acquisition positions of the high-altitude acquisition device and the ground acquisition device, information about the target object can be obtained from multiple dimensions, thereby increasing the richness and comprehensiveness of the data, which is beneficial to improving the accuracy of reconnaissance. By performing spatio-temporal synchronization calibration processing on the visual data collected by different devices, the errors caused by different acquisition times and spatial positions can be eliminated, making the subsequent object recognition and position determination more accurate. By using the recognition ability of the model to analyze the data set, interference objects and target objects can be effectively distinguished, improving the accuracy and efficiency of recognizing target objects. Determining the position information of the target object based on the position information carried by the target visual data realizes the precise positioning of the target object in the virtual world.
[0043] Based on any of the above embodiments, before step S10, it further includes: Input the environmental data of the area to be operated into the trained path generation model to obtain the preset operation path; the environmental data includes the topographic and geomorphic data of the area to be operated, the obstacle distribution data, and the morphological data of the target object.
[0044] It can be understood that the environmental data includes topographic and geomorphic data (the topographic and geomorphic data is used to reflect the terrain undulation, mountain and river distribution, etc. in the area to be operated, which helps to plan the flight or movement route of the acquisition device and avoid colliding with terrain obstacles), obstacle distribution data (the obstacle distribution data clarifies the positions and shapes of various obstacles in the area to be operated, enabling the path planning to avoid these obstacles and ensuring the driving safety of the acquisition device), and morphological data of the target object (the morphological data of the target object includes the size, shape, appearance characteristics, etc. of the target object, which helps the path generation model to plan a path that is more conducive to collecting the visual data of the target object according to the characteristics of the target object).
[0045] It should be noted that the trained path generation model performs operations and inferences based on the input environmental data. The model may contain various algorithms and rules, such as path search algorithms based on map information, path planning algorithms, and optimization algorithms considering the characteristics of target objects. Through the analysis of environmental data, the model can generate one or more preset operation paths suitable for the high-altitude acquisition device and the ground acquisition device to operate in the virtual world. The preset operation paths are used to indicate the flight paths and ground deployment plans for the coordinated operation of the high-altitude acquisition device and the ground acquisition device. At the same time, this operation path can take into account avoiding terrain and geomorphic obstacles, bypassing obstacles, and approaching the target object to ensure that the acquisition device can operate safely and efficiently collect visual data containing the target object during the operation process. In addition, the path planning process needs to consider the flight performance, endurance, and payload limitations of the high-altitude acquisition device, as well as the traveling ability and endurance time of the ground acquisition device to ensure that the planned path is practically feasible and optimal.
[0046] Optionally, the spatial data of the area to be operated, the obstacle information, and the distribution characteristics of potential target objects are input into the model equipped with the path planning algorithm to formulate an efficient and reasonable operation path.
[0047] In this embodiment, by inputting the environmental data of the area to be operated into the specially trained path generation model to obtain the preset operation path, it can ensure that the acquisition device quickly, efficiently, and safely collects the visual data of the target object, thereby improving the accuracy and efficiency of the entire reconnaissance process.
[0048] Based on any of the above embodiments, before step S10, it further includes: Obtain the first environmental perception data collected by the high-altitude acquisition device at the first position on the preset operation path at the first acquisition angle, and obtain the second environmental perception data collected by the ground acquisition device at the second position on the preset operation path at the second acquisition angle; It should be noted that the environmental perception data is used to reflect the environmental conditions at the current position and perspective of the acquisition device.
[0049] The first acquisition angle can be the initial acquisition angle of the high-altitude acquisition device. The first environmental perception data can include various environmental information of the operation area at the first position and the first acquisition angle, such as terrain and geomorphic features, obstacle distribution, and partial features of objects. The first environmental perception data provides data from a high-altitude perspective for subsequent analysis and determination of coordinated operation parameters.
[0050] Similarly, the second acquisition angle can also be the initial acquisition angle of the ground acquisition device. The second environmental perception data supplements the environmental information of the operation area from the ground perspective, including details of the ground, the relationship between objects and the ground, specific features of objects, etc. The second environmental perception data and the first environmental perception data complement each other and jointly constitute a comprehensive environmental perception of the operation area.
[0051] Determine target collaborative operation parameters based on the first environmental perception data and the second environmental perception data, and control the aerial acquisition device to collect the first set of object visual data according to the target collaborative operation parameters, and control the ground acquisition device to collect the second set of object visual data according to the target collaborative operation parameters; wherein, the target collaborative operation parameters are used to indicate the first target position and the first target acquisition angle of the aerial acquisition device, and the second target position and the second target acquisition angle of the ground acquisition device.
[0052] It should be noted that the target collaborative operation parameters are used to indicate the target positions and acquisition angles of the two acquisition devices to achieve more efficient and accurate data acquisition. At the same time, the target collaborative operation parameters are also used to ensure that the two acquisition devices can cooperate with each other when collecting data, and obtain comprehensive and accurate visual data about the operation-related objects from different angles and positions, so as to improve the quality and usability of the data.
[0053] Exemplarily, based on the first environmental perception data and the second environmental perception data, through algorithm or model analysis and processing, comprehensively considering the environmental information from both the aerial and ground perspectives, calculate the parameters that can make the aerial acquisition device and the ground acquisition device work together to achieve the best effect, including the target positions and acquisition angles of the two acquisition devices. These parameters can ensure that in the subsequent data acquisition process, the two acquisition devices can comprehensively and without omission collect each object from different angles, avoiding repeated acquisition or acquisition blind spots.
[0054] In this embodiment, by analyzing and processing the environmental perception data collected by the aerial acquisition device and the ground acquisition device at different or the same positions and angles, determining the optimized collaborative operation parameters, and controlling the acquisition devices to perform subsequent acquisition work according to the optimized collaborative operation parameters, it can ensure that the acquisition devices can efficiently and without dead angles collect object visual data and construct a virtual environment with rich details and high restoration.
[0055] On the basis of any of the above embodiments, the aerial acquisition device is equipped with a first environmental perception sensor; the ground acquisition device is equipped with a second environmental perception sensor; the first environmental perception data is collected by the first environmental perception sensor; the second environmental perception data is collected by the second environmental perception sensor; The determining of the target collaborative operation parameter based on the first environmental perception data and the second environmental perception data includes: The first environmental perception data and the second environmental perception data are input into a trained operation parameter inference model to obtain the target collaborative operation parameters.
[0056] In the specific implementation, both the high-altitude acquisition device and the ground acquisition device are equipped with sensors of the types of laser radar, camera, infrared sensor, ultrasonic sensor, etc. The high-altitude acquisition device and the ground acquisition device transmit their respective sensor data, environmental data, and position data to the operation parameter reasoning model in real time. The operation parameter reasoning model dynamically calculates the target collaborative operation parameters of the two acquisition devices through real-time processing and reasoning. Among them, different sensors have different functional characteristics and can perceive the working environment from multiple dimensions. Specifically, the laser radar can accurately measure the distance and shape of the object; the camera can capture the visual image of the object, which is convenient for object recognition and feature analysis; the infrared sensor can detect the infrared radiation emitted by the object, which is suitable for sensing objects at night or in low light environments; the ultrasonic sensor can be used for the detection of close-range obstacles and the measurement of distance. Exemplarily, the operation parameter reasoning model will use algorithms and machine learning techniques to mine the laws and information contained in real-time data. For example, the position and posture of the target object are identified based on the laser radar data and camera images, and the current working state of the acquisition device is evaluated in combination with the environmental data and position data. Through data processing and reasoning, the target collaborative operation parameters between the high-altitude and ground acquisition devices are calculated.
[0057] In this embodiment, comprehensive and accurate perception data is collected by environmental perception sensors, and with the help of an operation parameter inference model, target collaborative operation parameters can be efficiently and accurately analyzed from a variety of data, thereby realizing scientific collaborative operation between high-altitude collection devices and ground collection devices.
[0058] On the basis of any of the above embodiments, the target collaborative operation parameters include visual range parameters, relative position parameters between the high-altitude acquisition device and the ground acquisition device, and formation parameters; the operation parameter reasoning model includes a visual range processing sub-model, a relative position processing sub-model, and a formation processing sub-model; the trained operation parameter reasoning model is trained by the following steps: Obtain a multi-task training data set; the multi-task training data set includes a visual range training data subset, a relative position training data subset, and a formation training data subset; the visual range training data subset includes corresponding visual range images and visual range labels one by one; the relative position training data subset includes corresponding acquisition device position data and relative position labels one by one; the formation training data subset includes corresponding acquisition device position data and formation labels one by one; the acquisition device position includes the position data of the high-altitude acquisition device and the position data of the ground acquisition device; Use the visual range training data subset to perform supervised training on the visual range processing sub-model, use the relative position training data subset to perform supervised training on the relative position processing sub-model, and use the formation training data subset to perform supervised training on the formation processing sub-model to obtain the trained operation parameter inference model, and the trained operation parameter inference model is used to generate the target collaborative operation parameters including the visual range parameters, the relative position parameters, and the formation parameters.
[0059] It should be noted that the visual range training data subset includes corresponding visual range images and visual range labels one by one. Among them, the visual range image can refer to the images obtained by the acquisition device in different scenarios, which reflects the area it can perceive, and the visual range label clarifies the actual visual range information corresponding to the image, such as the specific acquisition operation angle range, acquisition coverage range, etc. The visual range training data subset is used to train the model to recognize the visual range corresponding to different images, so as to generate reasonable visual range parameters subsequently.
[0060] The relative position training data subset includes corresponding acquisition device position data and relative position labels one by one. Among them, the acquisition device position data records the specific position information of the high-altitude acquisition device and the ground acquisition device at different times, and the relative position label is used to represent the relative position relationship between the two acquisition devices, such as distance. Through the relative position training data subset, the model can learn how to infer the relative position between the acquisition devices according to the position data of the acquisition devices.
[0061] In this embodiment, the number of high-altitude acquisition devices and ground acquisition devices is not limited, and each of the high-altitude acquisition devices and ground acquisition devices can be one or more. The formation training data subset includes corresponding acquisition device position data and formation labels one by one. Among them, the formation label describes the shape and layout information of the formation composed of multiple acquisition devices, such as a linear formation, a triangular formation or other specific formations. By learning the data in the formation training data subset, the model can generate a suitable formation according to the position of each acquisition device.
[0062] It is understandable that the operation parameter inference model includes, but is not limited to, a visual range processing sub-model, a relative position processing sub-model, and a formation processing sub-model. Each sub-model is responsible for processing and learning a corresponding part of the data in the dataset to generate different types of parameters. The actual carriers of each sub-model can be manifested as algorithm modules or model units. Among them, the visual range processing sub-model masters how to extract key features from the input image data and predict reasonable visual range parameters by learning the relationship between the visual range image and the corresponding label. The relative position processing sub-model can calculate the relative position parameters between them according to the real-time position data of the acquisition devices. The formation processing sub-model can calculate reasonable formation parameters according to the position data of the acquisition devices to ensure that multiple acquisition devices can work together in the best formation.
[0063] It should be understood that supervised training refers to training each sub-model using the method of supervised learning. Supervised learning means that during the training process, the model knows the input data (such as visual range images, acquisition device position data) and the corresponding expected outputs (such as visual range labels, relative position labels, formation labels), and by continuously adjusting the parameters of the model, the output of the model is made to approach the expected output as much as possible. After supervised training of each sub-model, a trained operation parameter inference model can be obtained. This model can comprehensively process the input first environmental perception data and second environmental perception data, generate target collaborative operation parameters including visual range parameters, relative position parameters, and formation parameters, so as to guide the high-altitude acquisition device and the ground acquisition device to perform collaborative operations and improve the efficiency and quality of data acquisition.
[0064] In this embodiment, on the one hand, through lightweight and modular model design, the computational burden and latency in the inference process can be reduced, and the real-time data processing and decision-making response capabilities can be enhanced. On the other hand, by constructing a multi-task training dataset containing subsets of visual range, relative position, and formation training data, and performing supervised training on the three sub-models of visual range processing, relative position processing, and formation processing respectively, an operation parameter inference model that can generate comprehensive target collaborative operation parameters is created. When this model faces the environmental perception data transmitted from the high-altitude acquisition device and the ground acquisition device, it can calculate appropriate visual range parameters (used to indicate the visual data for the acquisition device to clearly capture objects), relative position parameters (used to ensure the best relative position between the acquisition devices and avoid acquisition blind spots and missed target acquisitions), and formation parameters (used to indicate efficient collaborative operations), improving the comprehensiveness, accuracy, and collaborative efficiency of data acquisition by the acquisition devices in the virtual world.
[0065] Based on any of the above embodiments, the method further includes: In the case that the object recognition result indicates that the target object is not recognized, obtain an updated multi-task training data set including historical reconnaissance data, and use the updated multi-task training data set to train the trained operation parameter inference model to obtain a trained updated operation parameter inference model. Return to the step of executing the acquisition of the first object vision data set for collecting operation-related objects on the preset operation path of the aerial acquisition device in the virtual world until the aerial acquisition device and the ground acquisition device collect the vision data including the target object according to the updated target collaborative operation parameters generated by the trained updated operation parameter inference model.
[0066] It should be noted that in the reconnaissance operation in the virtual world, the target object has been included in the preset operation path in the path planning stage. Under normal circumstances, the target object can be recognized. However, if the target object fails to be recognized, it may be because the currently set target collaborative operation parameters for the acquisition device are not reasonable enough. Based on the unreasonable parameters, the time taken for the acquisition device to travel to the location of the target object for data collection is too long. During the process of the acquisition device traveling along the preset operation path according to the unreasonable target collaborative operation parameters, the target object may disappear due to various factors. For example, the target object may disappear due to the explosion special effect in the virtual world, or disappear from the field of view due to the switching of the scene special effect. In view of this, it is necessary to optimize the operation parameter inference model so that the target collaborative operation parameters generated by the model are more in line with the actual operation requirements, thereby reducing or eliminating the situation where the acquisition device fails to collect the vision data of the target object before it disappears.
[0067] Specifically, when the target object is not recognized, it is necessary to obtain an updated multi-task training data set containing historical reconnaissance data. The historical reconnaissance data records the acquisition situations of past acquisition devices under different operation parameters, as well as the target objects and environmental information in the corresponding scenarios. These data can help the model learn more appropriate parameter settings in various complex situations (especially in scenarios where the target object is likely to disappear). Then, the trained operation parameter inference model is retrained using the updated multi-task training data set. During the training process, the model adjusts its internal parameters and structure according to the new data. For example, by analyzing the historical data, the parameter rules such as the optimal relative position, visual range, and formation of the acquisition device in scenarios where the target object may disappear are obtained, so that the model can generate more suitable target collaborative operation parameters. After the model training is completed to obtain the trained updated operation parameter inference model, return to execute the step of obtaining the first object visual data set collected by the high-altitude acquisition device for the operation-related objects on the preset operation path. At this time, the acquisition device will travel in the virtual world according to the updated target collaborative operation parameters, and these parameters are optimized, making it more likely for the acquisition device to collect the visual data of the target object before it disappears. Continuously repeat the above processes of recognition, optimization, and acquisition. If the target object still cannot be recognized, update the training data set and train the model again until the acquisition device successfully collects the visual data containing the target object according to the updated target collaborative operation parameters, completing the reconnaissance task.
[0068] In a specific implementation, the object recognition model provides real-time feedback on the current reconnaissance results, thereby continuously optimizing the subsequent deployment paths and reconnaissance angles of the ground acquisition device and the high-altitude acquisition device. Using the mechanism of real-time feedback and reinforcement learning, the path generation model and the operation parameter inference model can quickly analyze and evaluate the information collected during the reconnaissance task, and dynamically adjust the task path, operation angle, and formation. Through continuous iterative optimization, it can adaptively handle the uncertain factors and emergencies in the complex environment, ensure the continuous and effective progress of the task, and at the same time effectively reduce the need for human intervention and improve the reconnaissance operation efficiency.
[0069] In this embodiment, by integrating the historical reconnaissance data into the generation of the updated multi-task training data set, the operation parameter inference model is optimized, prompting it to generate more suitable target collaborative operation parameters, increasing the probability of the acquisition device collecting the visual data of the target object before it disappears, reducing problems such as acquisition delays, and thus ensuring the completion degree of the reconnaissance task.
[0070] Based on any of the above embodiments, each of the first object visual data carries first acquisition time information, and each of the second object visual data carries second acquisition time information; Performing spatio-temporal synchronization calibration processing on the first set of object visual data and the second set of object visual data to obtain a processed visual data set, including: Based on the first acquisition time information and the second acquisition time information, sorting a plurality of the first object visual data and a plurality of the second object visual data according to the acquisition time to obtain a time-calibrated data set; For each of the first object visual data, determining target second object visual data that has feature pixel correlation with the first object visual data from a plurality of the second object visual data, and integrating the first object visual data and the target second object visual data in pixel space based on the feature pixel correlation to obtain a spatially fused data set; Based on the spatio-temporal correlation between the time-calibrated data set and the spatially fused data set, determining the processed visual data set.
[0071] It should be noted that the purpose of sorting based on time information is to make the visual data from different acquisition devices consistent in time order, facilitating subsequent processing and analysis. By sorting the time, it is possible to clearly understand the data situation collected at different times and avoid data processing errors caused by chaotic time order.
[0072] Feature pixel correlation means that the pixel features of key scene elements in two or more visual data are similar, such as the pixel features of the shape, color, texture, etc. of an object. By finding the second object visual data with feature pixel correlation, it is possible to ensure that the visually corresponding and complementary visual data to the first object visual data is selected in space.
[0073] Optionally, using techniques such as image stitching and feature fusion to merge two correlated visual data at the pixel level, so that the processed data can comprehensively reflect the information of the object and the scene, achieving spatial fusion and making up for the limitations of the perspectives of using only high-altitude acquisition devices or only ground acquisition devices. Further, based on the spatio-temporal connection between the data, data integration is performed to obtain a processed visual data set.
[0074] In a specific implementation, for multi-angle image or video data collected by high-altitude acquisition devices and ground acquisition devices, high-precision image registration techniques and data fusion algorithms are used to achieve precise alignment and fusion of cross-device and multi-perspective data. This process includes spatio-temporal synchronization calibration of multi-perspective data, eliminating time differences and spatial position information errors, and performing information compensation and enhancement to make the generated processed visual data set more comprehensive and accurate.
[0075] In this embodiment, through the processing and calibration in two dimensions of time and space, the spatio-temporal synchronization calibration of the first set of visual data of the object and the second set of visual data of the object is achieved, the quality of the visual data is improved, and the accuracy and reliability of the reconnaissance of the target object in the virtual world are enhanced.
[0076] Based on the method described in any of the above embodiments, the present application further provides a Figure 2 structural schematic diagram of an electronic device as shown in Figure 2 . At the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include other hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the method described in any of the above embodiments.
[0077] Based on the method described in any of the above embodiments, the present application further provides a computer storage medium storing a computer program, which can be used to execute the method described in any of the above embodiments when being executed by a processor.
[0078] Based on the method described in any of the above embodiments, the present application further provides a computer program product, which includes one or more computer programs or instructions. The computer program or instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. When the computer program is executed by a processor, it implements the method described in any of the above embodiments.
[0079] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are only illustrative. For example, the flowcharts and block diagrams in the drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code includes one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0080] In addition, in each embodiment of the present application, the functional modules can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part.
[0081] If the above-mentioned function is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.
[0082] The above are only the embodiments of the present application and are not used to limit the protection scope of the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application. It should be noted that similar reference numerals and letters indicate similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0083] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed by the present application, and all should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
[0084] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent in such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.
Claims
1. A method for detecting a target object, characterized in that: The method comprises: Acquire a first object visual data set collected by an aerial acquisition device on a preset operation path in a virtual world for an operation-related object, and acquire a second object visual data set collected by a ground acquisition device on the preset operation path for the operation-related object; wherein the operation-related objects include disappearing interference objects and disappearing target objects; each first object visual data in the first object visual data set and each second object visual data in the second object visual data set carry position information, and the position information is used to characterize the position of the operation-related object in a non-disappearing state in the virtual world; Performing spatiotemporal synchronization calibration processing on the first object visual data set and the second object visual data set to obtain a processed visual data set; Inputting the processed visual data set into a trained object recognition model to obtain an object recognition result, and if the object recognition result indicates that the target object is recognized, determining target visual data including the target object from the processed visual data set; The position information of the target object is determined based on the position information carried by the target visual data.
2. The method according to claim 1, characterized in that Before obtaining the first object visual data set collected by the high-altitude collection device on the preset operation path in the virtual world for the operation-related objects, the method further includes: The environmental data of the area to be operated is input into the trained path generation model to obtain the preset operation path; the environmental data includes the topographic data of the area to be operated, the obstacle distribution data, and the morphological data of the target object.
3. The method according to claim 1, characterized in that Before obtaining the first object visual data set collected by the high-altitude collection device on the preset operation path in the virtual world for the operation-related objects, the method further includes: Acquire first environmental perception data collected by the high-altitude acquisition device at a first acquisition angle at a first position on the preset operation path, and acquire second environmental perception data collected by the ground acquisition device at a second acquisition angle at a second position on the preset operation path; Based on the first environmental perception data and the second environmental perception data, target collaborative operation parameters are determined, and the high-altitude acquisition device is controlled to collect the first object visual data set according to the target collaborative operation parameters, and the ground acquisition device is controlled to collect the second object visual data set according to the target collaborative operation parameters; wherein the target collaborative operation parameters are used to indicate the first target position and the first target acquisition angle of the high-altitude acquisition device, and the second target position and the second target acquisition angle of the ground acquisition device.
4. The method according to claim 3, characterized in that The high-altitude collection device is equipped with a first environment perception sensor; the ground collection device is equipped with a second environment perception sensor; the first environment perception data is collected by the first environment perception sensor; the second environment perception data is collected by the second environment perception sensor; The determining of the target collaborative operation parameter based on the first environmental perception data and the second environmental perception data includes: The first environmental perception data and the second environmental perception data are input into a trained operation parameter inference model to obtain the target collaborative operation parameters.
5. The method according to claim 4, characterized in that The target collaborative operation parameters include visual range parameters, relative position parameters between the high-altitude acquisition device and the ground acquisition device, and formation parameters; the operation parameter reasoning model includes a visual range processing sub-model, a relative position processing sub-model, and a formation processing sub-model; the trained operation parameter reasoning model is trained by the following steps: Acquire a multi-task training data set; the multi-task training data set includes a visual range training data subset, a relative position training data subset, and a formation training data subset; the visual range training data subset includes a one-to-one corresponding visual range image and a visual range label; The relative position training data subset includes one-to-one corresponding acquisition device position data and relative position labels; The formation training data subset includes one-to-one correspondence between the acquisition device position data and the formation label; The location of the acquisition device includes the location data of the high-altitude acquisition device and the location data of the ground acquisition device; The visual range processing sub-model is supervisedly trained using the visual range training data subset, the relative position processing sub-model is supervisedly trained using the relative position training data subset, and the formation formation processing sub-model is supervisedly trained using the formation formation training data subset to obtain the trained operation parameter inference model, and the trained operation parameter inference model is used to generate the target collaborative operation parameters including the visual range parameters, the relative position parameters, and the formation formation parameters.
6. The method according to claim 5, characterized in that The method further comprises: When the object recognition result indicates that the target object is not recognized, an updated multi-task training data set including historical reconnaissance data is obtained, and the trained operation parameter inference model is trained using the updated multi-task training data set to obtain the trained updated operation parameter inference model, and the step of obtaining a first object visual data set collected by the high-altitude acquisition device on the preset operation path in the virtual world for the operation-related objects is returned to execute until the high-altitude acquisition device and the ground acquisition device collect visual data including the target object according to the updated target collaborative operation parameters generated by the trained updated operation parameter inference model.
7. The method according to claim 1, characterized in that Each of the first object visual data carries first acquisition time information, and each of the second object visual data carries second acquisition time information; The step of performing spatiotemporal synchronization calibration processing on the first object visual data set and the second object visual data set to obtain a processed visual data set includes: Based on the first acquisition time information and the second acquisition time information, sorting the plurality of first object visual data and the plurality of second object visual data according to acquisition time to obtain a time-calibrated data set; For each first object visual data, determine target second object visual data having feature pixel correlation with the first object visual data from a plurality of second object visual data, and integrate the first object visual data with the target second object visual data in pixel space based on the feature pixel correlation to obtain a spatially fused data set; The processed visual dataset is determined based on the spatiotemporal correlation between the temporally calibrated dataset and the spatially fused dataset.
8. An electronic device, characterized in that: The electronic device comprises: processor; a memory for storing processor-executable instructions; Wherein, when the processor calls the executable instruction, the method described in any one of claims 1-7 is implemented.
9. A computer-readable storage medium, characterized in that: Computer instructions are stored thereon, and when the computer instructions are executed by a processor, the steps of any method described in claims 1-7 are implemented.
10. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Game control system and method based on stereo vision
CN101086681A
Image processing method and device, and storage medium
CN108022301A
Method and device for adjusting virtual camera
CN113055589A
Virtual and real entity coordinate mapping method and system for building digital twinning
CN114329747A
Three-dimensional information determination method and device, equipment and storage medium
CN117115274A