Target detection method and device, storage medium and electronic equipment
By using trajectory prediction and perception models to align image feature data in a collaborative perception system, the inaccuracy problem of object detection caused by network delay is solved, and higher detection accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510338723.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-03-20
AI Technical Summary
In a collaborative perception system, the accuracy of target detection is reduced due to the time delay introduced by the network connection, and information of different agents cannot be aligned in time and space.
By acquiring the historical moment image feature data of the second agent, the pre-trained trajectory prediction model determines the trajectory data of the target object, and combines the trajectory perception model to perform feature alignment to ensure the accuracy of object detection.
It effectively avoids the adverse impact of data transmission delay on target detection, improves the accuracy and robustness of target detection, and reduces false detection and missed detection.
Smart Images

Figure CN120299000A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of computer technology, and in particular, to an object detection method, apparatus, storage medium, and electronic device. Background Art
[0002] With the rapid development of vehicle networking technology, collaborative perception systems have been widely used in the field of intelligent transportation. The collaborative perception system constructs a global environmental perception model and performs object recognition by sharing and fusing sensor data (such as LiDAR point clouds, camera images) of multiple agents (such as vehicles, roadside units, etc.) in real time, thereby improving the safety and efficiency of the transportation system.
[0003] However, in the process of collaborative perception, due to the existence of network connections, time delays will inevitably be introduced, such as network communication delays, data processing delays, and clock synchronization errors. These time delays cause the information currently received by agents in the driving environment to actually be lagged information collected by other agents before, and the information of different agents cannot be aligned in time and space, seriously affecting the accuracy of object recognition results.
[0004] Therefore, how to avoid the impact of time delay in the collaborative perception process on object detection and ensure the accuracy of object detection is an urgent problem to be solved. Summary of the Invention
[0005] This specification provides an object detection method, apparatus, storage medium, and electronic device to partially solve the above problems existing in the prior art.
[0006] This specification adopts the following technical solutions:
[0007] An object detection method, which is applied to a first agent in a driving environment, and the method includes:
[0008] Obtain image feature data at a historical moment sent by a second agent in the driving environment; wherein, the image feature data contains a time series of image feature data of a target object;
[0009] Input the image feature data into a pre-trained trajectory prediction model to determine trajectory data of the target object; wherein, the trajectory data contains position information and direction information of the target object in the driving environment;
[0010] Determine first image feature data corresponding to the target object at the current moment according to the trajectory data;
[0011] Based on the first image feature data and the second image feature data corresponding to the target object in the image collected by the first agent at the current moment, perform object detection.
[0012] Optionally, determining the first image feature data corresponding to the target object at the current moment according to the trajectory data specifically includes:
[0013] According to the trajectory data, determine each feature point on the movement trajectory of the target object;
[0014] Input the feature point data at each of the feature points into a preset trajectory perception model, so as to determine the first image feature data through the trajectory perception model.
[0015] Optionally, determining each feature point on the movement trajectory of the target object according to the trajectory data specifically includes:
[0016] Input the trajectory data into a pre-trained offset generation model, so as to determine the feature points at each attention position through the offset generation model according to the trajectory distribution of the target object and the chronological order.
[0017] Optionally, inputting the feature point data at each of the feature points into a preset trajectory perception model, so as to determine the first image feature data through the trajectory perception model specifically includes:
[0018] Through the trajectory perception model, based on the correlation relationship between the image feature data of the target object at the current moment and each feature point data, determine the weight between the image feature data of the target object at the current moment and each feature point data;
[0019] According to the weight between the image feature data of the target object at the current moment and each feature point data, perform weighted summation on each feature point data to obtain the first image feature data.
[0020] Optionally, training the trajectory prediction model specifically includes:
[0021] Obtain a historical image sequence containing a target object;
[0022] For each historical image in the historical image sequence, determine the spatio-temporal data of the target object under the historical image, and determine the labeled trajectory data according to the spatio-temporal data of the target object under each historical image;
[0023] Input the feature data corresponding to the historical image sequence into the trajectory prediction model to be trained, so as to determine the predicted trajectory data corresponding to the target object through the trajectory prediction model;
[0024] Determine the loss value of the trajectory prediction model according to the deviation between the predicted trajectory data and the labeled trajectory data, and train the trajectory prediction model according to the loss value.
[0025] Optionally, training the offset generation model specifically includes:
[0026] Obtain the historical trajectory data of the target object and the labeled feature point data corresponding to each actual feature point in the historical trajectory data;
[0027] Input the historical trajectory data into the offset generation model to be trained, so as to generate the predicted feature point data corresponding to the historical trajectory data through the offset generation model;
[0028] Determine the loss value of the offset generation model according to the labeled feature point data and the predicted feature point data, and train the offset generation model according to the loss value.
[0029] Optionally, determining the loss value of the offset generation model according to the labeled feature point data and the predicted feature point data specifically includes:
[0030] Determine the spatial distance between each feature point in the predicted feature point data and each feature point in the labeled feature point data, and determine the matching cost data between different feature points based on the spatial distance;
[0031] Determine the matching rate between each feature point in the predicted feature point data and each feature point in the labeled feature point data according to the matching cost data;
[0032] Taking the maximization of the matching rate between each feature point in the predicted feature point data and each feature point in the labeled feature point data as the goal, determine the feature point matching data;
[0033] Determine the loss value according to the matching cost data and the feature point matching data.
[0034] This specification provides an object detection device, including:
[0035] An acquisition module, configured to acquire the image feature data at a historical moment sent by a second intelligent agent in the driving environment; wherein, the image feature data includes a time series of image feature data of the target object;
[0036] An input module, configured to input the image feature data into a pre-trained trajectory prediction model to determine the trajectory data of the target object; wherein, the trajectory data includes the position information and direction information of the target object in the driving environment;
[0037] A determination module, configured to determine first image feature data corresponding to the target object at the current moment according to the trajectory data;
[0038] A detection module, configured to perform target detection based on the first image feature data and second image feature data corresponding to the target object in the image collected by the first intelligent agent at the current moment.
[0039] This specification provides a computer-readable storage medium storing a computer program, which when executed by a processor implements the above target detection method.
[0040] This specification provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor implements the above target detection method when executing the program.
[0041] At least one of the above technical solutions adopted in this specification can achieve the following beneficial effects:
[0042] In the target detection method provided in this specification, image feature data sent by a second intelligent agent in a historical moment in the driving environment is acquired; the image feature data is input into a pre-trained trajectory prediction model to determine trajectory data of the target object; wherein, the trajectory data includes position information and direction information of the target object in the driving environment at the historical moment and the current moment; according to the trajectory data, first image feature data corresponding to the target object at the current moment is determined; target detection is performed based on the first image feature data and image feature data corresponding to the image including the target object collected by the first intelligent agent at the current moment. This solution effectively avoids the adverse impact on target detection caused by transmission delay during data transmission between different intelligent agents, and ensures the accuracy of target detection.
[0043] As can be seen from the above method, after the first intelligent agent receives the historical image feature data sent by the second intelligent agent, it will first predict the spatio-temporal information including the current moment and the historical moment through a pre-trained trajectory prediction model, and determine the image feature data of the target object at the current moment based on the predicted trajectory data, thereby avoiding the lag of the image feature data sent by the second intelligent agent in the previous moment compared with the image feature data of the first intelligent agent at the current moment, and thus accurately performing feature alignment on multiple image feature data in terms of time and space, effectively improving the accuracy of target detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] The drawings described herein are used to provide a further understanding of this specification, and constitute a part of this specification. The illustrative embodiments and descriptions of this specification are used to explain this specification, and do not constitute an improper limitation of this specification. In the drawings:
[0045] Figure 1 Schematic diagram of the alignment of image features for object detection provided in this specification;
[0046] Figure 2 Schematic diagram of the process flow of an object detection method provided in this specification;
[0047] Figure 3 Schematic diagram of visual trajectory data provided in this specification;
[0048] Figure 4 Schematic diagram of the determination method of image feature data provided in this specification;
[0049] Figure 5 Schematic diagram of the framework of an object detection process provided in this specification;
[0050] Figure 6 Schematic diagram of an object detection device provided in this specification;
[0051] Figure 7 This is provided in this specification corresponding to Figure 2 Schematic diagram of the electronic device. Specific implementation manners
[0052] To make the objectives, technical solutions, and advantages of this specification clearer, the technical solutions of this specification will be clearly and completely described below in conjunction with specific embodiments of this specification and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this specification.
[0053] In a complex multi-agent collaborative system, the dynamic characteristics of network connections inevitably introduce an error accumulation effect in the time dimension. This systematic delay phenomenon is called time delay. Its existence not only weakens the accuracy of real-time interaction, but also poses a fundamental challenge to the effectiveness of collaborative decision-making. Specifically, the generation of time delay can be attributed to the following three core dimensions:
[0054] Communication delay: During wireless transmission, the signal needs to break through the dual shackles of medium propagation delay and channel congestion. When multiple nodes communicate concurrently in a dense spectrum environment, the MAC layer contention avoidance mechanism may cause queuing delays of hundreds of milliseconds. Superimposed on the routing hops and traffic shaping strategies in the transmission path, it ultimately results in an end-to-end communication delay fluctuation range of 50 - 400 ms.
[0055] Computational processing delay: Perception data needs to go through a double processing process before and after transmission. The front-end device needs to perform feature extraction (such as YOLOv5 single-frame processing on Jetson AGX is about 50ms), which involves layer-by-layer operations and data compression of convolutional neural networks; the receiving end needs to perform spatial coordinate system conversion calibration (typically, the conversion between UTM and WGS84 coordinate systems requires 20ms matrix operations). This type of computing load increases nonlinearly with data accuracy requirements and scene complexity.
[0056] Precision limitations of clock synchronization: Although GPS timing accuracy can reach nanoseconds, in actual applications, due to ionospheric interference, urban canyon effects, and receiver crystal drift, the clock deviation between nodes may accumulate to the order of 10ms. In high-speed mobile scenarios (such as 120km / h relative speed in the Internet of Vehicles), a 10ms time difference corresponds to a 33cm position error. This timing misalignment will seriously undermine the mathematical consistency of sensor fusion.
[0057] like Figure 1 As shown, due to the lag effect of the above multi-dimensional time delay, the same feature of the target cannot be aligned in time during target detection, which seriously affects the result of target detection.
[0058] Figure 1 A schematic diagram of image feature alignment for target detection provided in this specification.
[0059] Among them, spatial misalignment means that asynchronous observation causes the features to be unable to be aligned in spatial position, causing target positioning deviation, and semantic misalignment means that the difference in feature representation of the same target across agents leads to false detection and missed detection.
[0060] Based on this, this specification provides a target detection method to eliminate the cross-agent feature space and semantic misalignment caused by communication delays, improve the motion compensation accuracy of multi-frame point cloud features in dynamic scenes, and achieve end-to-end low-latency robust collaborative perception.
[0061] The technical solutions provided by the embodiments of this specification are described in detail below in conjunction with the accompanying drawings.
[0062] Figure 2 A target detection method provided in this specification is a flowchart, comprising the following steps:
[0063] S201: Acquire image feature data of historical moments sent by a second agent in the driving environment; wherein the image feature data includes a time series of image feature data of a target object.
[0064] In an intelligent driving environment, an intelligent agent is usually integrated with multiple sensors (such as radar, lidar, cameras, etc.), so as to perceive the surrounding environment in real time through these sensors, understand traffic rules, predict the behaviors of other road users, and make driving decisions accordingly. These intelligent agents can include driving devices (such as vehicles, aircraft, etc.) and roadside devices (such as signal lights, roadside units, parking systems, etc.).
[0065] In the process of target detection, each intelligent agent can collect images of target objects through its own sensors, so as to obtain image data. These image data can include images collected by different sensors, such as point cloud images, camera (camera) images, infrared images, and depth images, etc.
[0066] Among them, the above-mentioned target objects can be obstacles or other intelligent agents in the driving environment, and there can be one or more of these target objects.
[0067] These intelligent agents can encode the collected original image data to obtain image feature data, and send the image feature data to the first intelligent agent that needs to perform target detection.
[0068] For example, an intelligent agent can use a deep learning model (such as PointPillars) to encode the original data into low-dimensional features (BEV feature maps), and transmit the features to the first intelligent agent through the C-V2X or DSRC protocol.
[0069] The first intelligent agent can obtain the image feature data sent by other intelligent agents (second intelligent agents) in the driving environment at historical moments. The image feature data can contain a time series of image feature data containing target objects sent by the second intelligent agent at several previous moments.
[0070] It should be noted that due to the delay in data transmission between intelligent agents, for any moment, the image feature data received by the first intelligent agent is actually collected and sent by the second intelligent agent before this moment. Therefore, the image feature data obtained by the first intelligent agent is sent by the second intelligent agent at historical moments.
[0071] S202: Input the image feature data into a pre-trained trajectory prediction model to determine the trajectory data of the target object; wherein, the trajectory data contains the position information and direction information of the target object in the driving environment.
[0072] After receiving the image feature data, the first intelligent agent can input it into a pre-trained trajectory prediction model, so as to determine the trajectory data of the target object through this trajectory prediction model.
[0073] The above trajectory data not only includes the position information and direction information corresponding to the target object at several previous moments, but may also include the position information and direction information of the target object at the current moment predicted by the trajectory prediction model based on the historical trajectory data. For ease of understanding, this specification provides a visual schematic diagram of the trajectory data, such as Figure 3 shown
[0074] wherein, the above trajectory data may be the trajectory field of the target object, and the position information and direction information respectively correspond to the position field (trajectory heat) and the direction field (tangent direction of the trajectory).
[0075] Before using the above trajectory prediction model, the trajectory prediction model can be trained by the server first. After the training is completed, it can be deployed and trajectory prediction can be performed. The process of training the trajectory prediction model may include the following steps:
[0076] S301: Obtain a historical image sequence containing the target object;
[0077] S301: For each historical image in the historical image sequence, determine the spatio-temporal data of the target object under the historical image, and determine the labeled trajectory data according to the spatio-temporal data of the target object under each historical image.
[0078] For each frame of historical image, the server can determine the bounding box of the target object in the historical image, and based on the annotation information of the bounding box (such as time, position, etc.), determine the spatio-temporal data of the target object under the historical image.
[0079] Among them, the server can extract the annotation information of the bounding box of each object within the maximum time window. This time window starts from the first appearance of the target object and lasts until the current moment. By connecting the center points of these bounding boxes, the motion trajectory of the target object, that is, the labeled trajectory data, can be constructed.
[0080] Generally speaking, the generation process of the labeled trajectory data can be divided into two steps, namely, performing interpolation calculation on the trajectory; projecting the interpolation result into the feature space. Thus, the labeled trajectory data of the target object can be obtained. It should be noted that when the trajectories overlap, the trajectory with a newer timestamp can cover the old trajectory.
[0081] S302: Input the feature data corresponding to the historical image sequence into the trajectory prediction model to be trained, so as to determine the predicted trajectory data corresponding to the target object through the trajectory prediction model;
[0082] S303: Determine the loss value of the trajectory prediction model according to the deviation between the predicted trajectory data and the labeled trajectory data, and train the trajectory prediction model according to the loss value.
[0083] In this specification, the loss value of the trajectory prediction model can be composed of two parts, namely the position field loss L pos and the direction field loss L ori , where L pos can be determined based on the deviation between the predicted trajectory data and the position information corresponding to the labeled trajectory data, expressed as:
[0084]
[0085] where the position information of the predicted trajectory data is x xyc , and the position information of the labeled trajectory data is
[0086] L ori can be determined based on the deviation between the predicted trajectory data and the direction information corresponding to the labeled trajectory data, expressed as:
[0087]
[0088] where the direction information corresponding to the predicted trajectory data is y, and the direction information corresponding to the labeled trajectory data is M is a mask matrix composed of 0s or 1s, with the same dimension as the trajectory field. The positions with a value of 1 indicate that there are target trajectories at those positions.
[0089] The server can minimize L ori and L pos as the optimization objective to train the trajectory prediction model until the training objective is met (such as reaching a preset number of times or converging to a preset range).
[0090] S203: Determine the first image feature data corresponding to the target object at the current moment according to the trajectory data.
[0091] After determining the trajectory data of the target object, the first agent can determine the feature points on the motion trajectory of the target object according to the trajectory data.
[0092] Specifically, the first agent can input the trajectory data into a pre-trained offset generation model, so as to determine the feature points at each attention position through the offset generation model according to the trajectory distribution of the target object and the chronological order.
[0093] After determining the feature points, the first agent can input the feature point data at each feature point into a preset trajectory perception model, so as to determine the first image feature data through the trajectory perception model.
[0094] Specifically, the first intelligent agent can, through the trajectory perception model, determine the weights between the image feature data of the target object at the current moment and each feature point data based on the association relationship between the image feature data of the target object at the current moment and each feature point data. Then, according to the weights between the image feature data of the target object at the current moment and each feature point data, perform a weighted sum on each feature point data to obtain the first image feature data.
[0095] For ease of understanding, this specification provides a schematic diagram of a method for determining image feature data, as Figure 4 shown.
[0096] Among them, the feature representation of the feature point data at each feature point consists of k and v. Based on q and k, the association relationship between the image feature data of the target object at the current moment and each feature point data can be determined, and then the weight w can be obtained. Then, based on the weight w and the v corresponding to each feature point data, perform a weighted sum to obtain the first image feature data.
[0097] Furthermore, before using the above-mentioned offset generation model, the server can first train the offset generation model, and its training process can include the following steps:
[0098] S401: Obtain the historical trajectory data of the target object and the label feature point data corresponding to each actual feature point in the historical trajectory data.
[0099] The server can obtain the historical trajectory data of the target object and the label feature point data corresponding to each actual feature point in the historical trajectory data. These feature points can be the feature points on the trajectory of the target object expected by the user. When determining the label feature points, these label feature points need to meet the following conditions:
[0100] The feature point data at the feature point contains the target object;
[0101] The feature point is located on the trajectory of the target object;
[0102] The timestamp corresponding to the feature point at any moment is earlier than the next feature point (i.e., this feature point is located at the historical position of the target object).
[0103] S402: Input the historical trajectory data into the offset generation model to be trained, so as to generate the predicted feature point data corresponding to the historical trajectory data through the offset generation model;
[0104] S403: Determine the loss value of the offset generation model according to the label feature point data and the predicted feature point data, and train the offset generation model according to the loss value.
[0105] The server can determine the loss value of the offset generation model based on the predicted feature point data output by the model generated from the user-expected tag feature point data and the offset.
[0106] Specifically, the server can determine the spatial distance (such as the L1 distance) between each feature point in the predicted feature point data and each feature point in the tag feature point data, and based on this spatial distance, determine the matching cost data between different feature points. The cost matrix corresponding to this matching cost data can be expressed as:
[0107] C jk =∑|p j -p k |1
[0108] Where C jk represents the cost matrix of the matching cost data, p j represents the position vector of the predicted feature point data j, and p k represents the position vector of the tag feature point data k. By traversing all the predicted feature points and tag feature points, the matching cost data can be obtained.
[0109] After that, the server can determine the matching rate P between each feature point in the predicted feature point data and each feature point in the tag feature point data according to the matching cost data.
[0110] Furthermore, the server can determine the feature point matching data with the goal of maximizing the matching rate between each feature point in the predicted feature point data and each feature point in the tag feature point data.
[0111] In practical applications, the server can use the sinkhorn algorithm to calculate the matching rate. The goal is to find a matrix such that the sum of each row of P is 1, the sum of each column is also 1, and it satisfies:
[0112]
[0113] After that, the server can determine the loss value according to the matching cost data C jk and the feature point matching data P * , and this loss value can be expressed as:
[0114]
[0115] S204: Perform object detection based on the first image feature data and the second image feature data corresponding to the target object in the image collected by the first agent at the current moment.
[0116] Specifically, the first agent can collect the image data of the target object at the current moment through its own sensors, determine the second image feature data corresponding to the target object based on the image data, and then fuse the first image feature data and the second image feature data of the target object to obtain the fusion data, and then perform target detection on the target object based on the fusion data.
[0117] Among them, the first image feature data represents the feature data corresponding to the target object image collected by the second agent at the current moment, and the second image feature data represents the feature data corresponding to the target object image collected by the first agent at the current moment. The two are aligned in time (that is, both are the current moment). For the convenience of understanding, this specification provides a schematic diagram of the framework of the target detection process, as Figure 5 shown.
[0118] Among them, Agent i represents the second agent in the driving environment, and Ego represents the first agent. After the input historical image feature data (t-1) is processed by the trajectory prediction model, the offset prediction model, and the trajectory perception model, the image feature data at the current moment (time t) can be obtained. By fusing the first image feature data corresponding to the second agent at the current moment and the second image feature data corresponding to the first agent at the current moment, target detection can be performed based on the fusion feature data.
[0119] It should be noted that for each agent, the agent can also process the image data collected at the current moment t based on the locally deployed trajectory prediction model, offset prediction model, and trajectory perception model to obtain the image feature data at the future moment t+1 and send it to the first agent. Due to the data transmission delay, after the first agent receives the first image feature data sent by the second agent, the second image feature data corresponding to the image collected by the first agent itself is just aligned in time with the first image feature data, that is, both are the image feature data at the moment t+1.
[0120] In the vehicle-to-vehicle (V2V) scenario, the input of the trajectory prediction model can be 4 frames of LiDAR data (0.4m rasterized), and the temporal embedding dimension is 64. The model architecture can adopt a 5-layer UNet, and the number of output channels is [32, 64, 128, 256, 512].
[0121] The number of attention heads of the trajectory perception model can be set to 4, and the number of feature points can be set to 18.
[0122] In the vehicle-to-infrastructure (V2I scenario), the roadside unit (RSU) can adopt 4-frame point clouds, and the vehicle side adopts 2 frames. The timing compensation adopts ego-motion compensation based on GPS / IMU. The deployment platform is NVIDIA Jetson AGX Xavier (inference latency ≤ 100 ms).
[0123] As can be seen from the above method, this solution guides the high-dimensional feature alignment through low-dimensional trajectory prediction, avoids the instability of directly predicting high-dimensional features, dynamically distributes attention points along the trajectory, supports non-linear motion compensation, and eliminates feature ambiguity based on the trajectory consistency attention mechanism. It effectively improves the accuracy of object detection.
[0124] During the actual test process, in this solution, the AP50 / 70 on the V2V4Real dataset is increased to 74.28% / 44.14% (compared with the single-agent +25.53% / 13.61%), and the detection accuracy is effectively improved;
[0125] At a 400 ms delay, the AP50 only drops by 4.87% (compared with the ERMVP, which drops by 13.48%), improving the robustness of object detection;
[0126] The end-to-end inference speed reaches 25 FPS (NVIDIA RTX 3090), effectively improving the computing efficiency.
[0127] The above is one or more implementation object detection methods of this specification. Based on the same idea, this specification also provides a corresponding object detection device, such as Figure 6 shown.
[0128] Figure 6 It is a schematic diagram of an object detection device provided by this specification, including:
[0129] An acquisition module 601, configured to acquire the image feature data at a historical moment sent by a second intelligent agent in the driving environment; wherein, the image feature data includes a time series of image feature data of a target object;
[0130] An input module 602, configured to input the image feature data into a pre-trained trajectory prediction model to determine the trajectory data of the target object; wherein, the trajectory data includes the position information and direction information of the target object in the driving environment;
[0131] A determination module 603, configured to determine the first image feature data corresponding to the target object at the current moment according to the trajectory data;
[0132] The detection module 604 is configured to perform object detection based on the first image feature data and the second image feature data corresponding to the target object in the image collected by the first agent at the current moment.
[0133] Optionally, the determining module 603 is specifically configured to determine each feature point on the motion trajectory of the target object according to the trajectory data; input the feature point data at each feature point into a preset trajectory perception model, so as to determine the first image feature data through the trajectory perception model.
[0134] Optionally, the determining module 603 is specifically configured to input the trajectory data into a pre-trained offset generation model, so as to determine the feature points at each attention position through the offset generation model according to the trajectory distribution of the target object and the chronological order.
[0135] Optionally, the determining module 603 is specifically configured to determine the weights between the image feature data of the target object at the current moment and each feature point data based on the association relationship between the image feature data of the target object at the current moment and each feature point data through the trajectory perception model; perform weighted summation on each feature point data according to the weights between the image feature data of the target object at the current moment and each feature point data to obtain the first image feature data.
[0136] Optionally, the device further includes:
[0137] The training module 605 is configured to obtain a historical image sequence including a target object; for each historical image in the historical image sequence, determine the spatio-temporal data of the target object in the historical image, and determine the labeled trajectory data according to the spatio-temporal data of the target object in each historical image; input the feature data corresponding to the historical image sequence into a to-be-trained trajectory prediction model, so as to determine the predicted trajectory data corresponding to the target object through the trajectory prediction model; determine the loss value of the trajectory prediction model according to the deviation between the predicted trajectory data and the labeled trajectory data, and train the trajectory prediction model according to the loss value.
[0138] Optionally, the training module 605 is specifically configured to obtain the historical trajectory data of the target object and the labeled feature point data corresponding to each actual feature point in the historical trajectory data; input the historical trajectory data into a to-be-trained offset generation model, so as to generate the predicted feature point data corresponding to the historical trajectory data through the offset generation model; determine the loss value of the offset generation model according to the labeled feature point data and the predicted feature point data, and train the offset generation model according to the loss value.
[0139] Optionally, the training module 605 is specifically configured to determine the spatial distance between each feature point in the predicted feature point data and each feature point in the labeled feature point data, and determine the matching cost data between different feature points based on the spatial distance; determine the matching rate between each feature point in the predicted feature point data and each feature point in the labeled feature point data according to the matching cost data; and determine the feature point matching data with the goal of maximizing the matching rate between each feature point in the predicted feature point data and each feature point in the labeled feature point data.
[0140] Determine the loss value according to the matching cost data and the feature point matching data.
[0141] This specification also provides a computer-readable storage medium, which stores a computer program that can be used to execute the above Figure 2 provided object detection method.
[0142] This specification also provides Figure 7 the schematic structural diagram of an electronic device corresponding to Figure 2 . As Figure 7 described, at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include other hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the above Figure 3 described object detection method. Of course, in addition to the software implementation method, this specification does not exclude other implementation methods, such as logic devices or a combination of software and hardware. That is to say, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0143] The improvement of a technology can be clearly distinguished as either a hardware improvement (e.g., improvement of circuit structures such as diodes, transistors, switches, etc.) or a software improvement (improvement of method flows). However, with the development of technology, many improvements of method flows today can be regarded as direct improvements of hardware circuit structures. Almost all designers obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement of a method flow cannot be implemented with a hardware entity module. For example, a programmable logic device (PLD) (e.g., a field programmable gate array (FPGA)) is such an integrated circuit whose logic function is determined by the user's programming of the device. The designer programs it himself to "integrate" a digital system on a piece of PLD, without the need to ask a chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compiler used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a hardware description language (HDL), and there is not only one type of HDL, but many types, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. Currently, the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog.
[0144] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to make the controller implement the same function in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or the structures within the hardware component.
[0145] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.
[0146] For the convenience of description, the above devices are described by dividing them into various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0147] Those skilled in the art should understand that the embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, this specification can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.
[0148] This specification is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the specification. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices generate means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0149] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0150] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are performed on the computer or other programmable device to produce a computer-implemented process, thereby providing steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0151] In a typical configuration, a processing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0152] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.
[0153] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a processing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0154] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0155] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0156] This specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media including storage devices.
[0157] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and reference can be made to the corresponding parts of the method embodiments for relevant details.
[0158] The above description is only for the embodiments of this specification and is not intended to limit this specification. For those skilled in the art, various changes and modifications can be made to this specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification shall be included within the scope of the claims of this specification.
Claims
1. A target detection method, characterized in that, The method is applied to a first intelligent agent in a driving environment, and the method includes: Obtain the image feature data at a historical moment sent by a second intelligent agent in the driving environment; wherein, the image feature data includes a time series of image feature data of a target object; Input the image feature data into a pre-trained trajectory prediction model to determine the trajectory data of the target object; wherein, the trajectory data includes the position information and direction information of the target object in the driving environment; Determine the first image feature data corresponding to the target object at the current moment according to the trajectory data; Perform object detection based on the first image feature data and the second image feature data corresponding to the target object in the image collected by the first intelligent agent at the current moment.
2. The method according to claim 1, wherein Determine the first image feature data corresponding to the target object at the current moment according to the trajectory data, specifically including: Determine each feature point on the movement trajectory of the target object according to the trajectory data; Input the feature point data at each feature point into a preset trajectory perception model to determine the first image feature data through the trajectory perception model.
3. The method according to claim 2, wherein Determine each feature point on the movement trajectory of the target object according to the trajectory data, specifically including: Input the trajectory data into a pre-trained offset generation model to determine the feature points at each attention position through the offset generation model according to the trajectory distribution of the target object and the sequence of time series.
4. The method according to claim 2, wherein Input the feature point data at each feature point into a preset trajectory perception model to determine the first image feature data through the trajectory perception model, specifically including: Through the trajectory perception model, determine the weight between the image feature data of the target object at the current moment and each feature point data based on the correlation relationship between them; Perform weighted summation on each feature point data according to the weight between the image feature data of the target object at the current moment and each feature point data to obtain the first image feature data.
5. The method according to claim 1, characterized in that, Train the trajectory prediction model, specifically including: Obtain a historical image sequence including a target object; For each historical image in the historical image sequence, determine the spatio-temporal data of the target object in the historical image, and determine the labeled trajectory data according to the spatio-temporal data of the target object in each historical image; Input the feature data corresponding to the historical image sequence into the trajectory prediction model to be trained to determine the predicted trajectory data corresponding to the target object through the trajectory prediction model; Determine the loss value of the trajectory prediction model according to the deviation between the predicted trajectory data and the labeled trajectory data, and train the trajectory prediction model according to the loss value.
6. The method according to claim 3, wherein Train the offset generation model, specifically including: Obtain the historical trajectory data of the target object and the labeled feature point data corresponding to each actual feature point in the historical trajectory data; Input the historical trajectory data into an offset generation model to be trained, so as to generate prediction feature point data corresponding to the historical trajectory data through the offset generation model; Determine the loss value of the offset generation model according to the labeled feature point data and the prediction feature point data, and train the offset generation model according to the loss value.
7. The method according to claim 6, wherein Determining the loss value of the offset generation model according to the labeled feature point data and the prediction feature point data specifically includes: Determine the spatial distance between each feature point in the prediction feature point data and each feature point in the labeled feature point data, and determine the matching cost data between different feature points based on the spatial distance; Determine the matching rate between each feature point in the prediction feature point data and each feature point in the labeled feature point data according to the matching cost data; Taking the maximization of the matching rate between each feature point in the prediction feature point data and each feature point in the labeled feature point data as the goal, determine the feature point matching data; Determine the loss value according to the matching cost data and the feature point matching data.
8. An object detection device, characterized in that, Includes: An acquisition module for acquiring image feature data at a historical moment sent by a second intelligent agent in a driving environment; wherein, the image feature data includes a time series of image feature data of a target object; An input module for inputting the image feature data into a pre-trained trajectory prediction model to determine the trajectory data of the target object; wherein, the trajectory data includes the position information and direction information of the target object in the driving environment; A determination module for determining the first image feature data corresponding to the target object at the current moment according to the trajectory data; A detection module for performing target detection based on the first image feature data and the second image feature data corresponding to the target object in the image collected by the first intelligent agent at the current moment.
9. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, and when the computer program is executed by a processor, the method described in any one of claims 1 to 7 above is implemented.
10. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, the method described in any one of claims 1 to 7 above is implemented.
Citation Information
Patent Citations
Unmanned system air-ground collaborative navigation and obstacle avoidance method based on vision
CN116540784A
Automatic driving method and device, storage medium and electronic equipment
CN118665531A
Unmanned driving device control
US20220340174A1