Method and device for detection of state of target, and intelligent device and medium
By extracting and matching features and associating the current target within the current time frame and the first historical target within the historical time frame, the problems of insufficient accuracy of target state detection and high model complexity in the prior art are solved, and lightweight target state prediction and detection separation are achieved, and prediction accuracy is improved.
Patent Information
- Application Number
- PCT/CN2024/130391
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-08
- Filing Date
- 2024-11-07
- Publication Date
- 2025-06-12
AI Technical Summary
The existing target state detection methods have insufficient prediction accuracy and are highly complex in the model, making it difficult to achieve accurate state prediction under lightweight models.
By extracting the current target in the current time frame and the first historical target in the historical time frame, the first mask attention information is used to perform correlation matching, the second feature information is obtained, and state prediction is performed based on this.
The separation of object detection and state prediction is realized, making the model lightweight and able to process arbitrary modal data without the need to perform feature extraction and adaptation for different modal data, reducing the complexity of the network structure and improving the accuracy of state information.
Smart Images

Figure CN2024130391_12062025_PF_FP_ABST
Abstract
Description
Target state detection method, device, smart device, and medium
[0001] This application claims priority to Chinese patent application No. 202311691390.7, filed on December 8, 2023, with the invention name “Target state detection method, device, smart device and medium”. The entire contents of the above Chinese patent application are incorporated into this application by reference. Technical Field
[0002] The present application relates to the field of target detection technology, and specifically provides a target state detection method, device, intelligent device and medium. Background Art
[0003] Autonomous driving capabilities are gaining increasing recognition, and with advances in sensor and information technology, user expectations are also rising. Accurately tracking targets within a perception scenario provides complete and accurate environmental information for downstream functional modules like Planning and Control (Planning and Control) and Autonomous Emergency Braking (AEB), enhancing the user experience.
[0004] In the related art, target state detection methods are roughly divided into two categories: 1) detection methods based on rule filtering; 2) detection methods based on pre-fusion neural networks.
[0005] Detection methods based on rule filtering generally assume that the object moves at a constant speed or with a constant acceleration. Based on the state of the target in the historical time frame, the state of the current time frame is predicted using the classical physics motion formula and Kalman filtering. However, this method has the following disadvantages: (1) The state of the current time frame depends on the state of the target in the historical time frame and is easily affected by upstream noise, resulting in an error accumulation effect.
[0006] Detection methods based on pre-fusion neural networks typically use radar point cloud data and camera image data as input. After extracting features through different branches, the feature maps are fused within the network. Ultimately, the target is detected and its state is predicted. However, this method has the following drawbacks: target detection and state prediction are integrated into a single model, and corresponding branches are required for different modal data to adapt to changes in the modal data, resulting in a complex model.
[0007] Therefore, how to accurately predict the target state under a relatively lightweight model is a technical problem that needs to be solved urgently by those skilled in the art.
[0008] Summary of the Invention
[0009] In order to overcome the above-mentioned defects, the present application is proposed to provide a target state detection method, device, intelligent device and medium that solve or at least partially solve the technical problems of low accuracy in predicting target states and complex models for predicting target states.
[0010] In a first aspect, the present application provides a method for detecting a target state, the method comprising:
[0011] Performing feature extraction on each current target and each first historical target to obtain first feature information of the current target and historical feature information of the first historical target; wherein the current target is a target within a current time frame; and the first historical target is a target within each historical time frame;
[0012] Determining first masked attention information of a current time frame based on the first feature information and the historical feature information of the first historical target; wherein the first masked attention information is used to indicate a first historical target that matches the current target;
[0013] Obtaining second feature information of the current target based on the first mask attention information, the first feature information, and the historical feature information of the first historical target;
[0014] Based on the second feature information, a state prediction is performed on the current target to obtain state information of the current target; wherein the state information includes speed information and / or direction information.
[0015] In a second aspect, the present application provides a target state detection device, which includes a processor and a storage device, wherein the storage device is suitable for storing multiple program codes, and the program codes are suitable for being loaded and run by the processor to execute any of the target state detection methods described above.
[0016] In a third aspect, a smart device is provided. The smart device may include the target state detection device as described above.
[0017] In a fourth aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a plurality of program codes, wherein the program codes are suitable for being loaded and run by a processor to execute any one of the target state detection methods described above.
[0018] The above one or more technical solutions of this application have at least one or more of the following beneficial effects:
[0019] In the technical solution of the present application, by extracting features from each current target in the current time frame and each first historical target in the historical time frame, the first feature information of the current target and the historical feature information of the first historical target are obtained; based on the first feature information of the current target and the historical feature information of the first historical target, the first mask attention information of the current time frame is determined; based on the first mask attention information, the first feature information of the current target and the historical feature information of the first historical target, the second feature information of the current target is obtained; based on the second feature information of the current target, the state of each current target is predicted to obtain the state information of each current target. In this way, the separation of target detection and state prediction is achieved, making each model lightweight, and any modal data can be used as input, without the need to adapt different modal data for feature extraction, reducing the complexity of the network structure, and in each detection, the state of the current target of the current time frame can be independently detected under the first mask attention information, without being affected by the upstream detection results, and the obtained state information is more accurate. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The disclosure of this application will be more easily understood with reference to the accompanying drawings. Those skilled in the art will readily appreciate that these drawings are for illustrative purposes only and are not intended to limit the scope of protection of this application. Furthermore, similar numbers in the figures represent similar components, where:
[0021] FIG1 is a schematic flow chart of main steps of a method for detecting a target state according to an embodiment of the present application;
[0022] FIG2 is a schematic diagram of a network structure for detecting target status of the present application;
[0023] FIG3 is a schematic diagram of an actual scenario using a conventional target state detection method;
[0024] FIG4 is a schematic diagram of an actual scenario using the target state detection method of the present application;
[0025] FIG5 is a schematic diagram of another actual scenario using the target state detection method of the present application;
[0026] FIG6 is a schematic diagram of another actual scenario using the target state detection method of the present application;
[0027] FIG7 is a schematic diagram of another actual scenario using the target state detection method of the present application;
[0028] FIG8 is a schematic diagram of yet another actual scenario using the target state detection method of the present application;
[0029] FIG9 is a main structural block diagram of a target state detection device according to an embodiment of the present application. DETAILED DESCRIPTION
[0030] Some embodiments of the present application are described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to explain the technical principles of the present application and are not intended to limit the scope of protection of the present application.
[0031] In the description of this application, "module" and "processor" may include hardware, software, or a combination of both. A module may include hardware circuitry, various suitable sensors, communication ports, and memory. It may also include software components, such as program code, or a combination of software and hardware. A processor may be a central processing unit, a microprocessor, an image processor, a digital signal processor, or any other suitable processor. A processor has data and / or signal processing capabilities. A processor may be implemented in software, hardware, or a combination of both. Non-transitory computer-readable storage media include any suitable medium capable of storing program code, such as magnetic disks, hard disks, optical disks, flash memory, read-only memory, random access memory, etc. The term "A and / or B" refers to all possible combinations of A and B, such as only A, only B, or both A and B. The terms "at least one of A or B" or "at least one of A and B" have similar meanings to "A and / or B" and may include only A, only B, or both A and B. The singular forms "a" and "the" may also include the plural forms.
[0032] An automated driving system (ADS) is a system that continuously performs all dynamic driving tasks (DDT) within its operational domain design (ODD). Specifically, the system is only allowed to fully assume the task of autonomous vehicle control under specified appropriate driving scenarios. When the vehicle meets the ODD conditions, the system is activated, replacing the human driver as the vehicle's primary driver. The DDT refers to the continuous lateral (left and right steering) and longitudinal motion control (acceleration, deceleration, and constant speed) of the vehicle, as well as the detection and response to objects and events in the vehicle's driving environment. The ODD refers to the conditions under which the automated driving system can operate safely. These conditions can include geographic location, road type, speed range, weather, time of day, and national and local traffic laws and regulations.
[0033] Correctly predicting the target's status in the ADS perception scenario provides complete and accurate environmental information for downstream functional modules such as PNC and AEB, helping to improve the user experience.
[0034] Therefore, in order to accurately predict the state of the target, this application provides the following technical solutions:
[0035] Referring to FIG1 , FIG1 is a flow chart showing the main steps of a method for detecting a target state according to an embodiment of the present application. As shown in FIG1 , the method for detecting a target state in the embodiment of the present application mainly includes the following steps 101 to 104 .
[0036] Step 101: Extract features from each current target and each first historical target to obtain first feature information of the current target and historical feature information of the first historical target.
[0037] In a specific implementation, cameras, radars, and other devices can be used to collect multi-frame perception data of the current scene, such as images and point clouds. This multi-frame perception data is then input into a pre-trained target detection model to detect targets, obtaining the targets in the current environment in each time frame. The current target in the current time frame and the first historical target in multiple historical time frames are then derived. The multi-frame perception data includes the perception data for the current time frame and the perception data for each historical time frame.
[0038] Specifically, after obtaining the target detection results in each time frame output by the target detection model, the preset time sliding window can be used to slide and intercept the output of the target detection model to obtain the current target in the current time frame and the first historical target in multiple historical time frames, so as to obtain all observations in a past period of time more comprehensively to comprehensively extract features and ensure the accuracy of the later state prediction. Among them, the time series length of the preset time sliding window can be set according to actual needs. In theory, the larger the time series length T of the preset time sliding window, the more information the model refers to the historical frame, and the smoother and more stable the output. However, the benefits of increasing the time series length T are marginally decreasing, the benefits brought by historical frames that are too far away are relatively limited, and the computing power overhead increases exponentially. Therefore, this method can set the time series length to, but not limited to, T=10, so that the accuracy of the prediction results and the balance between computing power are achieved.
[0039] In a specific implementation process, based on the network structure shown in Figure 2, an encoder can be used to extract features of each current target in the current time frame and each first historical target in the historical time frame to obtain the first feature information of each current target and the feature information of each first historical target. That is, the current target and the first historical target can be encoded to obtain the first feature information and the historical feature information of the first historical target.
[0040] Figure 2 is a schematic diagram of the network structure for detecting the target state of the present application. As shown in Figure 2, the network structure may include K encoders, J decoders, a first masked attention information network (Asso Head network), a second masked attention information network and a state prediction network (State Head network). The input is the current target of the current time frame and the first historical target of the historical time frame in step 101. The mapping vector of the input data is obtained through the multi-layer perceptron MLP. Under the second masked attention information (ST Attention Mask), the feature information Memory of the target Key in all time frames can be output through the softmax activation function, Add&LN network, FFN network, etc. The feature information Memory of the target Key in all time frames can include the first feature information of the current target in the current time frame and the historical feature information of the first historical target in the historical time frame. Tgt is the first feature information of the current target Query in the current time frame, and the last M can be taken from the Memory. The Asso Head network outputs the association matching result of the current target Query (that is, the association relationship between the current target Query and all target Keys) as the first mask attention information (Association Mask), wherein the first mask attention information is used to indicate the first historical target that matches the current target. In this way, the first mask attention information can be used to guide the decoder Decoder to pay attention only to the first historical target that matches itself, that is, only pay attention to its own historical state information. The multi-layer perceptron MLP, softmax activation function, Add&LN network, FFN network, etc. in the decoder Decoder can obtain the second feature information of each current target and input it into the State Head network, and then output the state information of each current target through the State Head network.
[0041] The input of the network structure of this embodiment is the target detected by the target detection model, that is, the network structure only performs state prediction. In this way, the separation of target detection and state prediction is achieved, making each model lightweight, so that it can be more suitable for actual product development practice, easier to iterate, and solve problems in modules. It can be affiliated with two teams for optimization and then merged at the time of final deployment. Even if there are multiple modal data at the same time, the detection results of multiple modal data are used as a whole time series data to input the network structure for subsequent state prediction. There is no need to set different branches for feature extraction for different modal data, which reduces the complexity of the network structure. In Figure 2, M represents the maximum number of targets in a time frame, a total of T time frames (current time frame T and historical time frame T-1), and N=T*M represents the maximum number of targets in all time frames. C represents the dimension after decoding the current target.
[0042] In a specific implementation process, when predicting a matching relationship, if the similarity of the feature information of the two targets in time, space, category, size, etc. meets the preset conditions, it is unlikely that the two will form a matching relationship, and they can be eliminated from network training and reasoning. Specifically, the above-mentioned preset conditions can be that, for two targets, if part of the feature information of the two targets is similar (such as the type, size, etc. feature information is similar), but the time interval between the two targets is detected exceeds the preset time length, it can be determined that it is impossible to form a matching relationship between the two. For another example, if part of the feature information of the two targets is similar, but the two targets appear in different spaces, it can also be determined that it is impossible to form a matching relationship between the two. For another example, if part of the feature information of the two targets is similar (such as the time and space feature information is similar), but the two types are completely different, or the size difference is greater than the preset difference value, it can also be determined that it is impossible to form a matching relationship between the two.
[0043] Based on this, the matrix corresponding to the matching possibility can be calculated first as the preset second mask attention information. The second mask attention information is used to indicate that there is no matching possibility with the current target. In this way, the current target and the first historical target can be feature extracted under the preset second mask attention information to obtain the first feature information and the historical feature information of the first historical target. In this way, some first historical targets that have no matching possibility with the current target can be eliminated as much as possible, and then feature extraction can be performed on the first historical target.
[0044] In a specific implementation process, spatiotemporal feature embedding can be implemented by the spatiotemporal positional embedding (STPE) module. Through STPE, the Transformer network can further distinguish elements of different positions and time frames, and better handle the dependencies in the sequence in the attention mechanism. STPE can be fixed or learnable. Fixed STPE can be generated by some mathematical functions (such as sine and cosine functions), while learnable STPE can be updated by back propagation during training. This embodiment can use fixed sine and cosine functions.
[0045] Therefore, spatiotemporal feature embedding and encoding can also be performed on each current target in the current time frame and each first historical target in the historical time frame to obtain first feature information of each current target and feature information of each historical target. The spatiotemporal position embedding is used to enhance the feature representation of the current target and the feature representation of the first historical target.
[0046] It should be noted that, in this embodiment, required feature information can be extracted according to actual needs, so that the acquisition of richer feature information can help improve the accuracy of Association Mask prediction and further help the prediction of the state.
[0047] Step 102: Determine first mask attention information of the current time frame based on the first feature information and the historical feature information of the first historical target;
[0048] In a specific implementation process, the state prediction of the target is generally only related to its own historical state information. Therefore, based on the first feature information of each current target and the historical feature information of each first historical target, each current target can be associated and matched, and the association matching result of each current target is obtained as the first mask attention information of the current time frame. Among them, the Association Mask can be, but is not limited to, a 0 / 1 matrix of shape = (M, N), which is learned by the Asso Head network. Each element in the Mask represents the [matching] between the current target Query and all Keys (1 represents no match, 0 represents a match). For a current target Query, there is at most one associated target at each moment, and it must be associated with itself at the current moment. Therefore, each row of the Mask has at most T zeros and at least 1 zero. That is, when all historical targets of all historical frames can be associated with the current target, there can be T zeros, and when it is only associated with itself, there is only one 0.
[0049] Specifically, under the preset second mask attention information, based on the first feature information of each current target and the historical feature information of each first historical target, the historical feature information of the first historical target can be filtered out to obtain the historical feature information of the second historical target, so as to facilitate the training convergence and reasoning performance of the encoder, and based on the current feature information of the current target and the historical feature information of the second historical target, an associated matching result of the current target can be obtained. The second historical target is a historical target that has a possibility of matching with the current target.
[0050] Step 103: Obtain second feature information of the current target based on the first mask attention information, the first feature information, and the historical feature information of the first historical target;
[0051] In a specific implementation process, after the Asso Head network obtains the Association Mask, it can be input into the decoder to guide the decoder to only focus on the first historical target that matches itself, and decode based on the first feature information of each current target and the historical feature information of each first historical target to obtain the second feature information of each current target, so as to input the second feature information of each current target into the State Head network, and then output the state information of each current target through the State Head network.
[0052] Step 104: Based on the second feature information, perform state prediction on the current target to obtain state information of the current target.
[0053] In a specific implementation process, the state information of the current target may include orientation information. At this time, the second feature information can be input into the orientation head network in the state prediction network for prediction to obtain the orientation information of the current target. The orientation information of the current target may include the orientation angle, yaw angle, pitch angle, roll angle, etc. of the current target. The orientation angle refers to the angle formed by rotating the target direction line of the target object with the north or south direction as the starting direction, with the position of the target object as the center. The target direction line can point to the direction of movement of the target object. The pitch angle refers to the angle between the direction of movement of the target object and the horizontal plane. The yaw angle refers to the angle between the projection direction of the direction of movement of the target object on the horizontal plane and the predetermined direction on the horizontal plane. The predetermined direction can be set to the road direction. The roll angle is used to represent the lateral inclination angle.
[0054] The target state detection method of this embodiment obtains the first feature information of the current target and the historical feature information of the first historical target by extracting features from each current target in the current time frame and each first historical target in the historical time frame; determines the first mask attention information of the current time frame based on the first feature information of the current target and the historical feature information of the first historical target; obtains the second feature information of the current target based on the first mask attention information, the first feature information of the current target and the historical feature information of the first historical target; and performs state prediction on each current target based on the second feature information of the current target to obtain the state information of each current target. In this way, the separation of target detection and state prediction is achieved, making each model lightweight, and any modal data can be used as input, without the need to adapt different modal data for feature extraction, reducing the complexity of the network structure, and in each detection, the state of the current target of the current time frame can be independently detected under the first mask attention information, without being affected by the upstream detection results, and the obtained state information is more accurate.
[0055] Figure 3 illustrates a practical scenario using conventional target state detection methods. As shown in Figure 3, the right side of the image shows the actual scene in which the ego vehicle is located. In this scene, a stationary electric bicycle is located to the left and in front of the ego vehicle. However, due to a perception error, the speed of the electric bicycle entering or exiting the lane is predicted to be intrusive, resulting in AEB, which affects the user's autonomous driving experience and, more seriously, may cause personal injury.
[0056] Figure 4 is a schematic diagram of an actual scenario using the target state detection method of the present application. As shown in Figure 4, the input of the target state detection method of the present application is the current target in the current time frame and the first historical target in the historical time frame (not shown in the figure) detected by the target detection model, and the output is the speed information and / or direction information of the current target in the current time frame. In this scenario, there is a moving target A in the lane r1 where the vehicle is located, and its direction is the same as the direction of travel of the vehicle. There is a moving target B in the opposite lane r2 of the lane where the vehicle is located, and its direction is opposite to the direction of travel of the vehicle. There are a first stationary object C and a second stationary object D on the roadside. The target state detection methods of the present application all made correct predictions, as shown in Figure 4, where the speeds of the moving target A and the moving target B are marked with arrows, and the moving directions of the moving target A and the moving target B are opposite.
[0057] Figure 5 is a schematic diagram of another actual scenario using the target state detection method of the present application. As shown in Figure 5, the right side of Figure 5 is the actual scene where the vehicle is located. In this scene, there is a stationary electric bicycle in front of the right side of the vehicle, and there are multiple moving vehicles in front of the left side and in front of the vehicle. The left side of Figure 5 shows the state information of multiple targets detected by the vehicle in this scene. It can be determined from Figure 5 that a single stationary vulnerable road user (Vulnerable Road User, VRU) is observed in front of the right side of the vehicle, and there is no abnormal speed that invades the lane (there is no detection box with an arrow in the figure), so AEB will not be triggered. The vehicles in front of the left side and in front of the vehicle have speed (detection boxes with arrows in the figure), which can provide complete and accurate environmental information for PNC.
[0058] FIG6 is a schematic diagram of another actual scenario using the target state detection method of the present application. As shown in FIG6 , the right side of FIG6 is the actual scene in which the ego vehicle is located. In this scene, there are some stationary electric bicycles and pedestrians in front of the right side of the ego vehicle, and there are multiple moving vehicles in front of the left side of the ego vehicle. The left side of FIG6 shows the state information of the multiple targets detected by the ego vehicle in this scene. FIG6 shows that a group of VRUs is observed in front of the right side of the ego vehicle, and there is no abnormal speed that would intrude into the lane, so AEB will not be triggered. However, the vehicles in front of the right side of the ego vehicle and in front of the ego vehicle have speeds that can provide complete and accurate environmental information for PNC.
[0059] Figure 7 is a schematic diagram of another actual scenario employing the target state detection method of the present application. As shown in Figure 7 , the right side of Figure 7 shows the actual scene in which the ego vehicle is located, with an electric bicycle moving at a certain speed to the right of the ego vehicle. The left side of Figure 7 shows the state information of multiple targets detected by the ego vehicle in this scene. Figure 7 indicates that a single VRU is observed to the right of the ego vehicle, exhibiting an abnormal speed that intrudes into the ego vehicle's lane, thus triggering AEB.
[0060] Figure 8 is a schematic diagram of another practical scenario employing the target state detection method of the present application. As shown in Figure 8 , the right side of Figure 8 shows the actual scene in which the ego vehicle is located, with multiple vehicles traveling at various angles in front of the ego vehicle. The left side of Figure 8 shows the state information of multiple targets detected by the ego vehicle in this scene. Figure 8 can be used to determine the speeds of vehicles at various angles in front of the ego vehicle, providing complete and accurate environmental information for the PNC.
[0061] It should be pointed out that although the various steps in the above embodiments are described in a specific order, those skilled in the art will understand that in order to achieve the effect of the present application, different steps do not have to be performed in such an order. They can be performed simultaneously (in parallel) or in other orders. These changes are within the scope of protection of the present application.
[0062] It will be understood by those skilled in the art that all or part of the processes in the method for implementing the above embodiment of the present application can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, it can implement the steps of each of the above method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file or some intermediate form. The computer-readable storage medium can include: any entity or device, medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory, random access memory, electric carrier signal, telecommunication signal and software distribution medium that can carry the computer program code. It should be noted that the content contained in the computer-readable storage medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable storage media do not include electric carrier signals and telecommunication signals.
[0063] Furthermore, the present application also provides a target state detection device.
[0064] 9 , which is a block diagram of the main structure of a target state detection device according to an embodiment of the present application. As shown in FIG9 , the target state detection device in the embodiment of the present application may include a processor 91 and a storage device 92 .
[0065] The storage device 92 may be configured to store a program for executing the target state detection method of the above-described method embodiment, and the processor 91 may be configured to execute the program in the storage device 92, including but not limited to a program for executing the target state detection method of the above-described method embodiment. For ease of illustration, only the portion related to the embodiment of the present application is shown. For specific technical details not disclosed, please refer to the method section of the embodiment of the present application. The target state detection device may be a control device formed by various electronic devices.
[0066] In a specific implementation process, the number of the storage device 92 and the processor 91 can be multiple. The program for executing the target state detection method of the above method embodiment can be divided into multiple subroutines, and each subroutine can be loaded and run by the processor 91 to execute different steps of the target state detection method of the above method embodiment. Specifically, each subroutine can be stored in a different storage device 92, and each processor 91 can be configured to execute the program in one or more storage devices 92 to jointly implement the target state detection method of the above method embodiment, that is, each processor 91 executes different steps of the target state detection method of the above method embodiment to jointly implement the target state detection method of the above method embodiment.
[0067] The multiple processors 91 may be processors deployed on the same device. For example, the device may be a high-performance device composed of multiple processors, and the multiple processors 91 may be processors configured on the high-performance device. Furthermore, the multiple processors 91 may be processors deployed on different devices. For example, the device may be a server cluster, and the multiple processors 91 may be processors on different servers in the server cluster.
[0068] Furthermore, the present application also provides an intelligent device, which includes the target tracking device of the above embodiment. The intelligent device may specifically include a driving device, an autonomous vehicle, a smart car, a robot, an unmanned aircraft, etc.
[0069] In some embodiments of the present application, the smart device further includes at least one sensor configured to sense information. The sensor is communicatively coupled to any of the processors described herein. Optionally, the smart device further includes an autonomous driving system configured to guide the smart device to autonomously drive or provide assisted driving. The processor communicates with the sensor and / or autonomous driving system to perform the method described in any of the above embodiments.
[0070] Furthermore, the present application also provides a computer-readable storage medium. In a computer-readable storage medium embodiment according to the present application, the computer-readable storage medium can be configured to store a program for executing the target state detection method of the above-mentioned method embodiment, and the program can be loaded and run by the processor to implement the above-mentioned target state detection method. For ease of explanation, only the parts related to the embodiment of the present application are shown. For specific technical details not disclosed, please refer to the method part of the embodiment of the present application. The computer-readable storage medium can be a storage device formed by various electronic devices. Optionally, the computer-readable storage medium in the embodiment of the present application is a non-temporary computer-readable storage medium.
[0071] Furthermore, it should be understood that since the configuration of each module is merely for the purpose of illustrating the functional units of the apparatus of the present application, the physical devices corresponding to these modules may be the processor itself, or a portion of the software in the processor, a portion of the hardware, or a combination of software and hardware. Therefore, the number of modules in the figure is merely illustrative.
[0072] Those skilled in the art will appreciate that the various modules in the device can be adaptively split or merged. Such splitting or merging of specific modules will not cause the technical solution to deviate from the principles of this application. Therefore, the technical solutions after splitting or merging will fall within the scope of protection of this application.
[0073] It should be noted that the relevant user personal information that may be involved in the various embodiments of this application is strictly in accordance with the requirements of laws and regulations, follows the principles of legality, legitimacy and necessity, and is based on the reasonable purposes of business scenarios to process personal information that users actively provide during the use of products / services or generated due to the use of products / services, as well as personal information obtained with the user's authorization.
[0074] The user personal information processed by this application will vary depending on the specific product / service scenario and must be based on the specific scenario in which the user uses the product / service. This may involve the user's account information, device information, driving information, vehicle information, or other related information. This application will treat the user's personal information and its processing with a high degree of diligence.
[0075] This application attaches great importance to the security of user personal information and has taken reasonable and feasible security protection measures that comply with industry standards to protect user information and prevent personal information from being accessed, disclosed, used, modified, damaged or lost without authorization.
[0076] Thus far, the technical solutions of the present application have been described in conjunction with the embodiments shown in the accompanying drawings. However, it is readily understood by those skilled in the art that the scope of protection of the present application is obviously not limited to these specific embodiments. Without departing from the principles of the present application, those skilled in the art may make equivalent changes or substitutions to the relevant technical features, and the technical solutions after such changes or substitutions will fall within the scope of protection of the present application.
Claims
1. A method for detecting a target state, characterized in that: include: Perform feature extraction on each current target and each first historical target to obtain first feature information of the current target and historical feature information of the first historical target; wherein the current target is a target within a current time frame; and the first historical target is a target within each historical time frame; Determine first mask attention information of the current time frame based on the first feature information and the historical feature information of the first historical target; wherein the first mask attention information is used to indicate a first historical target that matches the current target; Based on the first mask attention information, the first feature information and the historical feature information of the first historical target, obtaining second feature information of the current target; Based on the second feature information, a state prediction is performed on the current target to obtain state information of the current target; wherein the state information includes speed information and / or direction information.
2. The method for detecting a target state according to claim 1, characterized in that: Determining first mask attention information of a current time frame based on the first feature information and the historical feature information of the first historical target includes: Based on the first feature information and the historical feature information of the first historical target, the current target is associated and matched, and the associated matching result of the current target is obtained as the first mask attention information.
3. The target state detection method according to claim 2, characterized in that: Based on the first feature information and the historical feature information of the first historical target, the current target is associated and matched to obtain an associated and matched result of the current target, including: Under the preset second mask attention information, based on the first feature information and the feature information of the first historical target, the first historical target is filtered out to obtain the feature information of the second historical target; wherein the second mask attention information is used to indicate that there is no possibility of matching with the current target; and the second historical target is a historical target that has a possibility of matching with the current target; The association matching result is obtained based on the first feature information and the feature information of the second historical target.
4. The method for detecting a target state according to claim 1, characterized in that: Performing feature extraction on each current target and each first historical target to obtain first feature information of the current target and historical feature information of the first historical target includes: The current target and the first historical target are encoded to obtain the first characteristic information and historical characteristic information of the first historical target.
5. The method for detecting a target state according to claim 1, characterized in that: Performing feature extraction on each current target and each first historical target to obtain first feature information of the current target and historical feature information of the first historical target includes: Embed and encode the current target and the first historical target in time and space to obtain the first feature information and historical feature information of the first historical target; The spatiotemporal position embedding is used to enhance the feature representation of the current target and the feature representation of the first historical target.
6. The method for detecting a target state according to claim 1, characterized in that: Performing feature extraction on each current target and each first historical target to obtain first feature information of the current target and historical feature information of the first historical target includes: Under the preset second mask attention information, feature extraction is performed on the current target and the first historical target to obtain the first feature information and historical feature information of the first historical target; Among them, the second mask attention information is used to indicate that there is no matching possibility with the current target.
7. The method for detecting a target state according to any one of claims 1 to 6, characterized in that: Before extracting features for each current target and each first historical target, the following steps are also included: The multi-frame perception data is input into the target detection model for detection, and the current target is obtained. and the first historical target; wherein the multi-time frame perception data includes perception data of the current time frame and perception data of each historical time frame.
8. A target state detection device, characterized in that: The method comprises a processor and a storage device, wherein the storage device is suitable for storing a plurality of program codes, and the program codes are suitable for being loaded and run by the processor to execute the target state detection method according to any one of claims 1 to 7.
9. A smart device, characterized in that: A device for detecting a target state comprising the device as claimed in claim 8.
10. A computer-readable storage medium, characterized in that: A plurality of program codes are stored, and the program codes are suitable for being loaded and run by a processor to execute the target state detection method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Multi-target tracking method based on Mask R-CNN and apparent feature fusion
CN113506317A
Vehicle track prediction method based on deep learning
CN115293237A
Track prediction method based on target local end point and multi-head attention mechanism
CN115376099A
Target information detection method and device, driving device and medium
CN115965944A
Target state detection method and device, intelligent device and medium
CN117668573A
Cited By
Target tracking method and device
CN120526406A