Multi-mode fusion target tracking method, device and electronic equipment based on situational awareness
Through the situational awareness multi-mode fusion target tracking method, through object detection, position matching and feature matching, the problems of low accuracy and excessive computing resource consumption in the object tracking method are solved, and more efficient object tracking is achieved.
Patent Information
- Application Number
- CN202310996592.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-09
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2043-08-09
AI Technical Summary
In the prior art, the object tracking method in the image has problems such as low accuracy in object position detection and excessive consumption of computing resources. Especially when objects overlap, it is difficult to detect and has high computational complexity when the number of feature extraction times is large.
A multi-mode fusion target tracking method based on situational awareness is adopted to generate object tracking information through object detection, position matching and feature matching, reduce the number of feature extractions, improve accuracy and reduce computing resource consumption.
It improves the accuracy of object tracking information, reduces the consumption of computing resources, and solves the problem of detection difficulty and feature extraction complexity when objects overlap.
Smart Images

Figure CN117197186B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present disclosure relate to the field of computer technology, and in particular to a multi-mode fusion target tracking method, device, and electronic device based on situational awareness. Background Art
[0002] Tracking the object information in each image can determine the state of the object information in the image. When tracking objects in images, the commonly used methods are to match the position of the objects in the image to achieve tracking of the object information in each image, or to extract features from each frame of the image and track the object information in each image information based on the extracted feature information.
[0003] However, the inventors have discovered that when the above method is used to track object information, the following technical problems often occur:
[0004] First, tracking information is generated only based on the position information of the object. When objects in the image overlap, it is difficult to detect the position information of each overlapping object, resulting in low accuracy of the detected object position information, and low accuracy of the tracking information generated based on the object position information; if feature extraction is performed on each frame of the image, the number of feature extractions is large, and if the lighting environment and image resolution of the captured image are poor, the recognition rate of the extracted feature information is low, resulting in large time complexity and space complexity required to extract image features, resulting in high consumption of computing resources.
[0005] Second, the object information in the image is tracked directly without considering the prediction and processing of the state of the object information in the image, and the tracking results are not generated based on the processed object information. As a result, tracking all the object information consumes more computing resources and takes a long time to generate the tracking results.
[0006] The above information disclosed in this Background section is only for enhancement of understanding of the background of the inventive concept and therefore it may contain information that does not form the prior art that is already known in this country to a person of ordinary skill in the art. Summary of the Invention
[0007] The content of this disclosure is used to briefly introduce concepts that will be described in detail in the detailed description section below. The content of this disclosure is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0008] Some embodiments of the present disclosure propose a multi-mode fusion target tracking method, device, electronic device and computer-readable medium based on situational awareness to solve one or more of the technical problems mentioned in the above background technology section.
[0009] In a first aspect, some embodiments of the present disclosure provide a multi-modal fusion target tracking method based on situational awareness, the method comprising: performing object detection on a target image to obtain a target object information group; in response to determining that a forward target image that meets a preset forward condition is detected, performing the following matching steps based on the obtained target object information group and the forward target object information group corresponding to the detected forward target image: performing position matching on the forward target object information group and the obtained target object information group to obtain position matching result information; determining each unmatched target object information based on the above position matching result information and the obtained target object information group; performing feature matching on each unmatched target object information based on an object information set to obtain feature matching result information; updating status information of each target object information corresponding to the obtained position matching result information and the feature matching result information; determining target object information that meets a preset error-prone condition and target object information that meets a preset stationary condition based on the obtained position matching result information and the feature matching result information; and generating object tracking information corresponding to the obtained target object information group based on the target object information that meets the above preset error-prone condition, the target object information that meets the above preset stationary condition, and the forward target image that meets the above preset forward condition.
[0010] In a second aspect, some embodiments of the present disclosure provide a multi-mode fusion target tracking device based on situational awareness, the device comprising: a detection unit, configured to perform object detection on a target image to obtain a target object information group; an execution unit, configured to, in response to determining that a forward target image that meets a preset forward condition is detected, perform the following matching steps based on the obtained target object information group and the forward target object information group corresponding to the detected forward target image: position matching the forward target object information group and the obtained target object information group to obtain position matching result information; determining each unmatched target object information group based on the above position matching result information and the obtained target object information group; information; perform feature matching on each unmatched target object information according to the object information set to obtain feature matching result information; update the status information of each target object information corresponding to the obtained position matching result information and feature matching result information; determine the target object information that meets the preset error-prone condition and the target object information that meets the preset still condition according to the obtained position matching result information and feature matching result information; a generation unit is configured to generate object tracking information corresponding to the obtained target object information group according to the target object information that meets the above-mentioned preset error-prone condition, the target object information that meets the above-mentioned preset still condition and the forward target image that meets the above-mentioned preset forward condition.
[0011] In a third aspect, some embodiments of the present disclosure provide an electronic device comprising: one or more processors; a storage device on which one or more programs are stored, and when the one or more programs are executed by one or more processors, the one or more processors implement the method described in any implementation of the first aspect above.
[0012] In a fourth aspect, some embodiments of the present disclosure provide a computer-readable medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the method described in any implementation of the first aspect is implemented.
[0013] The above-mentioned embodiments of the present disclosure have the following beneficial effects: Through the multi-modal fusion target tracking method based on situational awareness of some embodiments of the present disclosure, the consumption of computing resources is reduced and the accuracy of the generated object tracking information is improved. Specifically, the reasons for the low accuracy of the generated object tracking information and the high consumption of computing resources are: when tracking information is generated only based on the position information of the object, when objects in the image overlap, it is difficult to detect the position information of each overlapping object, resulting in low accuracy of the detected object position information, and thus low accuracy of the tracking information generated based on the object position information; if feature extraction is performed on each frame of the image, the number of feature extractions is large, and if the lighting environment and image resolution of the captured image are poor, the recognition rate of the extracted feature information is low, resulting in high time and space complexity required for extracting image features, resulting in high consumption of computing resources. Based on this, the multi-modal fusion target tracking method based on situational awareness of some embodiments of the present disclosure first performs object detection on the target image to obtain a target object information group. In this way, the information of each object in the target image can be detected. Then, in response to determining that a forward target image meeting a preset forward condition has been detected, the following matching steps are performed based on the obtained target object information group and the forward target object information group corresponding to the detected forward target image: First, position matching is performed on the forward target object information group and the obtained target object information group to obtain position matching result information. This results in object information in the forward target image and position matching result information for the target object information group. Next, based on the position matching result information and the obtained target object information group, each piece of target object information that did not match is determined. This results in each piece of target object information that did not successfully match, which can be used to generate feature matching result information. Next, based on the object information set, feature matching is performed on each piece of target object information that did not match, to obtain feature matching result information. This results in a feature matching result that meets the feature matching condition. Next, the obtained position matching result information and the state information of each piece of target object information corresponding to the feature matching result information are updated. This results in the updated state information of each piece of target object information. Next, based on the obtained position matching result information and the obtained feature matching result information, target object information that meets the preset error-prone condition and target object information that meets the preset stationary condition are determined. Thus, target object information representing information that satisfies the preset error-prone condition and target object information that satisfies the preset stationary condition can be obtained, which can be used to generate object tracking information corresponding to the target object information group. Finally, based on the target object information that satisfies the preset error-prone condition, the target object information that satisfies the preset stationary condition, and the forward target image that satisfies the preset forward condition, object tracking information corresponding to the obtained target object information group is generated. Thus, tracking results for the target object information group and the forward target object information group can be obtained.Because the tracking results are generated based on both location and feature information, rather than solely on location, the accuracy of the generated tracking results is improved. Furthermore, because feature extraction is performed on individual unmatched target objects, rather than on all target objects, the number of feature extractions is reduced, thus reducing the computational resources consumed. This improves the accuracy of the generated object tracking information and reduces computational resource consumption. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that components and elements are not necessarily drawn to scale.
[0015] Figure 1 is a flowchart of some embodiments of the multi-mode fusion target tracking method based on situational awareness according to the present disclosure;
[0016] Figure 2 is a schematic structural diagram of some embodiments of a multi-mode fusion target tracking device based on situational awareness according to the present disclosure;
[0017] Figure 3 It is a structural diagram of an electronic device suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION
[0018] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0019] It should also be noted that, for ease of description, only the parts related to the invention are shown in the drawings. In the absence of conflict, the embodiments and features in the embodiments of the present disclosure may be combined with each other.
[0020] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0021] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".
[0022] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0023] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.
[0024] Figure 1 The process 100 of some embodiments of the multi-mode fusion target tracking method based on situational awareness according to the present disclosure is shown. The multi-mode fusion target tracking method based on situational awareness includes the following steps:
[0025] Step 101: perform object detection on a target image to obtain a target object information group.
[0026] In some embodiments, the execution subject (such as a computing device) of the multi-modal fusion target tracking method based on situational awareness can perform object detection on the target image to obtain a target object information group. The above-mentioned target image can be any frame image in the target video stream. The above-mentioned target video stream can be any captured video. The above-mentioned target object information group can characterize the information of each object included in the above-mentioned target image. The target object information in the above-mentioned target object information group can include the category and code of the object. For example, the category of the object can be "car". The code can be 01. In practice, the above-mentioned execution subject can perform image recognition on the above-mentioned target image to obtain the information of each object included in the above-mentioned target image.
[0027] Optionally, after step 101, first, the execution entity may further detect the change target image in response to determining that a change target image that satisfies a preset change condition is detected, and obtain a change target object information group. The preset change condition may be that the frame sequence of the change target image is greater than the frame sequence of the target image. The change target image may represent the detected updated target image. The change target object information group may represent information about each object in the change target image. For example, the change target object information in the change target object information group may include the category and code of the object. For example, the category of the object may be "building." The code may be 23.
[0028] Then, the obtained changed target object information group can be used as the target object information group to perform the above matching step again.
[0029] Optionally, before step 101, the execution entity may also, in response to determining that no forward target image satisfying the preset forward condition has been detected, add each target object information in the obtained target object information group as object information to the object information set, thereby obtaining a changed object information set as the object information set. The preset forward condition may be that the frame sequence of the forward target image is less than the frame sequence of the target object image. The forward target image may represent the previous frame image of the target image. The object information set may represent each object information extracted from the image. The initial value of the object information set may be empty. The changed object information set may represent the changed object information set.
[0030] Step 102: In response to determining that a forward target image that meets a preset forward condition is detected, the following matching steps are performed based on the obtained target object information group and the forward target object information group corresponding to the detected forward target image:
[0031] Step 1021 : Perform position matching on the forward target object information group and the obtained target object information group to obtain position matching result information.
[0032] In some embodiments, the execution entity may perform position matching on the forward target object information group and the obtained target object information group to obtain position matching result information. The position matching result information may indicate a matching result between the coordinates of the forward target object information in the forward target object information group and the coordinates of the target object information in the target object information group. The position matching result information may indicate, but is not limited to, one of the following: a successful match between the forward target object information and the target object information; an unsuccessful match between the forward target object information and the target object information; or an unsuccessful match between the target object information and the forward target object information.
[0033] In some optional implementations of some embodiments, the execution entity may perform position matching on the forward target object information group and the obtained target object information group through the following steps to obtain position matching result information:
[0034] In the first step, the target image is input into a pre-trained object center position information generation model to obtain first position information corresponding to the obtained target object information group. The pre-trained object center position information generation model may be a model that generates object center coordinates. The first position information in each of the first position information may represent the two-dimensional coordinates of the object center in the target image. For example, the object center position information generation model may be a convolutional neural network model.
[0035] In the second step, the forward target image corresponding to the forward target object information group is input into the object center position information generation model to obtain each piece of second position information corresponding to the forward target object information group. The second position information in each piece of second position information may represent the two-dimensional coordinates of the object center in the forward target image corresponding to the forward target object information group. The second position information in each piece of second position information corresponds to the forward target object information in the forward target object information group.
[0036] In the third step, for each target object information in the obtained target object information group, perform the following position matching steps:
[0037] The first sub-step is to determine the first position information corresponding to the target object information.
[0038] In the second sub-step, based on the first position information, it is determined whether any of the obtained second position information satisfies a preset position condition. The preset position condition may be that the distance between the first position information and the second position information is less than a preset distance value. The specific numerical value of the preset distance value is not limited. In practice, first, the execution entity may determine the distance value between the coordinates of the first position information and each of the second position information to obtain each distance value. Then, it may be determined whether any of the obtained distance values satisfies the preset position condition.
[0039] In the third sub-step, in response to determining that there is second location information that meets the preset location condition among the obtained second location information, the second location information that meets the preset location condition is determined as the target location information.
[0040] In a fourth sub-step, the forward target object information corresponding to the target position information in the obtained forward target object information group is determined as the target forward target object information, wherein the target forward target object information may be the forward target object information that meets the preset position condition.
[0041] In a fifth sub-step, the determined target forward target object information and the target object information are input into a pre-trained object information matching model to obtain matching information corresponding to the target object information and the determined target forward target object information. The object information matching model may be a model that generates matching information between forward target object information and target object information. For example, the object information matching model may be a convolutional neural network model. The matching information may be a degree of matching between the forward target object information and the target object information. The degree of matching may represent a similarity between the forward target object information and the target object information.
[0042] In a sixth sub-step, in response to determining that the obtained matching information satisfies a preset matching information condition, determining the first preset matching result information as the first matching result information corresponding to the target object information. The preset matching information condition may be that the matching information represents a degree of match greater than 80 percent. The first preset matching result information may indicate a successful match between the forward target object information and the target object information.
[0043] The seventh sub-step is to add the obtained first matching result information to the first matching result information set to update the first matching result information set. The first matching result information set may represent each matching result in which the forward target object information successfully matches the target object information.
[0044] In an eighth sub-step, the determined target forward target object information and the target object information are deleted from the obtained forward target object information group and the target object information group, respectively, to obtain updated forward target object information group and target object information group. In practice, the execution entity may, based on the ID codes included in the target forward target object information and the target object information, delete from the obtained forward target object information group the forward target object information corresponding to the ID code of the target forward target object information and the target object information from the target object information group.
[0045] In a fourth step, the second preset matching result information is determined as the second matching result information corresponding to each piece of forward target object information in the updated forward target object information group, thereby obtaining each piece of second matching result information. The second preset matching result information may indicate that the forward target object information was not successfully matched. The forward target object information in the updated forward target object information group may indicate unsuccessfully matched forward target object information.
[0046] In the fifth step, each second matching result information is added to the second matching result information set to update the second matching result information set. The second matching result information in the second matching result information set may indicate that the forward target object information is not successfully matched.
[0047] In a sixth step, the third preset matching result information is determined as the third matching result information for each target object information in the updated target object information, thereby obtaining each piece of third matching result information. The third preset matching result information may indicate that the target object information was not successfully matched. The target object information in the updated target object information may be target object information that was not successfully matched.
[0048] In the seventh step, each obtained third matching result information is added to the third matching result information set to update the third matching result information set. The third matching result information set may represent each matching result of each target object information that was not successfully matched.
[0049] In the eighth step, the updated first matching result information set, the second matching result information set, and the third matching result information set are determined as position matching result information.
[0050] Step 1022 : Determine each unmatched target object information based on the position matching result information and the obtained target object information group.
[0051] In some embodiments, the execution entity may determine each unmatched target object information based on the position matching result information and the obtained target object information group. The unmatched target object information may represent each unmatched forward target object information in the forward target image and each unmatched target object information in the target image.
[0052] In some optional implementations of some embodiments, the execution entity may determine each unmatched target object information based on the position matching result information and the obtained target object information group through the following steps:
[0053] In the first step, in response to determining that the third matching result information set and the second matching result information set in the above-mentioned position matching result information are not empty, the target object information corresponding to the above-mentioned third matching result information set and the forward target object information corresponding to the above-mentioned second matching result information set in the obtained target object information group are determined as unmatched target object information, thereby obtaining each unmatched target object information.
[0054] Step 1023 : Perform feature matching on each unmatched target object information based on the object information set to obtain feature matching result information.
[0055] In some embodiments, the execution entity may perform feature matching on each unmatched target object information based on the object information set to obtain feature matching result information. The feature matching result information may be the result of matching based on the object's feature information. The feature matching result information may represent, but is not limited to, one of the following: a successful match between the forward target object information and the target object information, an unsuccessful match between the forward target object information and the target object information, or an unsuccessful match between the forward target object information and the target object information.
[0056] In some optional implementations of some embodiments, the execution entity may perform feature matching on each unmatched target object information based on the object information set to obtain feature matching result information through the following steps:
[0057] In the first step, feature extraction processing is performed on each piece of object information in the obtained object information set to obtain feature information of each object corresponding to the obtained object information set. In practice, the execution entity may input each piece of object information in the obtained object information set into a pre-trained feature extraction model to obtain feature information of each object corresponding to the object information set. The feature extraction model may be a model for extracting feature information of objects in an image. The object feature information in each piece of object feature information may include, but is not limited to, one of the following: shape or color of the object. For example, the feature extraction model may be a neural network model.
[0058] The second step is to perform feature extraction processing on each piece of unmatched target object information to obtain feature information of each target object. In practice, the execution entity may input each piece of unmatched target object information into the feature extraction model to obtain feature information of each object as the feature information of each target object. The feature information of each piece of unmatched target object information may represent the characteristics of the target object information. The feature information of each piece of unmatched target object information may include, but is not limited to, one of the following: shape or color of the object.
[0059] In the third step, for each target object feature information obtained, perform the following steps:
[0060] The first sub-step is to determine whether there is object feature information that meets a preset feature condition among the obtained object feature information. The preset feature condition may be object feature information that is the same as the target object feature information.
[0061] In a second sub-step, in response to determining that object feature information that satisfies the preset feature condition exists among the obtained object feature information, fourth preset matching result information is determined as fourth matching result information corresponding to the target object feature information. The fourth matching result information may indicate a successful match between the forward target object information and the target object information.
[0062] In a third sub-step, the obtained fourth matching result information is added to the fourth matching result information set to obtain the changed fourth matching result information set as the fourth matching result information set. The fourth matching result information set may be each matching result information that is successfully matched based on the object feature information.
[0063] In a fourth sub-step, the target object information corresponding to the target object feature information is deleted from each unmatched target object information, so as to update each unmatched target object information.
[0064] In a fifth sub-step, the object feature information satisfying the above-mentioned preset feature conditions is deleted from the obtained feature information of each object, so as to update the obtained feature information of each object.
[0065] In a sixth sub-step, in response to determining that no object feature information that satisfies the preset feature condition exists in the obtained object feature information, the fifth preset matching result information is determined as the fifth matching result information corresponding to the target object feature information. The fifth matching result information may indicate that the forward target object information was not successfully matched or that the target object information was not successfully matched.
[0066] In a seventh sub-step, the fifth matching result information corresponding to the target object feature information is added to the fifth matching result information set, thereby obtaining a modified fifth matching result information set as the fifth matching result information set. The fifth matching result information set may represent each fifth matching result information that was not successfully matched based on the feature information.
[0067] In the fourth step, the obtained fifth matching result information set and the fourth matching result information set are determined as feature matching result information.
[0068] Step 1024: Update the status information of each target object information corresponding to the obtained position matching result information and feature matching result information.
[0069] In some embodiments, the execution subject may update the status information of each target object information corresponding to the obtained position matching result information and feature matching result information. The status information may characterize the real-time status of the target object. For example, the status information may include the position information of the target object and whether the target object in the image still exists. In practice, the execution subject may update the position information of the target object information corresponding to the first matching result information according to the filter, or update the status of the forward target object information corresponding to the second matching result information of the matching result information to "object missing", or add the target object information corresponding to the third matching result information as object information to the object information set. For example, the filter may be a Kalman filter.
[0070] Step 1025 : Determine target object information that meets a preset error-prone condition and target object information that meets a preset stationary condition based on the obtained position matching result information and feature matching result information.
[0071] In some embodiments, the execution entity may determine target object information that satisfies a preset error-prone condition and target object information that satisfies a preset stationary condition based on the obtained position matching result information and feature matching result information. In practice, the execution entity may determine target object information that satisfies a preset error-prone condition and target object information that satisfies a preset stationary condition based on the obtained position matching result information and feature matching result information in various ways.
[0072] In some optional implementations of some embodiments, the execution entity may determine target object information that meets a preset error-prone condition and target object information that meets a preset stationary condition based on the obtained position matching result information and feature matching result information through the following steps:
[0073] The first step is to determine each target object information that meets the preset matching conditions based on the obtained position matching result information and feature matching result information. The preset matching conditions can be that the position matching result information of the target object information indicates that the forward target object information successfully matches the target object information, or the feature matching result information of the target object information indicates that the forward target object information successfully matches the target object information, or the position matching result information of the forward target object information indicates that the forward target object information is not successfully matched, or the feature matching result information of the forward target object information indicates that the forward target object information is not successfully matched. In practice, the execution entity can select each target object information that meets the preset matching conditions from each target object information corresponding to the position matching result information. The execution entity can also select each target object information that meets the preset matching conditions from each target object information corresponding to the feature matching result information. In this way, each target object information that meets the preset matching conditions can be obtained.
[0074] The second step is to determine the state information of each target object information in each target object information that meets the preset matching conditions, and obtain each state information. In practice, the above-mentioned execution entity can obtain the state information of each target object information in each target object information based on a computer vision algorithm to obtain each state information. The above-mentioned computer vision algorithm can be a local tracker. The above-mentioned local tracker can be used to track the position and movement of the target object in the video. Each state information in the above-mentioned state information can include but is not limited to one of the following: position change information, historical speed information, historical intersection-over-union information, and position information between objects.
[0075] In the third step, for each of the determined status information, the following verification steps are performed:
[0076] In the first sub-step, the target object information corresponding to the state information is determined as the first target object information.
[0077] The second sub-step involves determining velocity information corresponding to the first target object information based on the position change information included in the status information. The position change information may represent the distance value by which the coordinates of the first target object information in the forward target image and the target image change. The distance value may include a first target distance value and respective second target distance values. The first target distance value may represent the distance difference between the center point of the first target object in the forward target image and the center point of the first target object in the target image. The second target distance difference value may represent the distance difference between the velocity component of the first target object information in the target direction in the forward target image and the velocity component of the first target object information in the target direction in the target image. The target direction may be, but is not limited to, one of the following: east, south, west, or north. The respective second target distance values may represent respective distance differences in the four directions of east, south, west, and north. The velocity information may represent the velocity of the center point of the first target object information and the velocity components of the first target object information in the four directions of east, south, west, and north. In practice, the execution entity may first determine the coordinate change information of the first target object information. Then, the interval between the forward target image and the target image can be determined as the target interval. Subsequently, the ratio of the distance difference corresponding to the center point of the first target object information to the target interval can be determined as the velocity of the center point of the first target object information. Next, the ratios of the distance differences corresponding to the first target object information in the four directions of east, south, west, and north to the target interval can be determined as the velocity components of the first target object information in those directions. The distance differences in the distance differences correspond to the ratios in the ratios.
[0078] The third sub-step is to determine the maximum intersection-and-union (IoU) information corresponding to the first target object information based on the inter-object position information included in the state information. The information of each target object in the current frame image that is different from the first target object information is used as different target object information. The inter-object position information may include the coordinate information of the first target object information and the coordinate information of each different target object information. The maximum IoU information may represent the maximum IoU information among the IoU information of the first target object information and each different target object information. The maximum IoU information may be the ratio of the intersection area and the union area of the bounding box of the first target object information and the bounding box of the different target object information. In practice, the execution entity may first determine the IoU information of the first target object information and each different target object information to obtain the respective IoU information. Then, the maximum IoU information among the obtained IoU information may be determined as the maximum IoU information.
[0079] A fourth sub-step is to determine, based on the historical speed information included in the status information, whether the projection angle information of the first target object information satisfies a preset angle value condition. The historical speed information may include the speed of the center point of the first target object information within the forward target image or forward numerical frame image, as well as the velocity components of the first target object information in the four directions of east, south, west, and north. The projection angle information may be the angle value of the projection of the first target object information on any vertical component in the coordinate system. The preset angle value condition may be that the angle value of the projection of the first target object information on any vertical component in the coordinate system is greater than 0 and less than a preset angle value. The value of the preset angle value is not limited herein. The number of forward numerical frame images within the forward numerical frame image is not limited herein. For example, the coordinate system of the first target object information may be a Cartesian coordinate system. In practice, first, in response to determining that the angle value represented by the projection angle information of the first target object information is greater than 0 and less than the preset angle value, the execution entity may determine that the projection angle information of the first target object information satisfies the preset angle value condition. Then, in response to determining that the angle value represented by the projection angle information of the first target object information is less than 0 or greater than or equal to the preset angle value, the execution entity may determine that the projection angle information of the first target object information does not meet the preset angle value condition.
[0080] A fifth sub-step, in response to determining that the projection angle information of the first target object information satisfies the preset angle value condition, determines whether the historical speed information satisfies a first preset historical speed condition. The first preset historical speed condition may be that the speed components of the first target object information in the four directions of east, south, west, and north included in the historical speed information are all less than preset speed values. The value of the preset speed value is not limited herein.
[0081] A sixth sub-step, in response to determining that the historical speed information satisfies the first preset historical speed condition, determines the first target object information as target object information that satisfies the preset error-prone condition.
[0082] The fourth step is to perform blocking processing on each piece of determined first target object information. In practice, the execution entity may perform blocking processing on the feature information of the first target object information within a preset time period. The blocking processing may indicate that the state information of the first target object information is determined to be a suspended state. The suspended state may indicate that the execution entity does not need to perform feature extraction on each piece of first target object information within the preset time period, and does not need to perform position information matching processing or feature information matching processing on the first target object information within the preset time period. There is no limitation on the specific time period of the preset time period.
[0083] Optionally, in the above-mentioned inspection step, first, the above-mentioned execution entity may also determine whether the line segment information corresponding to the above-mentioned first target object information satisfies a preset intersection condition based on the historical speed information included in the above-mentioned status information in response to determining that the projection angle information of the above-mentioned first target object information does not satisfy the above-mentioned preset angle value condition. The above-mentioned preset intersection condition may be that the line segment information of the above-mentioned first target object information and the line segment information of the different target object information have an intersection point. The line segment information of the above-mentioned first target object information may be a line starting from the center point of the above-mentioned first target object information in the forward target image and ending at the center point of the above-mentioned first target object information in the target image. The line segment information of the above-mentioned different target object information may be a line starting from the center point of the above-mentioned different target object information in the forward target image and ending at the center point of the above-mentioned different target object information in the target image.
[0084] Then, based on the historical I / O information included in the above-mentioned status information, it can be determined whether the target I / O information corresponding to the above-mentioned first target object information meets the preset change condition. Among them, the above-mentioned historical I / O information can include the individual I / O information corresponding to the above-mentioned first target object information in the preset numerical frame image before the above-mentioned target image and the I / O information corresponding to the first target object information in the above-mentioned target image. The above-mentioned target I / O information can include the maximum I / O information corresponding to the above-mentioned first target object information in each image in the above-mentioned preset numerical frame image. The above-mentioned preset change condition can be that the difference between the maximum I / O information and the minimum I / O information in the individual I / O information corresponding to the preset numerical frame image is greater than the preset I / O difference. The specific value of the above-mentioned preset I / O difference is not limited here. The above-mentioned preset numerical frame image can be a series of frame images. The number of images corresponding to the preset numerical frame image is not limited here.
[0085] Afterwards, in response to determining that the target intersection-over-union (IoU) information corresponding to the first target object information satisfies the preset change condition, the execution entity may determine whether the edge speed information corresponding to the first target object information satisfies the preset edge speed change condition based on the historical speed information included in the status information. The edge speed information may be a speed value corresponding to any direction of the first target object information. For example, if the arbitrary direction is east, then the opposite direction of the arbitrary direction may be west. The preset edge speed change condition may be that the change in the edge speed information of the first target object information in a preset numerical frame image is greater than a preset speed value. The value of the preset speed value is not specifically limited. In practice, the execution entity may first obtain each piece of edge speed information corresponding to the first target object information in the preset numerical frame image. Then, the execution entity may determine whether the difference between the maximum and minimum edge speed information among the obtained pieces of edge speed information is greater than a preset speed value. Then, in response to determining that the difference between the maximum and minimum edge speed information is greater than the preset speed value, the execution entity may determine that the edge speed information corresponding to the first target object information satisfies the preset edge speed change condition. Finally, in response to determining that the difference between the maximum edge speed information and the minimum edge speed information is less than or equal to the preset speed value, the execution entity may determine that the edge speed information corresponding to the first target object information does not meet the preset edge speed change condition.
[0086] Next, based on the historical I / O information of the first target object information included in the state information, it can be determined whether the maximum I / O information of the first target object information meets a preset increase condition. The maximum I / O information can be the I / O information with the highest value. The preset increase condition can be that the difference between the maximum I / O information of the corresponding target image and the maximum I / O information of the corresponding forward target image is greater than a preset I / O value. The specific value of the preset I / O value is not limited. In practice, first, the execution entity can obtain the maximum I / O information of the corresponding target image and the maximum I / O information of the corresponding forward target image from the historical I / O information. Then, in response to determining that the difference between the maximum I / O information of the target image and the maximum I / O information of the corresponding forward target image is greater than the preset I / O value, it can be determined that the maximum I / O information of the first target object information meets the preset increase condition. Then, in response to determining that the difference between the maximum I / O information of the target image and the maximum I / O information of the corresponding forward target image is less than or equal to the preset I / O value, it can be determined that the maximum I / O information of the first target object information does not meet the preset increase condition.
[0087] Secondly, in response to determining that the maximum intersection-over-joint information of the first target object information satisfies the preset increase condition, or the line segment information of the first target object information satisfies the preset intersection condition, or the edge speed information of the first target object information satisfies the preset edge speed change condition, the first target object information may be determined as first fallible target object information. The preset fallible condition may be that the maximum intersection-over-joint information of the first target object information satisfies the preset increase condition, or the line segment information of the first target object information satisfies the preset intersection condition, or the edge speed information of the first target object information satisfies the preset edge speed change condition.
[0088] Then, feature extraction processing can be performed on the determined first error-prone target object information to obtain object feature information corresponding to the determined first error-prone target object information. It should be noted that the method for performing feature extraction on the determined first error-prone target object information is the same as the method for performing feature extraction processing on each target object information in each unmatched target object information, and therefore, it will not be repeated here.
[0089] Then, the first error-prone target object information determined can be added to the object information set as object information to obtain a changed object information set as the object information set. In this way, the object information set can be updated to obtain an updated object information set.
[0090] Finally, the determined first fallible target object information may be determined as target object information that meets the above-mentioned preset fallible condition.
[0091] Optionally, in the above-mentioned inspection step, first, in response to determining that the above-mentioned first target object information does not meet the above-mentioned preset error-prone condition, the above-mentioned execution subject may determine whether the above-mentioned first target object information meets the preset combined speed condition based on the combined speed magnitude information included in the above-mentioned status information. The above-mentioned combined speed magnitude information may be the respective combined speed values of the above-mentioned first target object information in the preset numerical frame image. The above-mentioned preset combined speed condition may be that the respective combined speed values of the above-mentioned first target object information are all less than the preset combined speed value. There is no specific limitation on the size of the above-mentioned preset combined speed value. There is no limitation on the specific number of images in the preset numerical frame image.
[0092] Then, in response to determining that the first target object information satisfies the preset combined velocity condition, the first target object information may be determined as target object information that satisfies the preset stationary condition, and object feature information corresponding to the first target object information may be determined. The preset stationary condition may be that all combined velocity values corresponding to the first target object information are less than the preset combined velocity value.
[0093] Afterwards, the first target object information may be added as object information to the object information set to obtain a changed object information set as the object information set.
[0094] Finally, in response to detecting that the state information of the first target object information satisfies a preset normal matching condition within a preset time period, matching processing may be performed on the first target object information. The preset normal matching condition may be that the first target object information does not satisfy the preset static condition within the preset time period.
[0095] The above technical solution, as an inventive feature of an embodiment of the present disclosure, solves the second technical problem mentioned in the background technology: "Directly tracking object information in an image without considering the state prediction and processing of the object information in the image, and not generating tracking results based on the processed object information, resulting in a high consumption of computing resources for tracking all object information and a long time to generate tracking results." The high consumption of computing resources for tracking all object information and the long time to generate tracking results are caused by directly tracking the object information in the image without considering the state prediction and processing of the object information in the image, and not generating tracking results based on the processed object information. If these factors are resolved, the computing resources consumed for tracking all object information and the time to generate tracking results can be reduced. To achieve this effect, the present disclosure predicts each object information based on the object information's state information, processes object information that meets the preset error-prone condition and the preset stationary condition, and generates tracking result information based on the processed object information, rather than directly generating tracking result information based on the unprocessed object information. Therefore, the computational resources consumed in tracking all object information are reduced and the time consumed in generating tracking results is shortened.
[0096] Step 103 : generating object tracking information corresponding to the obtained target object information group based on the target object information meeting the preset error-prone condition, the target object information meeting the preset stationary condition, and the forward target image meeting the preset forward condition.
[0097] In some embodiments, the execution subject may generate object tracking information corresponding to the obtained target object information group based on the target object information that meets the preset error-prone condition, the target object information that meets the preset stationary condition, and the forward target image that meets the preset forward condition. The object tracking information may represent the matching result between the forward target object information group in the forward target image and the target object information group in the target image. The matching result represented by the object tracking information corresponding to the obtained target object information group may represent, but is not limited to, one of the following: the forward target object information successfully matches the target object information, the forward target object information does not successfully match, and the target object information does not successfully match. For example, the object tracking information may include that the target object information coded as 23 in the target object information group successfully matches the forward target object information coded as 23 in the forward target object information group.
[0098] In some optional implementations of some embodiments, the execution entity may generate object tracking information corresponding to the obtained target object information group based on the target object information that meets the preset error-prone condition, the target object information that meets the preset stationary condition, and the forward target image that meets the preset forward condition by performing the following steps:
[0099] The first step is to obtain the forward object tracking information corresponding to the forward target object image. The forward object tracking information can represent the matching result between the forward target object information group in the forward target object image and the upper frame target object information group in the upper frame target object image of the forward target object image. The upper frame target object information group can represent the information of each object in the upper frame target object image. The matching result represented by the forward object tracking information corresponding to the forward target object image can represent, but is not limited to, one of the following: the forward target object information and the upper frame target object information are successfully matched, the forward target object information is not successfully matched, and the upper frame target object information is not successfully matched. For example, the forward object tracking information can include that the forward target object information coded as 29 in the forward target object information group is not successfully matched.
[0100] In a second step, each target object information that does not meet the preset error-prone condition and the preset stationary condition is added to the forward object tracking information of the corresponding forward target object information group to obtain modified forward object tracking information. The modified forward object tracking information can represent the matching information of the target object information group.
[0101] In the third step, the obtained changed forward object tracking information is determined as the object tracking information corresponding to the obtained target object information group.
[0102] The above-mentioned embodiments of the present disclosure have the following beneficial effects: Through the multi-modal fusion target tracking method based on situational awareness of some embodiments of the present disclosure, the consumption of computing resources is reduced and the accuracy of the generated object tracking information is improved. Specifically, the reasons for the low accuracy of the generated object tracking information and the high consumption of computing resources are: when tracking information is generated only based on the position information of the object, when objects in the image overlap, it is difficult to detect the position information of each overlapping object, resulting in low accuracy of the detected object position information, and thus low accuracy of the tracking information generated based on the object position information; if feature extraction is performed on each frame of the image, the number of feature extractions is large, and if the lighting environment and image resolution of the captured image are poor, the recognition rate of the extracted feature information is low, resulting in high time and space complexity required for extracting image features, resulting in high consumption of computing resources. Based on this, the multi-modal fusion target tracking method based on situational awareness of some embodiments of the present disclosure first performs object detection on the target image to obtain a target object information group. In this way, the information of each object in the target image can be detected. Then, in response to determining that a forward target image meeting a preset forward condition has been detected, the following matching steps are performed based on the obtained target object information group and the forward target object information group corresponding to the detected forward target image: First, position matching is performed on the forward target object information group and the obtained target object information group to obtain position matching result information. This results in object information in the forward target image and position matching result information for the target object information group. Next, based on the position matching result information and the obtained target object information group, each piece of target object information that did not match is determined. This results in each piece of target object information that did not successfully match, which can be used to generate feature matching result information. Next, based on the object information set, feature matching is performed on each piece of target object information that did not match, to obtain feature matching result information. This results in a feature matching result that meets the feature matching condition. Next, the obtained position matching result information and the state information of each piece of target object information corresponding to the feature matching result information are updated. This results in the updated state information of each piece of target object information. Next, based on the obtained position matching result information and the obtained feature matching result information, target object information that meets the preset error-prone condition and target object information that meets the preset stationary condition are determined. Thus, target object information representing information that satisfies the preset error-prone condition and target object information that satisfies the preset stationary condition can be obtained, which can be used to generate object tracking information corresponding to the target object information group. Finally, based on the target object information that satisfies the preset error-prone condition, the target object information that satisfies the preset stationary condition, and the forward target image that satisfies the preset forward condition, object tracking information corresponding to the obtained target object information group is generated. Thus, tracking results for the target object information group and the forward target object information group can be obtained.Because the tracking results are generated based on both location and feature information, rather than solely on location, the accuracy of the generated tracking results is improved. Furthermore, because feature extraction is performed on individual unmatched target objects, rather than on all target objects, the number of feature extractions is reduced, thus reducing the computational resources consumed. This improves the accuracy of the generated object tracking information and reduces computational resource consumption.
[0103] Further references Figure 2 As an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of a multi-mode fusion target tracking method based on situational awareness. These device embodiments are similar to Figure 1 Corresponding to the method embodiments shown, the device can be specifically applied to various electronic devices.
[0104] like Figure 2 As shown, some embodiments of the multi-mode fusion target tracking generation device 200 based on situational awareness include: a detection unit 201, an execution unit 202 and a generation unit 203. The detection unit 201 is configured to perform object detection on the target image to obtain a target object information group; the execution unit 202 is configured to, in response to determining that a forward target image that meets a preset forward condition is detected, perform the following matching steps based on the obtained target object information group and the forward target object information group corresponding to the detected forward target image: perform position matching on the forward target object information group and the obtained target object information group to obtain position matching result information; determine each unmatched target object information based on the above position matching result information and the obtained target object information group; and match each unmatched target object information based on the object information set. Perform feature matching on the target object information to obtain feature matching result information; update the status information of each target object information corresponding to the obtained position matching result information and feature matching result information; determine the target object information that meets the preset error-prone condition and the target object information that meets the preset still condition based on the obtained position matching result information and feature matching result information; the generation unit 203 is configured to generate object tracking information corresponding to the obtained target object information group based on the target object information that meets the above-mentioned preset error-prone condition, the target object information that meets the above-mentioned preset still condition and the forward target image that meets the above-mentioned preset forward condition.
[0105] It is understood that the units described in the device 200 are similar to those described in the reference Figure 1 Therefore, the operations, features and beneficial effects described above for the method are also applicable to the device 200 and the units included therein, and will not be repeated here.
[0106] Reference below Figure 3 , which shows a structural diagram of an electronic device 300 (eg, a computing device) suitable for implementing some embodiments of the present disclosure. Figure 3 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0107] like Figure 3 As shown, the electronic device 300 may include a processing device 301 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. Various programs and data required for the operation of the electronic device 300 are also stored in the RAM 303. The processing device 301, the ROM 302, and the RAM 303 are connected to each other via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0108] Typically, the following devices may be connected to the I / O interface 305: an input device 306 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 308 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 309. The communication device 309 may allow the electronic device 300 to communicate with other devices wirelessly or by wire to exchange data. Figure 3 The electronic device 300 is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead. Figure 3 Each block shown in the figure may represent one device, or may represent multiple devices as needed.
[0109] In particular, according to some embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of the present disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In some such embodiments, the computer program can be downloaded and installed from a network via the communication device 309, or installed from the storage device 308, or installed from the ROM 302. When the computer program is executed by the processing device 301, the above-mentioned functions defined in the method of some embodiments of the present disclosure are performed.
[0110] It should be noted that the computer-readable medium described in some embodiments of the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or device. In some embodiments of the present disclosure, the computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0111] In some embodiments, the client and server can communicate using any currently known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.
[0112] The computer-readable medium may be included in the electronic device; or it may exist independently without being assembled into the electronic device. The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device: performs object detection on the target image to obtain a target object information group; in response to determining that a forward target image that meets a preset forward condition is detected, performs the following matching steps based on the obtained target object information group and the forward target object information group corresponding to the detected forward target image: performs position matching on the forward target object information group and the obtained target object information group to obtain position matching result information; determines each unmatched target object information based on the above position matching result information and the obtained target object information group. ; Based on the object information set, feature matching is performed on each unmatched target object information to obtain feature matching result information; the status information of each target object information corresponding to the obtained position matching result information and feature matching result information is updated; based on the obtained position matching result information and feature matching result information, the target object information that meets the preset error-prone condition and the target object information that meets the preset stationary condition are determined; based on the target object information that meets the above-mentioned preset error-prone condition, the target object information that meets the above-mentioned preset stationary condition and the forward target image that meets the above-mentioned preset forward condition, object tracking information corresponding to the obtained target object information group is generated.
[0113] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0114] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0115] The units described in some embodiments of the present disclosure may be implemented in software or hardware. The units described may also be provided in a processor, for example, they may be described as: a detection unit, an execution unit, and a generation unit. The names of these units do not, in some cases, constitute limitations on the units themselves. For example, the detection unit may also be described as "a unit that performs object detection on a target image and obtains a target object information group."
[0116] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0117] The above description is only an illustration of some preferred embodiments of the present disclosure and the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the above-mentioned inventive concept. For example, the above-mentioned features are replaced with (but not limited to) technical features with similar functions disclosed in the embodiments of the present disclosure.
Claims
1. A multi-mode fusion target tracking method based on situational awareness, comprising: Perform object detection on the target image to obtain a target object information group; In response to determining that a forward target image that meets the preset forward condition is detected, the following matching steps are performed according to the obtained target object information group and the forward target object information group corresponding to the detected forward target image: Performing position matching on the forward target object information group and the obtained target object information group to obtain position matching result information; determining each unmatched target object information based on the position matching result information and the obtained target object information group; According to the object information set, feature matching is performed on each unmatched target object information to obtain feature matching result information; Update the status information of each target object information corresponding to the obtained position matching result information and feature matching result information; Determining target object information that meets a preset error-prone condition and target object information that meets a preset stationary condition based on the obtained position matching result information and feature matching result information; Object tracking information corresponding to the obtained target object information group is generated according to the target object information that meets the preset error-prone condition, the target object information that meets the preset stationary condition, and the forward target image that meets the preset forward condition.
2. The method according to claim 1, wherein The method further comprises: In response to determining that a change target image that meets a preset change condition is detected, detecting the change target image to obtain a change target object information group; The obtained changed target object information group is used as the target object information group, and the matching step is performed again.
3. The method according to claim 1, wherein The method further comprises: In response to determining that no forward target image meeting the preset forward condition is detected, each target object information in the obtained target object information group is added as object information to the object information set to obtain a changed object information set as the object information set.
4. The method according to claim 1, wherein The performing position matching on the forward target object information group and the obtained target object information group to obtain position matching result information includes: Inputting the target image into a pre-trained object center position information generation model to obtain each first position information corresponding to the obtained target object information group; Inputting the forward target image corresponding to the forward target object information group into the object center position information generation model to obtain each second position information corresponding to the forward target object information group; For each target object information in the obtained target object information group, perform the following position matching steps: Determining first position information corresponding to the target object information; Determining, based on the first location information, whether there is second location information that meets a preset location condition among the obtained second location information; In response to determining that there is second location information that meets the preset location condition among the obtained second location information, determining the second location information that meets the preset location condition as target location information; determining the forward target object information corresponding to the target position information in the obtained forward target object information group as the target forward target object information; Inputting the determined target forward target object information and the target object information into a pre-trained object information matching model to obtain matching information corresponding to the target object information and the determined target forward target object information; In response to determining that the obtained matching information satisfies a preset matching information condition, determining the first preset matching result information as the first matching result information corresponding to the target object information; Adding the obtained first matching result information to the first matching result information set to update the first matching result information set; Deleting the determined target forward target object information and the target object information from the obtained forward target object information group and the target object information group, respectively, to obtain updated forward target object information group and target object information group; Determining the second preset matching result information as the second matching result information corresponding to each forward target object information in the updated forward target object information group, to obtain each second matching result information; Adding each obtained second matching result information to the second matching result information set to update the second matching result information set; Determining the third preset matching result information as the third matching result information of each target object information in the updated target object information to obtain each third matching result information; Adding each obtained third matching result information to the third matching result information set to update the third matching result information set; The updated first matching result information set, second matching result information set, and third matching result information set are determined as position matching result information.
5. The method according to claim 4, wherein The determining of each unmatched target object information according to the position matching result information and the obtained target object information group includes: In response to determining that the third matching result information set and the second matching result information set in the position matching result information are not empty, the target object information corresponding to the third matching result information set and the forward target object information corresponding to the second matching result information set in the obtained target object information group are determined as unmatched target object information, and each unmatched target object information is obtained.
6. The method according to claim 1, wherein The feature matching is performed on each unmatched target object information according to the object information set to obtain feature matching result information, including: Performing feature extraction processing on each object information in the obtained object information set to obtain feature information of each object corresponding to the obtained object information set; Performing feature extraction processing on each target object information in each unmatched target object information to obtain feature information of each target object; For each target object feature information obtained, perform the following steps: Determining whether there is object feature information that meets a preset feature condition in the obtained feature information of each object; In response to determining that object feature information that satisfies the preset feature condition exists in the obtained object feature information, determining fourth preset matching result information as fourth matching result information corresponding to the target object feature information; adding the fourth matching result information corresponding to the target object feature information to the fourth matching result information set, to obtain a changed fourth matching result information set as the fourth matching result information set; Deleting the target object information corresponding to the target object feature information from each unmatched target object information to update each unmatched target object information; Deleting the object feature information that meets the preset feature condition from the obtained feature information of each object, so as to update the obtained feature information of each object; In response to determining that no object feature information that satisfies the preset feature condition exists in the obtained object feature information, determining the fifth preset matching result information as the fifth matching result information corresponding to the target object feature information; adding the fifth matching result information corresponding to the target object feature information to the fifth matching result information set, to obtain a changed fifth matching result information set as the fifth matching result information set; The obtained fifth matching result information set and fourth matching result information set are determined as feature matching result information.
7. The method according to claim 1, wherein Generating object tracking information corresponding to the obtained target object information group according to the target object information satisfying the preset error-prone condition, the target object information satisfying the preset stationary condition, and the forward target image satisfying the preset forward condition includes: Obtaining forward object tracking information corresponding to the forward target object image; adding each target object information that does not satisfy the preset error-prone condition and the preset stationary condition to the forward object tracking information of the corresponding forward target object information group to obtain changed forward object tracking information; The obtained changed forward object tracking information is determined as the object tracking information corresponding to the obtained target object information group.
8. A multi-mode fusion target tracking device based on situational awareness, comprising: a detection unit, configured to perform object detection on a target image and obtain a target object information group; The execution unit is configured to, in response to determining that a forward target image meeting a preset forward condition is detected, perform the following matching steps based on the obtained target object information group and the forward target object information group corresponding to the detected forward target image: positionally match the forward target object information group and the obtained target object information group to obtain position matching result information; and determine each unmatched target object information based on the position matching result information and the obtained target object information group; According to the object information set, feature matching is performed on each unmatched target object information to obtain feature matching result information; Updating the status information of each target object information corresponding to the obtained position matching result information and feature matching result information; determining the target object information that meets the preset error-prone condition and the target object information that meets the preset static condition based on the obtained position matching result information and feature matching result information; The generation unit is configured to generate object tracking information corresponding to the obtained target object information group based on the target object information that meets the preset error-prone condition, the target object information that meets the preset stationary condition, and the forward target image that meets the preset forward condition.
9. An electronic device comprising: one or more processors; a storage device having one or more programs stored thereon, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 7.
10. A computer-readable medium having a computer program stored thereon, wherein: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Multi-target tracking method with fusion mechanism
CN113920161A
Multi-target tracking method and device
CN114219827A