An object trajectory recognition method, device, electronic device and storage medium

By performing lens segmentation and information fusion characteristics on the video, the problem of low accuracy of object re-recognition in complex scenes of small videos is solved, and efficient object trajectory recognition is achieved.

CN113591527BActive Publication Date: 2025-07-08TENCENT TECHNOLOGY (SHENZHEN) CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202110049271.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-01-14
Publication Date
2025-07-08
Estimated Expiration
2041-01-14

AI Technical Summary

Technical Problem

The prior art is difficult to achieve high accuracy of object re-identification in complex scenarios, especially in complex scenarios such as small videos, where the problem of object motion discontinuous caused by lens switching leads to interruption of tracking.

Method used

By segmenting the video lens, dividing it into multiple video clips, and object detection and tracking trajectory connection are performed in each clip, combining local and global semantic information fusion features, and finally object re-identification is performed in different video clips to form a complete motion trajectory.

Benefits of technology

It effectively improves the accuracy of object re-identification, avoids the problem of discontinuous object movement caused by lens switching, and realizes efficient object trajectory recognition in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113591527B_ABST
    Figure CN113591527B_ABST
Patent Text Reader

Abstract

This application relates to the field of computer technology, especially the field of artificial intelligence technology, and provides an object trajectory recognition method, device, electronic device and storage medium to improve the accuracy of object re-identification in videos. Among them, the method includes: performing shot segmentation on the video to be recognized to obtain multiple video segments, each video segment corresponding to a shot; performing object detection on each obtained video segment to respectively determine the detection frames of each object detected in each video segment; connecting the detection frames of the same object in different video frames of the same video segment to respectively obtain the tracking trajectories of each object in each video segment; for each object, connecting the tracking trajectories of the same object in different video segments to obtain the movement trajectories of each object in the video to be recognized. This application combines the characteristics of the video and divides the object trajectory recognition process into three parts: object detection, trajectory tracking, and re-identification, improving the accuracy of object re-identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to the field of artificial intelligence technology, and provides an object trajectory recognition method, device, electronic device and storage medium. Background Art

[0002] With the progress of society and technology, object recognition has increasingly become a reliable technology. For example, face recognition has been widely used as a security measure. For the object re-identification task, most of the related technologies focus on one or several specific categories, such as people, cars, etc., for re-identification, and most of the scenarios are surveillance cameras on campus or on the road, which are relatively simple. Based on these relatively simple methods, it is difficult to achieve better detection effects in complex scenarios such as small videos. Therefore, how to improve the accuracy of object re-identification in videos and achieve better detection effects in complex scenarios such as small videos is an urgent problem to be solved. Summary of the Invention

[0003] Embodiments of the present application provide an object trajectory recognition method, device, electronic device and storage medium to improve the accuracy of object re-identification in videos.

[0004] An object trajectory recognition method provided by an embodiment of the present application includes:

[0005] Performing shot segmentation on the video to be recognized to obtain a plurality of video segments, where each video segment corresponds to a shot, and each video segment includes at least one video frame;

[0006] Performing object detection on each of the obtained video segments, and respectively determining detection frames of each object detected in each of the video segments;

[0007] For each of the video segments, connecting the detection frames of the same object in different video frames within the same video segment to respectively obtain the tracking trajectories of each object in each of the video segments;

[0008] For each of the objects, connecting the tracking trajectories of the same object in different video segments to obtain the movement trajectories of each object in the video to be recognized.

[0009] An object trajectory recognition device provided by an embodiment of the present application includes:

[0010] A shot segmentation unit, configured to perform shot segmentation on the video to be recognized to obtain a plurality of video segments, where each video segment corresponds to a shot, and each video segment includes at least one video frame;

[0011] An object detection unit for performing object detection on each obtained video clip and respectively determining detection frames of each object detected in each video clip;

[0012] A trajectory tracking unit for respectively connecting the detection frames of the same object in different video frames within the same video clip for each video clip, and respectively obtaining the tracking trajectories of each object in each video clip;

[0013] A re-identification unit for connecting the tracking trajectories of the same object in different video clips for each object, and obtaining the movement trajectories of each object in the video to be identified.

[0014] Optionally, the object detection unit is specifically configured to:

[0015] For each video frame in each video clip, respectively obtain the detection frames of each object by performing the following operations:

[0016] Obtain the local semantic information, local position information, and global semantic information corresponding to one video frame in each video frame;

[0017] Fuse the local semantic information, local position information, and global semantic information corresponding to the one video frame and the video frames before the one video frame respectively to obtain corresponding local semantic fusion features, local position fusion features, and global semantic fusion features;

[0018] Based on the local semantic fusion features, the local position fusion features, and the global semantic fusion features, perform object detection on the one video frame, and determine the detection frames of each object detected in the one video frame.

[0019] Optionally, the object detection unit is specifically configured to:

[0020] In the video clip to which the one video frame belongs, obtain the first cross-correlation degree between the feature map of the region containing the target object in the video frames before the one video frame and the feature map corresponding to the one video frame;

[0021] Based on the first cross-correlation degree, determine the first region where the object corresponding to the target object is located in the one video frame, and obtain the corresponding local semantic fusion features by enhancing the features of the first region;

[0022] In the video segment to which the one video frame belongs, obtain the position information of the detection boxes in the video frames before the one video frame, determine the corresponding second region in the feature map corresponding to the one video frame based on the position information, and obtain the corresponding local position fusion feature by enhancing the features of the second region;

[0023] In the video segment to which the one video frame belongs, obtain the second cross-correlation degree between the average features of at least one object in the video frames before the one video frame and the feature map corresponding to the one video frame;

[0024] Based on the second cross-correlation degree, obtain the third region where the object corresponding to the at least one object is located in the one video frame, and obtain the corresponding global semantic fusion feature by enhancing the features of the third region.

[0025] Optionally, the trajectory tracking unit is specifically configured to:

[0026] Respectively obtain the object detection results within each detection box in each video frame within each video segment;

[0027] For each video frame within the same video segment, frame by frame obtain the feature similarity between the features of each object detection result corresponding to each video frame and the features corresponding to each existing tracking trajectory, where the existing tracking trajectory is the tracking trajectory of the object obtained based on the video frames before the current video frame;

[0028] For each of the object detection results, respectively perform the following operations:

[0029] For one object detection result among the object detection results, if the feature similarity between the one object detection result and the existing tracking trajectory with the highest corresponding feature similarity is greater than a first preset threshold, then splice the one object detection result with the existing tracking trajectory with the highest feature similarity to obtain the tracking trajectory of the object corresponding to the existing tracking trajectory with the highest feature similarity within the same video segment;

[0030] If the feature similarity between the one object detection result and the feature similarities of all the existing tracking trajectories is lower than the first preset threshold, then regard the object corresponding to the one object detection result as a new object, and create a new tracking trajectory with the current frame as the initial frame.

[0031] Optionally, the trajectory tracking unit is specifically configured to:

[0032] For each of the object detection results and each of the existing tracking trajectories, respectively perform the following operations:

[0033] For an object detection result among the respective object detection results and an existing tracking trajectory among the respective existing tracking trajectories, based on the features of the one object detection result and the features of the one existing tracking trajectory, obtain a corresponding response map, where the response map is used to characterize the similarity between the pixels in the feature map corresponding to the one object detection result and the pixels in the feature map corresponding to the one existing tracking trajectory;

[0034] Take the mean of the similarities corresponding to a set number of regions with the highest amplitudes in the response map as the feature similarity between the one object detection result and the one existing tracking trajectory.

[0035] Optionally, the re-identification unit is specifically configured to:

[0036] Initialize each tracking trajectory in the first video segment as a query set, and use each tracking trajectory in the remaining video segments as a retrieval library;

[0037] Traverse each tracking trajectory in the retrieval library one by one in chronological order, and respectively obtain the first trajectory similarity between each tracking trajectory in the retrieval library and each tracking trajectory in the query set;

[0038] Match each tracking trajectory in each video segment based on the first trajectory similarity, and connect the matching tracking trajectories belonging to the same object to obtain the motion trajectories of the respective objects.

[0039] Optionally, the re-identification unit is specifically configured to:

[0040] If the object is a person, based on the face recognition method and the person re-identification method, re-identify the people in different video segments to obtain the first object similarity of the people in different video segments; determine the same object in different video segments according to the first object similarity, and connect the tracking trajectories of the same object in different video segments to obtain the motion trajectories of the respective objects in the video to be recognized;

[0041] If the object is a non-person, based on the person re-identification method, re-identify the non-people in different video segments to obtain the second object similarity of the non-people in different video segments, determine the same object in different video segments according to the second object similarity, and connect the tracking trajectories of the same object in different video segments to obtain the motion trajectories of the respective objects in the video to be recognized;

[0042] Among them, the first object similarity is determined based on face similarity and re-identification similarity, the second object similarity is determined based on re-identification similarity, the face similarity is the similarity between the face features of objects in different video segments obtained through face recognition, and the re-identification similarity is the similarity between the pedestrian re-identification features of objects in different video segments obtained through pedestrian re-identification.

[0043] Optionally, the device further includes:

[0044] An update unit, configured to perform the following operations respectively for each tracking trajectory in the retrieval library:

[0045] For one tracking trajectory among the tracking trajectories in the retrieval library, respectively obtain the second trajectory similarity between the one tracking trajectory and each tracking trajectory in the query set;

[0046] Based on the obtained second trajectory similarity, find the tracking trajectory in the query set that best matches the one tracking trajectory, and determine the trajectory identifier of the best-matching tracking trajectory;

[0047] If the difference between the maximum similarity between each first tracking trajectory in the query set and the one tracking trajectory and the average similarity between each second tracking trajectory in the query set and the one tracking trajectory is greater than a second preset threshold, add the one tracking trajectory to the query set, where the first tracking trajectory is the tracking trajectory in the query set with the same identifier as the trajectory identifier, the second tracking trajectory is the tracking trajectory in the query set with a different identifier from the trajectory identifier, and the trajectory identifier of the one tracking trajectory in the query set is the same as the trajectory identifier.

[0048] Optionally, the object detection unit is specifically configured to:

[0049] Input the respective video segments into a trained re-identification model, and based on the object detection part in the re-identification model, perform object detection on the respective video segments to obtain detection frames of the respective objects detected in the respective video segments;

[0050] The trajectory tracking unit is specifically configured to:

[0051] Input the detected detection frames and the corresponding video segments into the object tracking part in the re-identification model, and based on the object tracking part, respectively connect the detection frames of the same object in different video frames within the same video segment to obtain the respective tracking trajectories of the respective objects in the respective video segments;

[0052] The re-identification unit is specifically configured to:

[0053] Based on the object re-identification part in the re-identification model, connect the tracking trajectories of the same object in different video segments to obtain the respective movement trajectories of each object in the video to be identified;

[0054] Among them, the re-identification model is obtained by training based on a training sample data set. The training samples in the training sample data set include multiple sample video segments obtained by segmenting a sample video. The sample video segment is a video containing at least one sample object, and the sample video segment contains at least one video frame.

[0055] Optionally, the device further includes:

[0056] A model training unit, configured to obtain a re-identification model by the following method:

[0057] Perform iterative training on the re-identification model according to the training samples in the training sample data set, and output the trained re-identification model when the training is completed; wherein, the following operations are performed during one iterative training process:

[0058] Select a training sample from the training sample data set;

[0059] Input each sample video segment in the training sample into the re-identification model respectively, and perform object detection on each sample video segment based on the object detection part in the re-identification model to obtain the detection frames of each sample object detected in each sample video segment;

[0060] Input the detected detection frames and the corresponding sample video segments into the object tracking part in the re-identification model, and connect the detection frames of the same sample object in different video frames within the same sample video segment based on the object tracking part respectively to obtain the respective tracking trajectories of each sample object in each sample video segment;

[0061] Based on the sample object re-identification part in the re-identification model, connect the tracking trajectories of the same sample object in different video segments to obtain the respective movement trajectories of each sample object in the sample video;

[0062] Construct a loss function based on the respective movement trajectories of each sample object, and adjust the parameters of the re-identification model based on the loss function.

[0063] Optionally, the device further includes:

[0064] A video screening unit, configured to screen each object in the video to be identified to obtain at least one candidate object;

[0065] Based on the motion trajectories of the at least one candidate object, other videos containing the at least one candidate object are obtained from a preset video set.

[0066] An electronic device provided by an embodiment of the present application includes a processor and a memory. Among them, the memory stores program codes. When the program codes are executed by the processor, the processor is caused to execute the steps of the above object trajectory recognition method.

[0067] An embodiment of the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps of any of the above object trajectory recognition methods.

[0068] An embodiment of the present application provides a computer-readable storage medium, which includes program codes. When the program codes run on an electronic device, the program codes are used to cause the electronic device to execute the steps of the above object trajectory recognition method.

[0069] The beneficial effects of the present application are as follows:

[0070] An embodiment of the present application provides an object trajectory recognition method, apparatus, electronic device, and storage medium. Since in the embodiment of the present application, in combination with the characteristics of the video, the video is first segmented into multiple video segments, and each video segment corresponds to a lens. After tracking the trajectories in each segmented lens, the tracking trajectories corresponding to the objects in each lens are matched and spliced, and then the complete motion trajectories of each object in the video to be recognized can be obtained, avoiding the problem that the object motion is discontinuous caused by the original lens switching in the video, and further avoiding the problem that the tracking will be interrupted. The embodiment of the present application divides the object trajectory recognition process into three parts: object detection, trajectory tracking, and re-identification, effectively improving the accuracy of object re-identification.

[0071] Other features and advantages of the present application will be described in the following description, and, in part, will be obvious from the description, or will be understood by implementing the present application. The objectives and other advantages of the present application can be realized and obtained by the structures specifically pointed out in the written description, claims, and drawings. Description of the Drawings

[0072] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation of the present application. In the drawings:

[0073] Figure 1 An optional schematic diagram of an application scenario in an embodiment of the present application;

[0074] Figure 2 A flowchart of a method for identifying an object trajectory in an embodiment of the present application;

[0075] Figure 3 A flowchart of a method for searching for similar objects in a short video in an embodiment of the present application;

[0076] Figure 4 A schematic diagram of the search results of similar objects in a short video in an embodiment of the present application;

[0077] Figure 5 A flowchart of a method for determining a detection frame in an embodiment of the present application;

[0078] Figure 6 A schematic diagram of the structure of a temporal information fusion network in an embodiment of the present application;

[0079] Figure 7 A schematic diagram of a local semantic information fusion process in an embodiment of the present application;

[0080] Figure 8 A schematic diagram of a local position information fusion process in an embodiment of the present application;

[0081] Figure 9 A schematic diagram of a global semantic information fusion process in an embodiment of the present application;

[0082] Figure 10 An example diagram of a point-to-point similarity calculation method in an embodiment of the present application;

[0083] Figure 11 A flowchart of a training method for a re-identification model in an embodiment of the present application;

[0084] Figure 12 A schematic diagram of the composition structure of an object trajectory recognition device in an embodiment of the present application;

[0085] Figure 13 A schematic diagram of the composition structure of the first electronic device in an embodiment of the present application;

[0086] Figure 14 A schematic diagram of the composition structure of the second electronic device in an embodiment of the present application. Detailed implementation manners

[0087] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the following will clearly and completely describe the technical solutions of this application in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are part of the technical solutions of this application, rather than all of them. Based on the embodiments described in this application document, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope protected by the technical solutions of this application.

[0088] The following introduces some concepts involved in the embodiments of this application.

[0089] Person re-identification (Re-ID): Also known as pedestrian re-identification, it is a technology that uses computer vision technology to determine whether a specific pedestrian exists in an image or video sequence. It is widely regarded as a sub-problem of image retrieval. Given a surveillance pedestrian image, retrieve the pedestrian image across devices. It aims to make up for the visual limitations of fixed cameras in related technologies and can be combined with pedestrian detection / pedestrian tracking technologies, and can be widely applied in fields such as intelligent video surveillance and intelligent security.

[0090] Short video: Refers to a short video, that is, a short film video, which is a way of spreading Internet content. Generally, it is a video with a duration of less than 5 minutes spread on Internet new media; with the popularization of mobile terminals and the acceleration of the network, short, flat, and fast large-flow dissemination content has gradually gained the favor of major platforms, fans, and capital.

[0091] Semantic information: It is one of the forms of information expression, referring to the information with certain meaning that can eliminate the uncertainty of things. For the information recipient, information can be expressed at three levels: syntactic information, semantic information, and pragmatic information. Semantic information can be understood and interpreted with the help of natural language. Only the information of human society contains semantic information. All scientific information belongs to semantic information. Due to differences in personal knowledge level and cognitive ability, the understanding of semantic information often has a strong subjective color. Different people obtain significantly different semantic information and pragmatic information from the same syntactic information. In the embodiments of this application, semantic information is divided into local semantic information and global semantic information, both of which are for the feature map of video frame images. Among them, the time-domain fusion ranges corresponding to local and global are different. Specifically, the fusion range of the local is the previous tau frames of the current frame, and the fusion range of the global is all frames before the current frame.

[0092] Fusion features: refer to the features obtained by fusing local semantic information, local position information, and global semantic information in this article, including local semantic fusion features, global semantic fusion features, and local position fusion features. Among them, the local semantic fusion feature refers to the fusion feature obtained by fusing the local semantic information corresponding to the current video frame and the tau video frames before this video frame within the same shot. The local position fusion feature refers to the fusion feature obtained by fusing the local position information corresponding to the current video frame and the tau video frames before this video frame within the same shot. The global semantic fusion feature refers to the fusion feature obtained by fusing the global semantic information corresponding to the current video frame and all the video frames before this video frame within the same shot.

[0093] The embodiments of this application relate to artificial intelligence (AI) and machine learning technologies, and are designed based on computer vision technology and machine learning (ML) in artificial intelligence.

[0094] Artificial intelligence uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, and is a theory, method, technology, and application system that can perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence.

[0095] Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling machines to have the functions of perception, reasoning, and decision-making. Artificial intelligence technologies mainly include several major directions such as computer vision technology, natural language processing technology, and machine learning / deep learning. With the research and progress of artificial intelligence technologies, artificial intelligence has been studied and applied in many fields. For example, common ones include smart homes, intelligent customer service, virtual assistants, smart speakers, intelligent marketing, driverless, autonomous driving, robots, intelligent healthcare, etc. It is believed that with the development of technology, artificial intelligence will be applied in more fields and play an increasingly important role.

[0096] Machine learning is a multi-disciplinary cross-discipline that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Compared with data mining, which looks for mutual characteristics among big data, machine learning pays more attention to the design of algorithms, enabling computers to automatically "learn" rules from data and use the rules to predict unknown data.

[0097] Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence. Its applications cover all fields of artificial intelligence. Machine learning and deep learning generally include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, and inductive learning. Reinforcement learning (RL), also known as re-inforcement learning, evaluation learning, or enhancement learning, is one of the paradigms and methodologies of machine learning, used to describe and solve the problem that an agent maximizes rewards or achieves specific goals through learning strategies during the interaction with the environment.

[0098] In the embodiments of this application, when identifying the trajectory of an object in a video, a re-identification model of machine learning is adopted. The method for training the re-identification model proposed in the embodiments of this application can be divided into two parts, including a training part and an application part; among them, the training part involves the technical field of machine learning. In the training part, the re-identification model is trained through machine learning, so that the sample video containing at least one sample object given in the embodiments of this application is used as a training sample to train the re-identification model. After the training sample passes through the re-identification model, the output result of the re-identification model is obtained. Combining the output result, the model parameters are continuously adjusted through an optimization algorithm; the application part is used to use the re-identification model trained in the training part to perform object detection, trajectory tracking, and re-identification on each object in the video to be identified, and finally obtain the movement trajectories of each object in the video to be identified. In addition, it should be noted that the re-identification model in the embodiments of this application can be trained online or offline, and no specific limitation is made here. In the embodiments of this application, offline training is taken as an example for illustration.

[0099] The following briefly introduces the design concept of the embodiments of this application:

[0100] The object re-identification problem is to use computer vision technology to determine whether a specific target exists in an image or a video sequence. Specifically, when tracking a specific target using a video, since the video source comes from a fixed position, when the target leaves the field of view, cross-video relay tracking is required. At this time, the problem of detecting the specific target in other video sources belongs to the object re-identification problem.

[0101] However, for the object re-identification task, most of the practices in related technologies focus on one or several specific categories, such as people, cars, etc., for re-identification, and most of the scenes are surveillance cameras on campus or on the road, which are relatively simple. Based on these relatively simple methods, it is difficult to achieve better detection effects in complex scenes such as short videos.

[0102] In view of this, embodiments of the present application propose an object trajectory recognition method, device, electronic device, and storage medium. Since embodiments of the present application provide an object trajectory recognition method, device, electronic device, and storage medium. In embodiments of the present application, in combination with the characteristics of the video, the video is first segmented into multiple video segments, each video segment corresponding to a lens. After tracking the trajectories in each segmented lens, the tracking trajectories corresponding to the objects in each lens are then matched and spliced, so as to obtain the complete motion trajectories of each object in the video to be recognized, avoiding the problem that the discontinuous object motion caused by the original lens switching in the video leads to the interruption of tracking. Embodiments of the present application divide the object trajectory recognition process into three parts: object detection, trajectory tracking, and re-identification, effectively improving the accuracy of object re-identification.

[0103] The preferred embodiments of the present application are described below with reference to the accompanying drawings of the specification. It should be understood that the preferred embodiments described herein are only for the purpose of illustrating and explaining the present application, and are not used to limit the present application. And without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.

[0104] As Figure 1 shown, it is a schematic diagram of the application scenario of the embodiments of the present application. It is a schematic diagram of the application scenario of the embodiments of the present application. The application scenario diagram includes two terminal devices 110 and a server 120. The terminal device 110 and the server 120 can communicate through a communication network. The user can browse videos through the terminal device 110, and video-related applications can be installed on the terminal device 110, such as video software, short video software, etc. The applications involved in the embodiments of the present application can be software, or web pages, applets, etc. The background server is the background server corresponding to the software or web page, applet, etc., and the specific type of the client is not limited.

[0105] In an alternative embodiment, the communication network is a wired network or a wireless network. The terminal device 110 and the server 120 can be directly or indirectly connected through wired or wireless communication methods, and the present application does not limit this here.

[0106] In the embodiments of the present application, the terminal device 110 is an electronic device used by a user. The electronic device may be a computer device such as a personal computer, mobile phone, tablet computer, notebook, e-reader, smart home, etc., which has a certain computing ability and runs instant messaging software and websites or social software and websites. Each terminal device 110 is connected to the server 120 through a wireless network. The server 120 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It may also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.

[0107] Among them, the re-identification model can be deployed on the server 120 for training. A large number of training samples can be stored in the server 120 for training the re-identification model. Optionally, after the re-identification model is trained based on the training method in the embodiments of the present application, the trained re-identification model can be directly deployed on the server 120 or the terminal device 110. Generally, the re-identification model is directly deployed on the server 120. In the embodiments of the present application, the re-identification model is mainly used to re-identify each object in the video to be identified and obtain the movement trajectories of each object. According to this model, the movement trajectories and semantic features of key objects in the short video can be automatically extracted, and based on the movement trajectories and semantic features, other videos containing the same objects can also be found in multiple videos, etc.

[0108] In a possible application scenario, the training samples in the present application can be stored using cloud storage technology. Cloud storage is a new concept extended and developed from the concept of cloud computing. A distributed cloud storage system (hereinafter referred to as the storage system) refers to a storage system that combines a large number of different types of storage devices (storage devices are also called storage nodes) in the network through cluster applications, grid technology, and distributed storage file systems, and works together through application software or application interfaces to jointly provide data storage and service access functions to the outside world.

[0109] In a possible application scenario, to facilitate reducing communication latency, servers 120 can be deployed in various regions, or for load balancing, different servers 120 can serve the corresponding regions of each terminal device 10 respectively. Multiple servers 120 can share data through a blockchain, and multiple servers 120 are equivalent to a data sharing system composed of multiple servers 120. For example, the terminal device 110 is located at location a and is communicatively connected to server 120, and the terminal device 110 is located at location b and is communicatively connected to other servers 120.

[0110] For each server 120 in the data sharing system, it has a node identifier corresponding to that server 120. Each server 120 in the data sharing system can store the node identifiers of other servers 120 in the data sharing system, so that subsequently, according to the node identifiers of other servers 120, the generated blocks can be broadcast to other servers 120 in the data sharing system. Each server 120 can maintain a node identifier list as shown in the following table, and store the server 120 name and the node identifier correspondingly in this node identifier list. Among them, the node identifier can be an Internet Protocol (IP) address for interconnection between networks and any other information that can be used to identify the node. Only the IP address is taken as an example in Table 1 for illustration.

[0111] Table 1

[0112] Server Name Node Identifier Node 1 119.115.151.174 Node 2 118.116.189.145 … … Node N 119.124.789.258

[0113] Next, in combination with the application scenario described above, the object trajectory recognition method provided by the exemplary embodiments of the present application will be described with reference to the accompanying drawings. It should be noted that the above application scenario is only shown for the convenience of understanding the spirit and principle of the present application, and the embodiments of the present application are not limited in this regard.

[0114] Refer to Figure 2 As shown, it is a flowchart of the implementation of an object trajectory recognition method provided by an embodiment of the present application. The specific implementation process of this method is as follows:

[0115] S21: Perform shot segmentation on the video to be recognized to obtain multiple video segments, where each video segment corresponds to a shot, and each video segment contains at least one video frame;

[0116] In the embodiment of the present application, in order to address the problem that the original shot transitions in the video cause discontinuous object movement, which in turn leads to the interruption of tracking, in the embodiment of the present application, shot segmentation is performed on the video to be recognized, and the video to be recognized is cut according to shots and divided into multiple video segments.

[0117] Among them, when performing shot segmentation on the video to be recognized, some common shot segmentation methods are as follows:

[0118] The histogram method determines whether a shot transition has occurred by statistically analyzing whether the change amount of the histogram of the color distribution of each pixel in the front and back frames exceeds a threshold; the pixel method determines based on the difference in the sum of the brightness, grayscale, or colorfulness of all pixels in the front and back frames; or some methods based on deep network training and learning to make judgments, etc., which are not specifically limited here.

[0119] Taking the video to be recognized as a short video as an example, for a short video with a length of T input First, perform shot segmentation to obtain K sequences of video segments representing K shots, where represents the k-th shot starting from time t * where the parameter t represents time.

[0120] S22: Perform object detection on each obtained video segment, and respectively determine the detection frames of each object detected in each video segment;

[0121] In the embodiments of the present application, through object detection, objects included in each video frame of each video segment in the short video are recognized, and each object is located and classified.

[0122] In an alternative embodiment, step S22 can be performed based on a machine learning model. Specifically: input each video segment into a trained re-identification model, and based on the object detection part in the re-identification model, perform object detection on each video segment respectively to obtain the detection frames of each object detected in each video segment.

[0123] For example, as listed above, input each video segment into the object detection part in the re-identification model respectively, and perform object detection on each video segment based on this object detection part to obtain the detection frames of N objects in each video frame of each video segment

[0124] S23: For each video segment respectively, connect the detection frames of the same object in different video frames within the same video segment to respectively obtain the tracking trajectories of each object in each video segment;

[0125] In an alternative embodiment, step S23 can also be performed based on a machine learning model. Specifically, the detected detection boxes and the corresponding video segments are input into the object tracking part of the re-identification model. Based on the object tracking part, the detection boxes of the same object in different video frames within the same video segment are connected respectively to obtain the tracking trajectories of each object in each video segment.

[0126] Based on the above process, the detection boxes of the same object detected in different video frames at different times within different shots obtained by the object detection part can be connected to obtain the tracking trajectories of each object in each video segment. The tracking trajectory refers to the tracking trajectory obtained by tracking the trajectories of each object within each shot.

[0127] Specifically, the detection boxes detected in step S22 and the corresponding video segments are input into the object tracking part of the re-identification model together. Based on the object tracking part, the detection boxes of the same object detected in different video frames at different times within different shots obtained by the object detection part are connected to obtain the tracking trajectories of each object within each shot, denoted as

[0128] S24: For each object, connect the tracking trajectories of the same object in different video segments to obtain the motion trajectories of each object in the video to be recognized.

[0129] In an alternative embodiment, step S24 can also be performed based on a machine learning model. Specifically, based on the object re-identification part of the re-identification model, connect the tracking trajectories of the same object in different video segments to obtain the motion trajectories of each object in the video to be recognized.

[0130] In the object re-identification part, it is mainly used to connect the tracking trajectories of the same object in different shots into a complete motion trajectory. Therefore, when identifying the same object in different shots, the specific workflow is to initialize the tracking trajectories in the first shot as the query set, and the remaining tracking trajectories as the gallery, and then traverse the tracking trajectories in the gallery one by one in chronological order, calculate the feature similarity with the tracking trajectories in the query set respectively for matching and connection, and finally obtain the complete motion trajectories of each object.

[0131] It should be noted that the cross-shot general object re-identification part in the embodiments of this application needs to be divided into two branches to process the cross-shot re-identification tasks of people and non-human objects respectively. Therefore, the trajectories can be divided into human trajectories and non-human object trajectories, and the complete motion trajectories of people and non-human objects in the video can be obtained respectively. And output.

[0132] In an alternative embodiment, each object in the video to be recognized can also be screened to obtain at least one candidate object; based on the movement trajectories of the at least one candidate object, other videos containing the at least one candidate object can be obtained from a preset video set. Additionally, other videos containing the same candidate objects can also be searched from the preset video set based on the movement trajectories and the extracted features of the at least one candidate object, realizing the search for similar objects in small videos.

[0133] Refer to Figure 3 As shown, it is a flowchart of a method for searching for similar objects in small videos in an embodiment of the present application. For the small video to be recognized, first, the small video is segmented into multiple video clips, denoted as Fq, and object detection is performed on each video clip to obtain each detection box, denoted as B 1:t , and then, object tracking is performed based on the detection box and the video clip to obtain each tracking trajectory, denoted as traj 1:n , where 1, 2, and 3 are divided according to the category of the object and correspond to 3 objects respectively. The trajectories are divided into person trajectories (PersonTracklets) and non-person object trajectories (Non-Person Tracklets). As shown in, object 2 is a person, and objects 1 and 3 are non-person objects.

[0134] In the embodiment of the present application, after merging the tracking trajectories of each object in different shots, the complete movement trajectories (Tracklets) of each object can be obtained. Next, major object selection can be further performed to screen out the major objects, and then object recognition detection (Indentical ObjectRetrieval) can be performed according to the movement trajectories traji∈major(N) of these major objects (the features of these major objects can also be further referred to), and other small video sets {Fg} containing these objects can be screened out.

[0135] In the above embodiment, aiming at the particularity of small videos, a system combining three modules of video object detection, video object tracking, and video object re-identification is proposed, so as to find the major objects in small videos, extract the object trajectories and features for the complex scenes of small videos, and use them to find the consistent objects in other videos, effectively realizing the search for similar objects in small videos.

[0136] Refer to Figure 4As shown, it is a schematic diagram of the search results of similar objects in a short video in an embodiment of the present application. Among them, based on the short video on the left side of the arrow, other short videos containing the same person as this short video can be searched, Figure 4 and three are listed, namely the three short videos on the right side of the arrow. In addition, more short videos with the same or similar objects can also be searched, but not all are listed in Figure 4 this figure.

[0137] Based on the above method, short videos containing similar objects or the same objects as the target video can be effectively identified from a pile of short videos, realizing the search for similar objects in short videos.

[0138] In an optional implementation manner, when performing object detection on each obtained video segment and respectively determining the detection frames of each object detected in each video segment, specifically, separate detection is performed on each video segment, and each video frame in the same video segment is detected respectively to obtain the detection frames of each object in each video segment. Specifically, for each video segment among each video segment, object detection needs to be performed on each video frame in this video segment, and the following operations are respectively used to obtain the detection frames of each object in each video frame. Refer to Figure 5 As shown, it is a schematic flowchart of a method for determining a detection frame in an embodiment of the present application:

[0139] S51: Obtain the local semantic information, local position information, and global semantic information corresponding to one video frame among each video frame;

[0140] In the embodiment of the present application, the global semantic information and the local semantic information mainly refer to the semantic information of the object, and the local position information mainly refers to the position information of the detection frame.

[0141] S52: Respectively fuse the local semantic information, local position information, and global semantic information corresponding to one video frame and the video frames before this video frame to obtain the corresponding local semantic fusion feature, local position fusion feature, and global semantic fusion feature;

[0142] In the embodiments of the present application, local and global refer to local and global in the time domain. Among them, the fusion range of the local is the previous tau frames of the current frame, and the fusion range of the global is all frames before the current frame. A video frame in this step may refer to the current frame video frame in the frame-by-frame detection process. For the current frame, this step specifically fuses the local semantic information corresponding to the current frame and several previous video frames (such as tau frames) of the current frame to obtain the corresponding local semantic fusion feature; fuses the local position information corresponding to the current frame and several previous video frames (such as tau frames) of the current frame to obtain the corresponding local position fusion feature; fuses the global semantic information corresponding to the current frame and all previous video frames of the current frame to obtain the corresponding global semantic fusion feature.

[0143] S53: Based on the local semantic fusion feature, local position fusion feature, and global semantic fusion feature, perform object detection on a video frame to determine the detection boxes of each object detected in a video frame.

[0144] Specifically, this step refers to performing object detection on the current frame video frame based on the local semantic fusion feature, local position fusion feature, and global semantic fusion feature determined in step S52 to obtain the detection boxes of each object detected in this frame video frame.

[0145] In the above embodiment, by continuously repeating the above steps to perform object detection on each video frame in each video segment, the detection boxes of each object can be obtained.

[0146] Refer to Figure 6 As shown, it is a schematic structural diagram of a Temporal Information Fusion Network (TIFN) in the embodiments of the present application. Among them, the video frames at times from t * -τ to t * are input into the TIFN. Assuming the current frame is to After corresponding processing through residual modules, spatial pyramid modules, temporal information fusion modules, etc., and in addition, there are also downsampling operations, upsampling operations, splicing operations, etc. as shown in Figure 6 Finally, the corresponding local semantic fusion feature, local position fusion feature, and global semantic fusion feature can be obtained.

[0147] ​In the embodiments of the present application, by using TIFN proposed for the small video scenario for information fusion, it mainly uses local semantic information, local position information, and global semantic information to aggregate and enhance the features of the current frame in the small video, separates semantic similarity and position continuity, and introduces global features, enhancing the robustness of video object detection, better adapting to the complex scene characteristics of small videos, and enabling the video object detection algorithm to still maintain good detection effects in complex scenarios such as small videos.

[0148] Specifically, three different sub-modules are used in TIFN in the embodiments of the present application to fuse local semantic information, local position information, and global semantic information respectively. Among them Figure 6 "×3" therein indicates that the fusion is performed through three different sub-modules respectively. The following combines Figures 7 to 9 these several drawings to introduce in detail the process of information fusion listed in the embodiments of the present application:

[0149] Refer to Figure 7 As shown, it is a schematic diagram of a local semantic information fusion process listed in the embodiments of the present application. For local semantic information, specifically: in the video segment to which a video frame belongs, obtain the feature map of the region containing the target object in the video frames before this video frame, and the cross-correlation degree between it and the feature map corresponding to this video frame, that is, the first cross-correlation degree. Based on the first cross-correlation degree, determine the first region where the object corresponding to the target object is located in this video frame, and obtain the corresponding local semantic fusion feature by enhancing the features of the first region.

[0150] Among them, in this sub-module, the feature of the object region detected in the previous few frames (referring to the feature map of the region containing the target object) is mainly used to calculate the first cross-correlation degree with the current frame feature map, that is Figure 7 As shown, can represent the feature map of frame t * -τ, can represent the feature map of frame t * frame, and calculate the first cross-correlation degree through the cross-correlation (xcorr) operation, represents the detection box in frame t * -τ. In actual code implementation, in the present application, the feature map of the region containing the target object is used as the convolution kernel, that is, the in the figure, perform a convolution operation on the current frame feature map to obtain a similarity response map Attn_ls representing the similarity between each region of the current frame and the target object, so as to find the region where the corresponding object is located in the current frame and enhance the features of this region, and obtain the local semantic fusion feature corresponding to the current frame video frame. Finally, the detection box in

[0151] Refer to Figure 8 As shown, it is a schematic diagram of a local position information fusion process in an embodiment of the present application. For local position information, specifically: in a video segment to which a video frame belongs, obtain the position information of the detection boxes in the video frames before the video frame, determine the corresponding second region in the feature map corresponding to the video frame based on the position information, and obtain the corresponding local position fusion feature by enhancing the features of the second region.

[0152] Among them, in this sub-module, mainly use the position information of the object detection boxes detected in the previous few frames, and in the corresponding region of the feature map of the current frame, by creating a two-dimensional Hanning window, that is, in the figure Obtain an attention map Attn_ll with attention concentrated on the region where the target object is located in the previous few frames, so as to enhance the features of this region and obtain the corresponding local position fusion feature. Finally, the detection box in can be determined based on this fusion feature

[0153] Refer to Figure 9 As shown, it is a schematic diagram of a global semantic information fusion process in an embodiment of the present application. For global semantic information, specifically: in a video segment to which a video frame belongs, obtain the second cross-correlation degree between the average features of at least one object in the video frames before a video frame and the feature map corresponding to the video frame; based on the second cross-correlation degree, obtain the third region where the object corresponding to at least one object is located in a video frame, and obtain the corresponding global semantic fusion feature by enhancing the features of the third region.

[0154] Among them, in this sub-module, mainly use a global pool to store the average features {crop n} G of the main objects that have appeared in all the frames before the current frame, and use these features to calculate the cross-correlation degree with the feature map of the current frame, that is, the second cross-correlation degree. In the same way as the local semantic information sub-module, obtain a similarity response map Attn_gs, so as to find the region where the corresponding object is located in the current frame and enhance the features of this region, and obtain the corresponding global semantic fusion feature. Finally, the detection box in can be determined based on this fusion feature

[0155] In the above embodiment, after respectively determining the detection boxes of each object detected in each video segment, object tracking and re-identification can be performed.

[0156] In an alternative embodiment, when performing step S23, it specifically includes the following processes:

[0157] First, obtain the object detection results within each detection box in each video frame of each video segment. Further, for each video frame within the same video segment, frame by frame, obtain the feature similarity between the features of each object detection result corresponding to each video frame and the features of each existing tracking trajectory, where the existing tracking trajectory is the tracking trajectory of an object obtained based on the video frames before the current video frame.

[0158] In the embodiments of the present application, for each object detection result obtained for the same video segment, the following operations are respectively performed:

[0159] For an object detection result among each object detection result, if the feature similarity between the object detection result and the existing tracking trajectory with the highest corresponding feature similarity is greater than the first preset threshold, then splice the object detection result and the existing tracking trajectory with the highest feature similarity, that is, assign the object detection result to the existing tracking trajectory with the highest feature similarity to it, and then splice the result and the existing tracking trajectory. Finally, use the spliced tracking trajectory as the tracking trajectory of the object corresponding to the existing tracking trajectory with the highest feature similarity of the object detection result with the highest feature similarity within the video segment. If the feature similarity between the object detection result and the feature similarities of all existing tracking trajectories is very low, all lower than the first preset threshold, it indicates that the object is a newly emerged object in the current video frame. Take the object corresponding to the object detection result as a new object, and create a new tracking trajectory with the current frame as the initial frame.

[0160] In the embodiments of the present application, a ResNet50 network with the last fully connected layer removed is used as the backbone network for the object tracking part to extract the features of the object within the detection box, and frame by frame, calculate the feature similarity between the features of each object detection result in the current frame and the features of each existing tracking trajectory, which can be represented in the form of a similarity matrix. Finally, based on the magnitude of the feature similarity, assign each object detection result, splice it with the corresponding existing tracking trajectory, or create a new tracking trajectory to achieve the effective detection of the tracking trajectories of objects within the same shot.

[0161] In the embodiments of the present application, different from the conventional tracking algorithms, in order to obtain a more fine-grained feature similarity and achieve the mutual matching between the local regions of the detected objects, the present application adopts a point-to-point similarity calculation method. For the convenience of calculation, the present application uses the cosine distance to measure the feature similarity of the local regions. Of course, other distances can also be used to measure the feature similarity, such as the Euclidean distance, which is not specifically limited herein.

[0162] In an alternative embodiment, when calculating the feature similarity between an object detection result among various object detection results and an existing tracking trajectory among various existing tracking trajectories, based on the feature of the object detection result and the feature of the existing tracking trajectory, a corresponding response map is obtained, where the response map is used to represent the similarity between the pixels in the feature map corresponding to the object detection result and the pixels in the feature map corresponding to the existing tracking trajectory; the mean of the similarities corresponding to a set number of regions with the highest amplitudes in the response map is used as the feature similarity between the object detection result and the existing tracking trajectory.

[0163] For example, when the set number is k, the calculation process is shown in Formulas 1 and 2:

[0164]

[0165] where, f det , f tr respectively represent the feature of the object detection result in the current frame and the feature of the existing tracking trajectory, φ(f det , f tr ) is the response map of the two features point by point, top k (·) represents the k regions with the highest amplitudes in the response map, and the final similarity result is obtained by taking the mean of the k regions with the highest amplitudes in the response map.

[0166] Refer to Figure 10 shown, which is an example diagram of a point-to-point similarity calculation method in an embodiment of the present application. Among them, F d represents the feature map of an object detection result in the current frame, P 11 , P 12 respectively represent two feature points in the feature map; F t represents the feature map of an existing tracking trajectory, P 21 , P 22 respectively represent two feature points in the feature map. As Figure 10 shown, Rensponse Map represents the corresponding response map, where the final feature similarity result is obtained by taking the mean (mean) of the k regions with the highest amplitudes in the response map, that is, similarity in the figure.

[0167] In the above embodiment, after obtaining the tracking trajectories of each object in each shot through trajectory tracking, the tracking trajectories of the same object in different shots can be spliced to obtain the complete motion trajectories of each object in the video to be recognized.

[0168] In an alternative embodiment, when connecting the tracking trajectories of the same object in different video segments, it is necessary to match the same object in different shots. In the embodiments of the present application, this is mainly achieved based on trajectory similarity. Specifically, first, initialize each tracking trajectory in the first video segment as a query set, and use each tracking trajectory in the remaining video segments as a retrieval library. Then, traverse each tracking trajectory in the retrieval library one by one in chronological order, and respectively obtain the first trajectory similarity between each tracking trajectory in the retrieval library and each tracking trajectory in the query set. Finally, based on the first trajectory similarity, match each tracking trajectory in each video segment, and connect the matched tracking trajectories belonging to the same object to obtain the complete motion trajectory of each object.

[0169] If the first trajectory similarity between a tracking trajectory B1 in the retrieval library and a tracking trajectory A1 in the query set is relatively high, for example, greater than a certain similarity threshold, it can be considered that these two tracking trajectories are the tracking trajectories of the same object, that is, A1 and B1 are matched. If there are other tracking trajectories B2 and B3 in the retrieval library in addition to B1, and their first trajectory similarities with A1 are also relatively high, and these two tracking trajectories belong to the same object, then B1, B2, B3, and A1 can be connected to obtain the complete motion trajectory of the object corresponding to A1. Among them, when connecting B1, B2, B3, and A1, they can be spliced according to the time sequence between the corresponding shots.

[0170] In an alternative embodiment, when matching each tracking trajectory in each video segment based on the first trajectory similarity and connecting the matched tracking trajectories belonging to the same object to obtain the motion trajectory of each object, the objects can be divided into human and non-human objects, and the two cases will be discussed separately. The following is a detailed introduction to these two cases:

[0171] Case 1: The object is a human.

[0172] In this case, based on the face recognition method and the person re-identification method, re-identify the people in different video segments to obtain the first object similarity of the people in different video segments. Determine the same object in different video segments according to the first object similarity, and connect the tracking trajectories of the same object in different video segments to obtain the motion trajectory of each object in the video to be recognized.

[0173] Among them, the first object similarity is determined based on the face similarity and the re-identification similarity. The face similarity is the similarity between the face features of the objects in different video segments obtained through face recognition, and the re-identification similarity is the similarity between the person re-identification features of the objects in different video segments obtained through person re-identification.

[0174] That is, in the embodiments of the present application, for the person re-identification branch, both the face recognition algorithm and the ReID algorithm are used to jointly re-identify the people in different shots. Since conventional pedestrian re-identification algorithms often rely on the clothing of people for re-identification, and there will be cases where the clothing of people changes in many short videos. In addition, because the application scenario of the pedestrian re-identification algorithm is the video obtained by surveillance cameras, the people in it are often far away from the camera, and it is difficult to obtain facial details. Only features such as clothing, decoration, and human posture with more obvious appearance information can be relied on. On the contrary, in the short video scenario, the clothing, decoration, and posture of people are more complex and have a larger variation range. Since the face of the person is closer to the camera, it has instead become the most distinguishable and consistent feature of the person in different shots. Therefore, in addition to the conventional pedestrian re-identification algorithm, the person re-identification branch in the embodiments of the present application additionally introduces the face recognition algorithm. In the case where the face cannot be detected, the pedestrian re-identification feature is used as the judgment benchmark. In the embodiments of the present application, the Part-based Convolutional Baseline (PCB) model based on local information is directly cited as the pedestrian re-identification feature extractor, and the Additive Angular Margin Loss for Deep Face Recognition (ArcFace) algorithm (referred to as the angular face recognition algorithm for short) model is cited as the face feature extractor. The similarity of people is jointly determined by the cosine similarity of the features extracted by the above two models, as shown in Formula 3:

[0175]

[0176] Among them, λ1 and λ2 are the similarities calculated by the pedestrian re-identification algorithm and the face recognition algorithm respectively. Generally speaking, in the short video scenario, λ2 > λ1. Simi person That is, the first object similarity in the embodiments of the present application.

[0177] Case 2: The object is non-human.

[0178] In this case, based on the pedestrian re-identification method, non-humans in different video segments are re-identified to obtain the second object similarity of non-humans in different video segments. Based on the second object similarity, the same object in different video segments is determined, and the tracking trajectories of the same object in different video segments are connected to obtain the motion trajectories of each object in the video to be recognized. Among them, the second object similarity is determined based on the re-identification similarity.

[0179] For the re-identification branch of non-human objects, as described above, objects of the same variety or brand that are difficult to distinguish by the naked eye can be regarded as the same instance, unless they appear in the same frame of the video at the same time. Therefore, the present application can regard the re-identification task of non-human objects as a fine-grained classification task. In the embodiment of the present application, the first 4 residual modules of the Residual Network 50 (ResNet50) pre-trained by the ImageNet (a computer vision system recognition project name) dataset and followed by an average pooling layer are used as the feature extractor for the non-human object re-identification branch. Since the ImageNet dataset contains 1,000 categories and many fine-grained categories, the feature extractor can extract non-human object features with good discrimination for different varieties and brands. Similarly, this branch uses the cosine distance to calculate the similarity between different non-human object features, that is, the second object similarity.

[0180] In addition, considering the conventional pedestrian re-identification task, since the pedestrian re-identification scenario often has a fixed camera, a relatively stable background, and the appearance information of the same person under different cameras is also relatively consistent, it is feasible to maintain only a fixed query set. Unless a new object appears, the query set will always be the tracking trajectories of each object in the first shot. Even if the tracking trajectory in the subsequent retrieval library matches the trajectory in the query set, the query set will not be modified or updated. However, in a complex small video scenario, due to the complex changes in object appearance information, it is difficult for a fixed query set to be satisfied and it is prone to false matching. Therefore, the present application introduces a query set self-update mechanism in both the human and non-human object re-identification branches.

[0181] In an alternative embodiment, the query set is updated by the following method. For each tracking trajectory in the retrieval library, the following operations are performed respectively:

[0182] First, for one of the tracking trajectories in each tracking trajectory in the retrieval library, the second trajectory similarity between one tracking trajectory and each tracking trajectory in the query set is obtained respectively; this process is similar to the above process of connecting the tracking trajectories of the same object in different shots. First, the object tracking trajectories in the first shot are still initialized as the query set, and the remaining tracking trajectories are initialized as the retrieval library, and the tracking trajectories in the retrieval library are traversed one by one in chronological order to calculate the second trajectory similarity between each tracking trajectory in the query set and each tracking trajectory in the detection library. The calculation process of the second trajectory similarity here is the same as that of the first trajectory similarity. The difference is that after calculating the second trajectory similarity between the tracking trajectories in the retrieval library and each tracking trajectory in the query set, first find the tracking trajectory in the query set that best matches it, and then determine the matching trajectory ID, that is, based on the obtained second trajectory similarity, find the tracking trajectory in the query set that best matches one tracking trajectory, and determine the trajectory identification IDX of the best-matching tracking trajectory.

[0183] If the difference between the maximum similarity between each first tracking trajectory in the query set and one tracking trajectory and the average similarity between each second tracking trajectory in the query set and one tracking trajectory is greater than the second preset threshold, then add one tracking trajectory to the query set, where the first tracking trajectory is the tracking trajectory in the query set with the same identifier as the trajectory identifier IDX, the second tracking trajectory is the tracking trajectory in the query set with a different identifier from the trajectory identifier IDX, and the trajectory identifier of one tracking trajectory in the query set is the same as the trajectory identifier.

[0184] For example, there are four tracking trajectories in the query library, namely A1, A2, A3, and A4, where the trajectory identifiers (ID) of the first two tracking trajectories are the same, indicating that these two tracking trajectories may be of the same object. If the maximum similarity max(simi sameID ) between the retrieval library tracking trajectory B1 and all the first tracking trajectories in the query set with the same trajectory ID as IDX, and the average similarity avg(simi diffID ) between the second tracking trajectories in the query set with different trajectory IDs from IDX is greater than the second preset threshold Δ, then it is considered that the retrieval library tracking trajectory B1 is highly likely to match the corresponding query set tracking trajectory, and this retrieval library tracking trajectory is also added to the query set, and its trajectory ID is the same as the matching trajectory ID, that is, the trajectory identifier is all IDX. Therefore, the updated query set contains a total of five tracking trajectories, namely A1, A1, A2, A3, and B1.

[0185] For example, the trajectory identifiers of the tracking trajectories A1 and A2 in the query set are both IDX. Among them, the similarity between the retrieval library tracking trajectory B1 and the first type of tracking trajectory A1 in the query set is s1, and the similarity with the first type of tracking trajectory A2 in the query set is s2, where s1 > s2, then max(simi sameID ) = s1. The similarity between the retrieval library tracking trajectory B1 and the second type of tracking trajectory A3 in the query set is s3, and the similarity with the second type of tracking trajectory A4 in the query set is s4, where the average value avg(simi diffID ) = (s3 + s4) / 2. If the difference between s1 and (s3 + s4) / 2 is greater than Δ, then update the query set, and the trajectory identifiers of A1, A2, and B1 in the updated query set are all IDX.

[0186] Through the above implementation manners, it can be ensured that the matching between the retrieval set and the query set trajectories does not only depend on the trajectories within the first shot, but also calculates the matching similarity with the remaining highly reliable trajectories, reducing the false matching caused by the complex appearance changes in the short video.

[0187] After introducing the above embodiments, the training process of the re-identification model will be described in detail below.

[0188] In the embodiments of the present application, the re-identification model is obtained by training based on a training sample data set. The training samples in the training sample data set include multiple sample video segments obtained by segmenting the sample video. The sample video segment is a video containing at least one sample object, and the sample video segment contains at least one video frame.

[0189] Optionally, the re-identification model is obtained by the following method: performing iterative training on the re-identification model according to the training samples in the training sample data set, and outputting the trained re-identification model when the training is completed; wherein, the following operations are performed during one iterative training process, as Figure 11 shown.

[0190] Referring to Figure 11 shown, which is a flowchart of a training method of a re-identification model in the embodiments of the present application. One iterative training process specifically includes the steps:

[0191] S111: Select a training sample from the training sample data set;

[0192] S112: Input each sample video segment in the training sample into the re-identification model respectively, and perform object detection on each sample video segment based on the object detection part in the re-identification model to obtain the detection frames of each sample object detected in each sample video segment;

[0193] S113: Input the detected detection boxes and the corresponding sample video segments into the object tracking part of the re-identification model. Based on the object tracking part, connect the detection boxes of the same sample object in different video frames within the same sample video segment to obtain the respective tracking trajectories of each sample object in each sample video segment;

[0194] S114: Based on the sample object re-identification part in the re-identification model, connect the tracking trajectories of the same sample object in different video segments to obtain the respective movement trajectories of each sample object in the sample video;

[0195] S115: Construct a loss function based on the respective movement trajectories of each sample object, and adjust the parameters of the re-identification model based on the loss function.

[0196] In step S115, the specific process of constructing the loss function includes the construction of the loss function for the object detection part, the construction of the loss function for the object tracking part, and the construction of the loss function for the object re-identification part.

[0197] The object detection part used in the embodiments of the present application mainly adopts the YOLOv3-SPP (You Only Live Once version 3) model. Therefore, the loss function used in the training stage consists of two parts: classification error and regression error, which are respectively for the classification sub-task and the regression sub-task in the object detection framework.

[0198] Among them, the classification error includes two parts, namely the classification error loss of the foreground and background obj and the classification error loss of the specific fine classification cls , in the embodiments of the present application, the foreground refers to objects (including people and non-people), the background is non-objects (such as buildings, roads, etc.), and the specific fine classification refers to the division of object categories, such as people, cats, dogs, etc.

[0199] The classification errors listed in this article are calculated using cross-entropy, and the formulas are shown in Formulas 4 and 5 respectively:

[0200]

[0201] Among them, S 2 represents the feature Figure 1 There are a total of S×S units (also called pixels), B represents B candidate boxes generated by each unit, and Indicates whether the j-th candidate box of the i-th unit matches a certain ground truth box, that is, the Intersection-over-Union (IOU) between the two is greater than the specified threshold. If they match, then Otherwise and are the foreground confidence and class confidences predicted by the model. is the true foreground confidence, which is determined by If then Otherwise is the true class confidence.

[0202] In the embodiment of the present application, the regression error also includes two parts, namely the center coordinate error and the width-height coordinate error. In the model, the offset of the center coordinate and width-height predicted by the regression sub-network used to generate the detection box is combined with the center coordinate and width-height of the current position corresponding anchor box, and through inverse operation, the center coordinate and width-height of the detection box actually predicted by the model are obtained, and then the squared errors are calculated respectively with the center coordinate and width-height of the actual annotation box. The calculation formula is as shown in Formula 6:

[0203]

[0204] Among them, represents the center coordinate of the detection box actually predicted by the model. represents the width and height of the detection box. represents the center coordinate of the actual annotation box. represents the width and height of the actual annotation box.

[0205] Based on the classification error and regression error calculated above, the overall loss function of the object detection part can be calculated, as shown in Formula 7:

[0206] Loss = λ reg × loss reg + loss obj + loss cls Formula 7

[0207] Among them, λ reg is an artificially defined coefficient for balancing the regression error and the classification error.

[0208] In addition, since the appearance features are mainly trained in the embodiment of the present application, the loss function of the object tracking part is the same as that of the object re-identification part, and the Triplet Loss is used, as shown in Formula 8:

[0209] Loss = max(d(a, p) - d(a, n) + margin, 0) Equation 8

[0210] Among them, Triplet Loss is a loss function in deep learning, used to train samples with small differences, such as faces, etc. The Feed data includes anchor examples, positive examples, and negative examples. Here, a in Equation 8 is the Anchor, p is the Positive, and n is the Negative. By optimizing the distance between the anchor example and the positive example to be less than the distance between the anchor example and the negative example, the similarity calculation of the samples is achieved. d(a, p) represents the distance between the anchor example and the positive example, and d(a, n) represents the distance between the anchor example and the negative example.

[0211] In the embodiments of this application, setting a reasonable margin value is crucial, which is an important indicator for measuring similarity. Briefly speaking, the smaller the margin value is set, the easier it is for the loss to approach 0, but it is difficult to distinguish similar images. The larger the margin value is set, the more difficult it is for the loss value to approach 0, and it may even cause the network not to converge, but it can be more confident in distinguishing relatively similar images. The margin value can be specifically set according to experience.

[0212] The experimental results of the embodiments of this application are briefly introduced below:

[0213] In the embodiments of this application, the optimizer used in the object detection part is the Stochastic Gradient Descent (SGD) algorithm, with the momentum coefficient set to 0.9, and the rest of the training settings follow those in YOLOv3-SPP. The object tracking part uses the Adaptive Moment Estimation (Adam) optimizer, and the parameter settings are simple. Basically, the default parameters of Adam are followed, and then the learning rate is decreased at 30, 60, 90, and 120. The object re-identification part follows the training settings in the PCB algorithm.

[0214] In the embodiments of the present application, the temporal information fusion network proposed in the object detection part of the present application can achieve nearly a 5% improvement in the mean Average Precision (mAP) compared to the baseline model YOLOv3 on the Short Video Dataset for Identical Object Retrieval (SVD-IOR). Moreover, compared with other video detection algorithms on the commonly used evaluation dataset ImageNet VID for video object detection, it has comparable performance and better speed.

[0215] Table 2 Ablation experiments with the addition of local semantic information fusion, local position information fusion, and global semantic information fusion

[0216]

[0217] Table 3 Comparative experiments between TIFN and other video object detection algorithms on the ImageNet VID dataset

[0218]

[0219]

[0220] The overall results of the present application are also better than those of other algorithms in complex scenarios. Especially in the small video dataset SVD-IOR and the VidOR dataset which is very similar to the small video scenario, significantly better results are obtained. An example of the experimental result diagram for general object re-identification among multiple videos is as follows Figure 4 shown.

[0221] Table 4 Comparative experiments between the system proposed in the present application and other algorithms on different datasets

[0222]

[0223] In summary, the present application utilizes a system including object detection, object tracking, and object re-identification to extract the object trajectory and its features for complex scenarios of small videos, and searches for other videos containing the same object through its features, which can effectively improve the accuracy of object trajectory recognition.

[0224] Based on the same inventive concept, the embodiments of the present application also provide an object trajectory recognition device. As Figure 12 shown, it is a schematic structural diagram of an object trajectory recognition device 1200 in the embodiments of the present application, which may include:

[0225] The shot segmentation unit 1201 is configured to perform shot segmentation on the video to be recognized, obtaining multiple video segments, where each video segment corresponds to one shot, and each video segment contains at least one video frame;

[0226] The object detection unit 1202 is configured to perform object detection on each of the obtained video segments, respectively determining the detection frames of each object detected in each video segment;

[0227] The trajectory tracking unit 1203 is configured to, respectively for each video segment, connect the detection frames of the same object in different video frames within the same video segment, respectively obtaining the respective tracking trajectories of each object in each video segment;

[0228] The re-identification unit 1204 is configured to, for each object, connect the tracking trajectories of the same object in different video segments, obtaining the respective motion trajectories of each object in the video to be recognized.

[0229] Optionally, the object detection unit 1202 is specifically configured to:

[0230] For each video frame in each video segment, respectively perform the following operations to obtain the detection frames of each object:

[0231] Obtain the local semantic information, local position information, and global semantic information corresponding to one video frame in each video frame;

[0232] Fuse the local semantic information, local position information, and global semantic information corresponding to one video frame and the video frames before it, respectively obtaining the corresponding local semantic fusion feature, local position fusion feature, and global semantic fusion feature;

[0233] Based on the local semantic fusion feature, local position fusion feature, and global semantic fusion feature, perform object detection on one video frame, determining the detection frames of each object detected in one video frame.

[0234] Optionally, the object detection unit 1202 is specifically configured to:

[0235] In the video segment to which one video frame belongs, obtain the first cross-correlation degree between the feature map of the region containing the target object in the video frame before one video frame and the feature map corresponding to one video frame;

[0236] Based on the first cross-correlation degree, determine the first region where the object corresponding to the target object is located in one video frame, and by enhancing the features of the first region, obtain the corresponding local semantic fusion feature;

[0237] In a video segment to which a video frame belongs, obtain the position information of the detection boxes in the video frame before the video frame, determine the corresponding second region in the feature map corresponding to the video frame based on the position information, and obtain the corresponding local position fusion feature by enhancing the features of the second region;

[0238] In a video segment to which a video frame belongs, obtain the second cross-correlation degree between the average features of at least one object in the video frame before the video frame and the feature map corresponding to the video frame;

[0239] Based on the second cross-correlation degree, obtain the third region where the object corresponding to at least one object in the video frame is located, and obtain the corresponding global semantic fusion feature by enhancing the features of the third region.

[0240] Optionally, the trajectory tracking unit 1203 is specifically configured to:

[0241] Respectively obtain the object detection results within each detection box in each video frame within each video segment; for each video frame within the same video segment, frame by frame obtain the feature similarity between the features of each object detection result corresponding to the video frame and the features corresponding to each existing tracking trajectory, where the existing tracking trajectory is the tracking trajectory of the object obtained based on the video frames before the current video frame;

[0242] For each object detection result, respectively perform the following operations:

[0243] For an object detection result among each object detection result, if the feature similarity between an object detection result and the existing tracking trajectory with the highest corresponding feature similarity is greater than the first preset threshold, then splice the object detection result and the existing tracking trajectory with the highest feature similarity to obtain the tracking trajectory of the object corresponding to the existing tracking trajectory with the highest feature similarity within the same video segment;

[0244] If the feature similarity between an object detection result and the feature similarities of all existing tracking trajectories is lower than the first preset threshold, then use the object corresponding to the object detection result as a new object and create a new tracking trajectory with the current frame as the initial frame.

[0245] Optionally, the trajectory tracking unit 1203 is specifically configured to:

[0246] For each object detection result and each existing tracking trajectory, respectively perform the following operations:

[0247] For an object detection result among the various object detection results and an existing tracking trajectory among the various existing tracking trajectories, based on the features of an object detection result and the features of an existing tracking trajectory, obtain a corresponding response map, where the response map is used to characterize the similarity between the pixels in the feature map corresponding to an object detection result and the pixels in the feature map corresponding to an existing tracking trajectory;

[0248] Take the mean of the similarities corresponding to the regions with the highest magnitudes among a set number of regions in the response map as the feature similarity between an object detection result and an existing tracking trajectory.

[0249] Optionally, the re-identification unit 1204 is specifically configured to:

[0250] Initialize each tracking trajectory within the first video segment as a query set, and use each tracking trajectory within the remaining video segments as a retrieval library;

[0251] Traverse each tracking trajectory in the retrieval library one by one in chronological order, and respectively obtain the first trajectory similarities between each tracking trajectory in the retrieval library and each tracking trajectory in the query set;

[0252] Match each tracking trajectory within each video segment based on the first trajectory similarities, and connect the tracking trajectories that are matched to the same object to obtain the movement trajectories of each object.

[0253] Optionally, the re-identification unit 1204 is specifically configured to:

[0254] If the object is a person, based on the face recognition method and the person re-identification method, re-identify the people in different video segments to obtain the first object similarities of the people in different video segments; determine the same object in different video segments according to the first object similarities, and connect the tracking trajectories of the same object in different video segments to obtain the movement trajectories of each object in the video to be recognized;

[0255] If the object is a non-person, based on the person re-identification method, re-identify the non-people in different video segments to obtain the second object similarities of the non-people in different video segments, determine the same object in different video segments according to the second object similarities, and connect the tracking trajectories of the same object in different video segments to obtain the movement trajectories of each object in the video to be recognized;

[0256] Among them, the first object similarity is determined based on face similarity and re-identification similarity, the second object similarity is determined based on re-identification similarity, the face similarity is the similarity between the face features of objects in different video segments obtained through face recognition, and the re-identification similarity is the similarity between the pedestrian re-identification features of objects in different video segments obtained through pedestrian re-identification.

[0257] Optionally, the device further includes:

[0258] An update unit 1205, configured to perform the following operations respectively for each tracking trajectory in the retrieval library:

[0259] For one tracking trajectory among each tracking trajectory in the retrieval library, respectively obtain the second trajectory similarity between one tracking trajectory and each tracking trajectory in the query set;

[0260] Based on the obtained second trajectory similarity, find the tracking trajectory in the query set that best matches one tracking trajectory, and determine the trajectory identifier of the best-matching tracking trajectory;

[0261] If the difference between the maximum similarity between each first tracking trajectory in the query set and one tracking trajectory and the average similarity between each second tracking trajectory in the query set and one tracking trajectory is greater than the second preset threshold, then add one tracking trajectory to the query set, where the first tracking trajectory is the tracking trajectory in the query set with the same identifier as the trajectory identifier, the second tracking trajectory is the tracking trajectory in the query set with a different identifier from the trajectory identifier, and the trajectory identifier of one tracking trajectory in the query set is the same as the trajectory identifier.

[0262] Optionally, the object detection unit 1202 is specifically configured to:

[0263] Input each video segment into a trained re-identification model, and based on the object detection part in the re-identification model, perform object detection on each video segment respectively to obtain the detection frames of each object detected in each video segment;

[0264] The trajectory tracking unit 1203 is specifically configured to:

[0265] Input the detected detection frames and the corresponding video segments into the object tracking part in the re-identification model, and based on the object tracking part, connect the detection frames of the same object in different video frames within the same video segment respectively to obtain the respective tracking trajectories of each object in each video segment;

[0266] The re-identification unit 1204 is specifically configured to:

[0267] Based on the object re-identification part in the re-identification model, connect the tracking trajectories of the same object in different video segments to obtain the respective movement trajectories of each object in the video to be recognized;

[0268] Among them, the re-identification model is obtained by training based on a training sample data set. The training samples in the training sample data set include multiple sample video segments obtained by segmenting the sample video. The sample video segment is a video containing at least one sample object, and the sample video segment contains at least one video frame.

[0269] Optionally, the device further includes:

[0270] A model training unit 1206, configured to obtain the re-identification model through the following method:

[0271] Perform iterative training on the re-identification model according to the training samples in the training sample data set, and output the trained re-identification model when the training is completed; among them, the following operations are performed during one iterative training process:

[0272] Select a training sample from the training sample data set;

[0273] Input each sample video segment in the training sample into the re-identification model respectively. Based on the object detection part in the re-identification model, perform object detection on each sample video segment respectively to obtain the detection frames of each sample object detected in each sample video segment;

[0274] Input the detected detection frames and the corresponding sample video segments into the object tracking part in the re-identification model. Based on the object tracking part, connect the detection frames of the same sample object in different video frames within the same sample video segment respectively to obtain the respective tracking trajectories of each sample object in each sample video segment;

[0275] Based on the sample object re-identification part in the re-identification model, connect the tracking trajectories of the same sample object in different video segments to obtain the respective movement trajectories of each sample object in the sample video;

[0276] Construct a loss function based on the respective movement trajectories of each sample object, and adjust the parameters of the re-identification model based on the loss function.

[0277] Optionally, the device further includes:

[0278] A video screening unit, configured to screen each object in the video to be recognized to obtain at least one candidate object;

[0279] Based on the movement trajectories of at least one candidate object, obtain other videos containing at least one candidate object from a preset video set.

[0280] Based on the same inventive concept as the above method embodiments, an electronic device is also provided in the embodiments of the present application. This electronic device can be used for object trajectory recognition. In one embodiment, the electronic device can be a server, such as Figure 1 the server 120 shown. In this embodiment, the structure of the electronic device can be as Figure 13 shown, including a memory 1301, a communication module 1303, and one or more processors 1302.

[0281] The memory 1301 is used to store the computer program executed by the processor 1302. The memory 1301 mainly includes a program storage area and a data storage area. Among them, the program storage area can store the operating system and programs required to run the instant messaging function, etc.; the data storage area can store various instant messaging information and operation instruction sets, etc.

[0282] The memory 1301 can be a volatile memory, such as a random-access memory (RAM); the memory 1301 can also be a non-volatile memory, such as a read-only memory, a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD); or the memory 1301 is any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited to this. The memory 1301 can be a combination of the above memories.

[0283] The processor 1302 can include one or more central processing units (CPUs) or be a digital processing unit, etc. The processor 1302 is used to implement the above object trajectory recognition method when calling the computer program stored in the memory 1301.

[0284] The communication module 1303 is used to communicate with terminal devices and other servers.

[0285] In the embodiments of the present application, the specific connection medium between the above memory 1301, communication module 1303, and processor 1302 is not limited. In the embodiments of the present disclosure, Figure 13 it is shown that the memory 1301 and the processor 1302 are connected through a bus 1304. The bus 1304 is shown as a thick line in Figure 13 The connection manners between other components are only for illustrative purposes and are not to be considered limiting. The bus 1304 can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation,Figure 13 It is represented by only one thick line in the figure, but it does not mean that there is only one bus or one type of bus.

[0286] The computer storage medium is stored in the memory 1301, and computer-executable instructions are stored in the computer storage medium. The computer-executable instructions are used to implement the object trajectory recognition method of the embodiments of the present application. The processor 1302 is used to execute the above object trajectory recognition method, as Figure 2 shown.

[0287] In another embodiment, the electronic device may also be other electronic devices, such as Figure 1 the terminal device 110 shown. In this embodiment, the structure of the electronic device may be as Figure 14 shown, including: communication component 1410, memory 1420, display unit 1430, camera 1440, sensor 1450, audio circuit 1460, Bluetooth module 1470, processor 1480 and other components.

[0288] The communication component 1410 is used to communicate with the server. In some embodiments, it may include a Wireless Fidelity (WiFi) module. The WiFi module belongs to short-range wireless transmission technology, and the electronic device can help users send and receive information through the WiFi module.

[0289] The memory 1420 can be used to store software programs and data. The processor 1480 executes various functions and data processing of the terminal device 110 by running the software programs or data stored in the memory 1420. The memory 1420 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, flash memory device, or other volatile solid-state storage devices. The memory 1420 stores an operating system that enables the terminal device 110 to run. In the present application, the memory 1420 can store the operating system and various application programs, and can also store the code for executing the object trajectory recognition method of the embodiments of the present application.

[0290] The display unit 1430 can also be used to display information input by the user or information provided to the user, as well as the graphical user interface (GUI) of various menus of the terminal device 110. Specifically, the display unit 1430 may include a display screen 1432 disposed on the front of the terminal device 110. Among them, the display screen 1432 may be configured in the form of a liquid crystal display, a light-emitting diode, etc. The display unit 1430 can be used to display the sub-application playback interface in the embodiments of the present application.

[0291] The display unit 1430 can also be used to receive input numerical or character information and generate signal inputs related to the user settings and function control of the terminal device 110. Specifically, the display unit 1430 may include a touch screen 1431 disposed on the front of the terminal device 110, which can collect touch operations of the user thereon or nearby, such as clicking buttons, dragging scroll boxes, etc.

[0292] Among them, the touch screen 1431 can cover the display screen 1432, or the touch screen 1431 and the display screen 1432 can be integrated to implement the input and output functions of the terminal device 110. After integration, it can be simply called a touch display screen. In this application, the display unit 1430 can display application programs and corresponding operation steps.

[0293] The camera 1440 can be used to capture static images, and the user can send the images captured by the camera 1440 to the user of the chat partner through the client. The camera 1440 can be one or more. An object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then transmits the electrical signal to the processor 1480 to be converted into a digital image signal.

[0294] The terminal device may further include at least one sensor 1450, such as an acceleration sensor 1451, a distance sensor 1452, a fingerprint sensor 1453, and a temperature sensor 1454. The terminal device may also be configured with other sensors such as a gyroscope, a barometer, a hygrometer, a thermometer, an infrared sensor, a light sensor, and a motion sensor.

[0295] The audio circuit 1460, the speaker 1461, and the microphone 1462 can provide an audio interface between the user and the terminal device 110. The audio circuit 1460 can transmit the electrical signal converted from the received audio data to the speaker 1461, and the speaker 1461 converts it into a sound signal for output. The terminal device 110 may also be configured with volume buttons for adjusting the volume of the sound signal. On the other hand, the microphone 1462 converts the collected sound signal into an electrical signal, which is received by the audio circuit 1460 and converted into audio data, and then the audio data is output to the communication component 1410 to be sent to another terminal device 110, for example, or the audio data is output to the memory 1420 for further processing.

[0296] The Bluetooth module 1470 is used to interact with other Bluetooth devices with Bluetooth modules through the Bluetooth protocol. For example, the terminal device can establish a Bluetooth connection with a wearable electronic device (such as a smart watch) that also has a Bluetooth module through the Bluetooth module 1470, so as to conduct data interaction.

[0297] The processor 1480 is the control center of the terminal device. It uses various interfaces and circuits to connect all parts of the entire terminal. By running or executing software programs stored in the memory 1420 and calling data stored in the memory 1420, it executes various functions of the terminal device and processes data. In some embodiments, the processor 1480 may include one or more processing units; the processor 1480 may also integrate an application processor and a baseband processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the baseband processor mainly processes wireless communication. It can be understood that the above baseband processor may not be integrated into the processor 1480. In this application, the processor 1480 can run the operating system, application programs, user interface display and touch response, as well as the object trajectory recognition method of the embodiments of this application. In addition, the processor 1480 is coupled to the display unit 1430.

[0298] In some possible implementation manners, various aspects of the object trajectory recognition method provided in this application can also be implemented in the form of a program product, which includes program code. When the program code runs on a computer device, the program code is used to enable the computer device to execute the steps in the object trajectory recognition method according to various exemplary embodiments of this application described above in this specification. For example, the computer device can execute the steps as Figure 2 shown in.

[0299] The program product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0300] The program product of the embodiments of the present application may be a portable compact disc read-only memory (CD-ROM) and include program code, and can run on a computing device. However, the program product of the present application is not limited thereto. In the embodiments of the present application, the readable storage medium may be any tangible medium that contains or stores a program, and the program can be used by or in conjunction with a command execution system, device, or component.

[0301] The readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable signal medium may also be any readable medium other than the readable storage medium, and the readable medium can send, propagate, or transmit a program for use by or in conjunction with a command execution system, device, or component.

[0302] The program code contained on the readable medium can be transmitted by any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the above.

[0303] The program code for performing the operations of the present application can be written in any combination of one or more programming languages. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, executed as an independent software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, by using an Internet service provider to connect through the Internet).

[0304] It should be noted that although several units or subunits of the device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present application, the features and functions of the two or more units described above can be embodied in one unit. Conversely, the features and functions of one unit described above can be further divided and embodied by multiple units.

[0305] In addition, although the operations of the method of the present application are described in a specific order in the drawings, this does not require or imply that these operations must be performed in that specific order, or that all of the shown operations must be performed to achieve the desired result. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution.

[0306] Those skilled in the art should understand that the embodiments of the present application may be provided as a method, a system, or a computer program product. Therefore, the present application may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0307] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0308] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that realizes the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0309] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are performed on the computer or other programmable device to generate a computer-implemented process, thereby providing steps for realizing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks by the instructions executed on the computer or other programmable device.

[0310] Although the preferred embodiments of the present application have been described, additional changes and modifications can be made to these embodiments by those skilled in the art once they learn the basic creative concept. Therefore, the appended claims are intended to be interpreted to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present application.

[0311] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations.

Claims

1. An object trajectory recognition method, characterized in that, The method includes: Performing shot segmentation on the video to be recognized to obtain a plurality of video segments, where each video segment corresponds to a shot, and each video segment contains at least one video frame; For each video frame in each of the obtained video segments, the following operations are respectively performed: In the video segment to which a video frame belongs, obtaining the first cross-correlation degree between the feature map of the region containing the target object in the video frame before the one video frame and the feature map corresponding to the one video frame, and obtaining the position information of the detection box in the video frame before the one video frame; Based on the first cross-correlation degree, determining the first region where the object corresponding to the target object in the one video frame is located, and obtaining the corresponding local semantic fusion feature by enhancing the feature of the first region; After determining the corresponding second region in the feature map corresponding to the one video frame based on the position information, obtaining the corresponding local position fusion feature by enhancing the feature of the second region; And fusing the global semantic information corresponding to the one video frame and the video frame before the one video frame respectively to obtain the corresponding global semantic fusion feature; Based on the local semantic fusion feature, the local position fusion feature, and the global semantic fusion feature, performing object detection on the one video frame to determine the detection boxes of the respective objects detected in the one video frame; For each of the respective video segments, connecting the respective detection boxes of the same object in different video frames within the same video segment, and respectively obtaining the respective tracking trajectories of the respective objects in each of the respective video segments; For each of the respective objects, connecting the tracking trajectories of the same object in different video segments to obtain the respective movement trajectories of the respective objects in the video to be recognized.

2. The method according to claim 1, wherein The fusing the global semantic information corresponding to the one video frame and the video frame before the one video frame respectively to obtain the corresponding global semantic fusion feature includes: In the video segment to which the one video frame belongs, obtaining the second cross-correlation degree between the average feature of at least one object in the video frame before the one video frame and the feature map corresponding to the one video frame; Based on the second cross-correlation degree, obtaining the third region where the object corresponding to the at least one object in the one video frame is located, and obtaining the corresponding global semantic fusion feature by enhancing the feature of the third region.

3. The method according to claim 1, wherein The connecting the respective detection boxes of the same object in different video frames within the same video segment for each of the respective video segments to respectively obtain the respective tracking trajectories of the respective objects in each of the respective video segments includes: Respectively obtaining the object detection results within each of the respective detection boxes in each of the respective video frames within each of the respective video segments; For each of the respective video frames within the same video segment, frame by frame obtaining the feature similarity between the features of the respective object detection results corresponding to each of the respective video frames and the features corresponding to each of the respective existing tracking trajectories, where the existing tracking trajectories are the tracking trajectories of the objects obtained based on the video frames before the current video frame; For each of the object detection results, perform the following operations respectively: For one object detection result among the object detection results, if the feature similarity between the one object detection result and the existing tracking trajectory with the highest corresponding feature similarity is greater than the first preset threshold, splice the one object detection result and the existing tracking trajectory with the highest feature similarity to obtain the tracking trajectory of the object corresponding to the existing tracking trajectory with the highest feature similarity within the same video segment; If the feature similarity between the one object detection result and all existing tracking trajectories is lower than the first preset threshold, take the object corresponding to the one object detection result as a new object and create a new tracking trajectory with the current frame as the initial frame.

4. The method according to claim 3, wherein The step of obtaining the feature similarity between the features of each object detection result corresponding to each video frame and the features corresponding to each existing tracking trajectory frame by frame includes: For each of the object detection results and each of the existing tracking trajectories, perform the following operations respectively: For one object detection result among the object detection results and one existing tracking trajectory among the existing tracking trajectories, based on the feature of the one object detection result and the feature of the one existing tracking trajectory, obtain a corresponding response map, where the response map is used to represent the similarity between the pixels in the feature map corresponding to the one object detection result and the pixels in the feature map corresponding to the one existing tracking trajectory; Take the mean value of the similarities corresponding to the regions with the highest amplitudes of a set number in the response map as the feature similarity between the one object detection result and the one existing tracking trajectory.

5. The method according to claim 1, wherein The step of connecting the tracking trajectories of the same object in different video segments to obtain the motion trajectories of each object in the video to be recognized includes: Initialize the tracking trajectories in the first video segment as the query set and use the tracking trajectories in the remaining video segments as the retrieval library; Traverse each tracking trajectory in the retrieval library one by one in chronological order, and respectively obtain the first trajectory similarity between each tracking trajectory in the retrieval library and each tracking trajectory in the query set; Match the tracking trajectories in each video segment based on the first trajectory similarity, and connect the tracking trajectories that match and belong to the same object to obtain the motion trajectories of each object.

6. The method according to claim 5, wherein The step of matching the tracking trajectories in each video segment based on the first trajectory similarity and connecting the tracking trajectories that match and belong to the same object to obtain the motion trajectories of each object includes: If the object is a person, based on the face recognition method and the person re-identification method, re-identify the people in different video segments to obtain the first object similarity of the people in different video segments; determine the same object in different video segments according to the first object similarity, and connect the tracking trajectories of the same object in different video segments to obtain the motion trajectories of each object in the video to be recognized; If the object is a non - human, based on the person re - identification method, re - identify the non - humans in different video segments to obtain the second object similarity of the non - humans in different video segments. Determine the same object in different video segments according to the second object similarity, and connect the tracking trajectories of the same object in different video segments to obtain the motion trajectories of each object in the video to be identified; Among them, the first object similarity is determined based on the face similarity and the re - identification similarity. The second object similarity is determined based on the re - identification similarity. The face similarity is the similarity between the face features of the objects in different video segments obtained through face recognition. The re - identification similarity is the similarity between the person re - identification features of the objects in different video segments obtained through person re - identification.

7. The method according to claim 5, wherein The method further includes: For each tracking trajectory in the retrieval library, perform the following operations respectively: For one tracking trajectory among the tracking trajectories in the retrieval library, respectively obtain the second trajectory similarity between the one tracking trajectory and each tracking trajectory in the query set; Based on the obtained second trajectory similarity, find the tracking trajectory in the query set that best matches the one tracking trajectory, and determine the trajectory identifier of the best - matching tracking trajectory; If the difference between the maximum similarity between each first tracking trajectory in the query set and the one tracking trajectory and the average similarity between each second tracking trajectory in the query set and the one tracking trajectory is greater than a second preset threshold, add the one tracking trajectory to the query set, where the first tracking trajectory is the tracking trajectory in the query set with the same identifier as the trajectory identifier, the second tracking trajectory is the tracking trajectory in the query set with a different identifier from the trajectory identifier, and the trajectory identifier of the one tracking trajectory in the query set is the same as the trajectory identifier.

8. The method according to claim 1 or 2, characterized in that, The detection frames of each object in each video segment are obtained by inputting each video segment into a trained re - identification model and performing object detection on each video segment respectively based on the object detection part in the re - identification model; The step of respectively connecting the detection frames of the same object in different video frames within the same video segment for each video segment to obtain the respective tracking trajectories of each object in each video segment includes: Input the detected detection frame and the corresponding video segment into the object tracking part of the re - identification model, and respectively connect the detection frames of the same object in different video frames within the same video segment based on the object tracking part to obtain the respective tracking trajectories of each object in each video segment; The step of connecting the tracking trajectories of the same object in different video segments for each object to obtain the respective motion trajectories of each object in the video to be identified includes: Based on the object re-identification part in the re-identification model, connect the tracking trajectories of the same object in different video segments to obtain the respective motion trajectories of each object in the video to be identified; Among them, the re-identification model is obtained by training based on a training sample dataset. The training samples in the training sample dataset include multiple sample video segments obtained by segmenting a sample video. The sample video segment is a video containing at least one sample object, and the sample video segment contains at least one video frame.

9. The method according to claim 8, wherein The re-identification model is obtained by training in the following manner: According to the training samples in the training sample dataset, perform iterative training on the re-identification model, and when the training is completed, output the trained re-identification model; among them, the following operations are performed during one iterative training process: Select a training sample from the training sample dataset; Input each sample video segment in the training sample into the re-identification model respectively. Based on the object detection part in the re-identification model, perform object detection on each sample video segment respectively to obtain the detection frames of each sample object detected in each sample video segment; Input the detected detection frames and the corresponding sample video segments into the object tracking part in the re-identification model. Based on the object tracking part, connect the detection frames of the same sample object in different video frames within the same sample video segment respectively to obtain the respective tracking trajectories of each sample object in each sample video segment; Based on the sample object re-identification part in the re-identification model, connect the tracking trajectories of the same sample object in different video segments to obtain the respective motion trajectories of each sample object in the sample video; Construct a loss function based on the respective motion trajectories of each sample object, and adjust the parameters of the re-identification model based on the loss function.

10. The method according to any one of claims 1 to 7 and 9, characterized in that The method further includes: Screen each object in the video to be identified to obtain at least one candidate object; Based on the motion trajectories of the at least one candidate object, obtain other videos containing the at least one candidate object from a preset video set.

11. An object trajectory recognition device, characterized in that, It includes: A shot segmentation unit for segmenting the video to be identified to obtain multiple video segments, where each video segment corresponds to one shot, and each video segment contains at least one video frame; An object detection unit is configured to perform the following operations respectively for each video frame in each obtained video clip: in the video clip to which a video frame belongs, obtain the first cross-correlation degree between the feature map of the region containing the target object in the video frame before the video frame and the feature map corresponding to the video frame, and obtain the position information of the detection box in the video frame before the video frame; determine the first region where the object corresponding to the target object is located in the video frame based on the first cross-correlation degree, and obtain the corresponding local semantic fusion feature by enhancing the feature of the first region; after determining the corresponding second region in the feature map corresponding to the video frame based on the position information, obtain the corresponding local position fusion feature by enhancing the feature of the second region; and fuse the global semantic information corresponding to the video frame and the video frame before the video frame respectively to obtain the corresponding global semantic fusion feature; perform object detection on the video frame based on the local semantic fusion feature, the local position fusion feature and the global semantic fusion feature, and determine the detection boxes of the respective objects detected in the video frame. A trajectory tracking unit is configured to connect the respective detection boxes of the same object in different video frames within the same video clip for each of the video clips respectively, and obtain the respective tracking trajectories of the respective objects in each of the video clips. A re-identification unit is configured to connect the tracking trajectories of the same object in different video clips for each of the objects, and obtain the respective movement trajectories of the respective objects in the video to be identified.

12. An electronic device, characterized in that, It includes a processor and a memory, wherein the memory stores program code, and when the program code is executed by the processor, the processor is caused to execute the steps of the method according to any one of claims 1 to 10.

13. A computer-readable storage medium, characterized in that, It includes program code, and when the program code runs on an electronic device, the program code is used to cause the electronic device to execute the steps of the method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Method for performing multi-face tracking on participating athletes in sports video

    CN106022220A

  • Cross-camera pedestrian detection tracking method based on depth learning

    CN108875588A

  • Multi-target tracking method and electronic device

    CN112070807A