Method and apparatus for recognizing video objects in a video

By acquiring video object traces and performing clustering, target object clusters in the video are identified, solving the problems of low recognition efficiency and misidentification in existing technologies, and achieving more efficient and accurate video object recognition.

CN116883881BActive Publication Date: 2025-11-28TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210299447.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-25
Publication Date
2025-11-28
Estimated Expiration
2042-03-25

AI Technical Summary

Technical Problem

Existing technologies for identifying video content are typically based on processing individual video frames, resulting in low recognition efficiency, wasted resources, and susceptibility to misidentification due to noise in individual video frames.

Method used

By extracting video object traces from the video, clustering them to obtain target object clusters, and matching target objects from the target object information database, redundant parsing of video frames is reduced, thereby improving recognition accuracy.

Benefits of technology

It effectively reduces the frequency of access to the target object information database, saves processing resources, reduces false recognition caused by noise in a single video frame, and improves the accuracy of video object recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116883881B_ABST
    Figure CN116883881B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method for identifying video objects in a video, comprising: obtaining at least one video object trace from the video, each video object trace comprising a plurality of video object image frames of a same video object within a segment time of the video; clustering the at least one video object trace to obtain a target object cluster corresponding to each video object, each target object cluster comprising at least one video object trace of a corresponding video object; determining a target object matching at least part of each target object cluster from a target object information library, and identifying target object information corresponding to the target object as object information of the video object corresponding to the at least part of the target object cluster, the target object information library comprising at least one target object and target object information corresponding thereto. The method avoids parsing each video frame of the video one by one, and can identify video objects from the video with high efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of image processing, and in particular, to a method and apparatus for identifying video objects in a video, a computing device, a computer storage medium, and a computer program product. BACKGROUND

[0002] With the progress and development of Internet technology, image acquisition, image processing and other technologies, video has been widely used in people's daily life and industrial production. For example, video has become an important information carrier, and in many industries, video has also become an important production factor, promoting the improvement of production efficiency. In many application scenarios, people need to find out the target video that meets their own needs or expectations from a large number of videos, or need to judge whether a video is a video that meets their own needs, and such judgment and analysis is often time-consuming. Therefore, it is necessary to efficiently identify the content or objects in the video. Although there are ways to machine-identify video content, the existing technical means for identifying video content are usually based on the identification of a single video frame, which requires frequent operation and processing of each video frame of the video, and the identification efficiency and accuracy of the video information are low, and a large amount of processing resources are wasted. SUMMARY

[0003] Therefore, embodiments of the present disclosure provide a method and apparatus for identifying video objects in a video, a computing device, a computer storage medium, and a computer program product, which are expected to overcome some or all of the defects of the related art mentioned above and other possible defects.

[0004] The method for identifying video objects in a video provided by the embodiments of the present disclosure comprises: obtaining at least one video object trace from the video, wherein each video object trace comprises a plurality of video object image frames of a same video object within a segment time of the video, and each video object image frame is at least a part of a video frame in the video comprising the same video object; clustering the at least one video object trace to obtain each target object cluster corresponding to each video object in the video, wherein each target object cluster comprises at least one video object trace of the video object corresponding to the target object cluster; determining a target object matching at least part of the target object clusters from a target object information library, and identifying target object information corresponding to the target object as object information of the video object corresponding to the at least part of the target object clusters, wherein the target object information library comprises at least one target object and target object information corresponding to the at least one target object.

[0005] According to some embodiments of the present disclosure, obtaining the at least one video object trace from the video comprises: decoding the video to obtain a plurality of video frames; identifying a plurality of video frames including video objects from the plurality of video frames; detecting each video object in the plurality of video frames including video objects to obtain a plurality of video object image frame sequences, wherein each video object image frame sequence in the plurality of video object image frame sequences comprises a plurality of video object image frames that are continuous in time; and generating the at least one video object trace based on the plurality of video object image frame sequences.

[0006] According to some embodiments of the present disclosure, wherein generating the at least one video object trace based on the plurality of video object image frame sequences comprises: removing a calibrated video object image frame from the plurality of video object image frame sequences to obtain the at least one video object trace, wherein the calibrated video object image frame comprises at least one of a first video object image frame in the plurality of video object image frame sequences whose definition is lower than a definition threshold and a second video object image frame in the plurality of video object image frame sequences whose position of the video object is unchanged within a predetermined time period.

[0007] According to some embodiments of the present disclosure, wherein generating the at least one video object trace based on the plurality of video object image frame sequences comprises: ordering each video object image frame in each video object image frame sequence in the plurality of video object image frame sequences in time sequence to form a video object image frame set; determining a similarity between each pair of adjacent video object image frames in the video object image frame set; and truncating the video object image frame set between any pair of adjacent video object image frames in the video object image frame set in response to a similarity between the pair of adjacent video object image frames being lower than a first similarity threshold, thereby obtaining at least two video object traces.

[0008] According to some embodiments of the present disclosure, determining the target objects matching at least part of the target object clusters from the target object information library comprises: determining a time length of each target object cluster, respectively, wherein the time length is a sum of durations of video object traces included in each target object cluster; and removing a target object cluster whose time length is less than a first time threshold from the target object clusters, thereby obtaining the at least part of the target object clusters.

[0009] According to some embodiments of the present disclosure, the clustering of the at least one video object trace to obtain each target object cluster corresponding to each video object in the video comprises: clustering the at least one video object trace based on spatio-temporal information corresponding to each video object trace in the at least one video object trace to obtain an original object cluster corresponding to each video object, wherein the spatio-temporal information comprises at least one of time information and position information of the video object trace in the video; and taking the original object cluster corresponding to each video object as the target object cluster.

[0010] According to some embodiments of the present disclosure, the clustering of the at least one video object trace to obtain each target object cluster corresponding to each video object in the video comprises: clustering the at least one video object trace based on spatio-temporal information corresponding to each video object trace in the at least one video object trace to obtain an original object cluster corresponding to each video object, wherein the spatio-temporal information comprises at least one of time information and position information of the video object trace in the video; determining a similarity between the original object cluster and a sample object cluster in an object cluster archive, wherein the object cluster archive comprises at least a sample object cluster feature of the sample object cluster; and merging the sample object cluster into the original object cluster to obtain the target object cluster in response to the similarity between the original object cluster and the sample object cluster being higher than a second similarity threshold.

[0011] According to some embodiments of the present disclosure, the determining of the similarity between the original object cluster and a sample object cluster in an object cluster archive comprises: determining an original object cluster feature of the original object cluster, wherein the original object cluster feature comprises an average feature of a plurality of video object image frames of each video object trace in the original object cluster; and determining a similarity between the original object cluster feature and a sample object cluster feature, wherein the sample object cluster feature comprises an average feature of a plurality of video object image frames of each video object trace in the sample object cluster.

[0012] According to some embodiments of the present disclosure, the clustering of the at least one video object trace based on the spatio-temporal information corresponding to each video object trace to obtain the original object cluster corresponding to each video object comprises: determining trace features of each video object trace, wherein the trace features comprise average features of multiple video object image frames in the video object trace; determining similarities between the trace features of different video object traces in the at least one video object trace and matching degrees of the spatio-temporal information corresponding to the different video object traces; establishing a same object probability matrix based on the similarities between the trace features and the matching degrees of the spatio-temporal information, wherein each element in the same object probability matrix represents a probability that the different video object traces point to the same video object; establishing an adjacency matrix based on the same object probability matrix, wherein each element in the adjacency matrix represents whether the different video object traces point to the same video object; and obtaining the original object cluster corresponding to each video object based on the adjacency matrix.

[0013] According to some embodiments of the present disclosure, the target object information at least comprises target object features, and determining target objects matching at least part of the target object clusters from a target object information library comprises: determining similarities between the target object clusters and target objects in the target object information library based on target object cluster features of each target object cluster in the at least part of the target object clusters and target object features of the target objects in the target object information library, wherein the target object cluster features comprise average features of multiple video object image frames included by each video object trace in the target object cluster; and obtaining, from the target object information library, target objects with similarities higher than a third similarity threshold to the target object clusters as the target objects matching the at least part of the target object clusters.

[0014] According to some embodiments of the present disclosure, the target object information library includes a first target object information sub-library and a second target object information sub-library, the first target object information sub-library includes at least one first target object and first target object information corresponding to the first target object, the second target object information sub-library includes at least one second target object and second target object information corresponding to the second target object, and the second target object information sub-library has a higher priority than the first target object information sub-library; and wherein determining the target objects matching at least part of the target object clusters from the target object information library includes: retrieving the first target objects and the second target objects matching the target object clusters from the first target object information sub-library and the second target object information sub-library, respectively; and determining the second target objects matching the target object clusters in the second target object information sub-library as the target objects matching at least part of the target object clusters.

[0015] Another embodiment of the present disclosure provides an apparatus for identifying video objects in a video, comprising: a video object trace obtaining module configured to obtain at least one video object trace from the video, wherein each of the at least one video object trace includes a plurality of video object image frames of a same video object within a segment time of the video, and each of the plurality of video object image frames is a part of a video frame of the video including the same video object; a clustering module configured to cluster the at least one video object trace to obtain respective target object clusters corresponding to respective video objects in the video, each target object cluster including at least one video object trace of the video object to which the target object cluster corresponds; and a target object determining module configured to determine target objects matching at least part of the target object clusters from a target object information library, and identify target object information corresponding to the target objects as object information of the video objects to which the at least part of the target object clusters correspond, the target object information library including at least one target object and target object information corresponding to the at least one target object.

[0016] Yet another embodiment of the present disclosure provides a computing device, comprising: a memory configured to store computer executable instructions; and a processor configured to execute the method as described in any of the preceding embodiments of the method for identifying video objects in a video when the computer executable instructions are executed by the processor.

[0017] Still another embodiment of the present disclosure provides a computer readable storage medium storing computer executable instructions, which when executed perform the method for identifying video objects in a video as described in any of the preceding embodiments.

[0018] Still another embodiment of the present disclosure provides a computer program product comprising a computer program, wherein the computer program, when executed by a processor, implements the steps of the method for identifying a video object in a video according to any one of the preceding embodiments.

[0019] For the method or device for identifying a video object in a video provided by the embodiments of the present disclosure, the object information in the video is identified by obtaining video object traces, clustering the video object traces, and determining the target object cluster based on the clustering of the video object traces, thereby avoiding the tedious individual analysis of each video frame of the video, effectively reducing the access frequency to the target object information library, and saving processing resources. Meanwhile, the matching between the target object cluster and the target object in the target object information library is determined based on the target object cluster, which can effectively reduce the misidentification caused by single video frame noise, and further improve the accuracy of identifying the object information of the video object in the video.

[0020] These and other advantages of the present disclosure will become apparent from the embodiments described herein after and with reference to the drawings. BRIEF DESCRIPTION OF DRAWINGS

[0021] Embodiments of the present disclosure will now be described in more detail and with reference to the drawings, in which:

[0022] Figure 1 An example implementation environment of the method for identifying a video object in a video according to some embodiments of the present disclosure is shown;

[0023] Figure 2 A flowchart of the method for identifying a video object in a video according to an embodiment of the present disclosure is shown;

[0024] Figure 3 An example of displaying object information by applying the method for identifying a video object in a video according to some embodiments of the present disclosure is shown;

[0025] Figure 4 A process of obtaining at least one video object trace from a video in the method for identifying a video object in a video according to some embodiments of the present disclosure is shown;

[0026] Figure 5 A process of generating at least one video object trace based on a plurality of video object image frame sequences according to some embodiments of the present disclosure is shown;

[0027] Figure 6 To illustrate the phenomenon of shot jump that may exist in a video;

[0028] FIGS. 7(a) to 7(c) show a process of clustering the at least one video object trace to obtain each target object cluster corresponding to each video object in the video according to some embodiments of the present disclosure;

[0029] Figure 8 A specific example of clustering the at least one video object trace to obtain each original object cluster corresponding to each video object based on the spatio-temporal information corresponding to each video object trace in the at least one video object trace according to some embodiments of the present disclosure is shown;

[0030] Figure 9 A process of determining target objects matching at least part of the target object clusters from the target object information library according to some embodiments of the present disclosure is shown;

[0031] Figure 10 A system for identifying object information of video objects in a video according to embodiments of the present disclosure is schematically outlined;

[0032] Figure 11 An example of an interface for operating and controlling the running of the method according to embodiments of the present disclosure on a terminal device is schematically shown;

[0033] Figure 12 A block diagram of an apparatus for identifying object information of video objects in a video according to embodiments of the present disclosure is shown; and

[0034] Figure 13 An example system is illustrated that includes an example computing device that is representative of one or more systems and / or devices that can implement various methods or apparatuses described herein. DETAILED DESCRIPTION

[0035] The following description provides specific details of various embodiments of the present disclosure in order to provide a thorough description of the various embodiments of the present disclosure and to comply with the written description requirements of the Patent Act. It is understood that the present disclosure can be practiced without these details. In some instances, well-known structures or functions have not been shown or described in detail in order to avoid unnecessarily obscuring the description of the embodiments of the present disclosure. It is also understood that the terminology used herein is for the purpose of describing particular embodiments of the present disclosure only and is not intended to limit the scope of the present disclosure.

[0036] Artificial Intelligence (AI) is the use of digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology of computer science, which attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines, so that machines have the functions of perception, reasoning and decision-making.

[0037] Artificial intelligence technology is a comprehensive discipline, involving a wide range of fields, both hardware and software technologies. Artificial intelligence basic technologies generally include technologies such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction system, mechatronics, etc. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning, automatic driving, intelligent transportation, automatic control, etc.

[0038] Machine Learning (ML) is a multi-disciplinary subject, involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory, etc. It is a subject that studies how computers simulate or implement human learning behavior to acquire new knowledge or skills, and reorganize existing knowledge structure to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental approach to making computers intelligent, and its applications are widespread in various fields of artificial intelligence. Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rule-based learning.

[0039] With the research and progress of artificial intelligence technology, artificial intelligence technology has been researched and applied in many fields, such as common smart home, smart wearable devices, virtual assistants, smart speakers, smart marketing, unmanned vehicles, autonomous vehicles, drones, robots, intelligent medical care, intelligent customer service, Internet of Vehicles, autonomous driving, intelligent transportation, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0040] In order to facilitate the understanding of the embodiments of the present disclosure, the following will first introduce several concepts.

[0041] Video Objects: The “video objects” mentioned in this article can be any object that appears in the video. Users can also determine the specific content of the video objects according to their own needs or application scenarios. Examples of video objects include, but are not limited to, items (e.g., lamps, balls, etc.), people, human body parts (e.g., faces, etc.), animals, animal body parts, items attached to the bodies of animals or people, etc.

[0042] Video object traces refer to multiple video object image frames of a specific video object within a segment of time in a video. For example, the position of the video object may move or change within this segment of time, and correspondingly, these multiple video object image frames can form a video object trace within that segment of time. Therefore, a video object trace can correspond to a specific segment of time in the video. For the same video object, the video may contain multiple video object traces corresponding to different segment times.

[0043] Video object image frame: At least a portion of video frames that include the same video object. For example, if the video object is a face, the video object image frame is the region corresponding to the face in the video frame that includes at least the face.

[0044] Figure 1 The illustration depicts an exemplary implementation environment of a method for identifying video objects in a video according to some embodiments of the present disclosure. For example... Figure 1 As shown, various types of terminal devices communicate with servers via networks. Examples of terminal devices include, but are not limited to, mobile phones, personal computers, tablets, laptops, smart voice interaction devices, smart home appliances, in-vehicle terminals, and aircraft. Servers can be, for example, independent physical servers, server clusters or distributed systems composed of multiple physical servers, or cloud servers providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms.

[0045] A terminal device can acquire or receive video and send it to a server. The server can obtain video object traces for video objects based on the received video, and cluster the obtained video object traces to obtain target object clusters corresponding to each video object. Each target object cluster includes at least one video object trace of the video object corresponding to that target object cluster. The server can store a target object information database, and determine target objects that match at least a portion of the target object clusters from the target object information database. The server then identifies the target object information corresponding to the target object as the object information of the video object corresponding to the at least a portion of the target object clusters. The target object information database includes at least one target object and target object information corresponding to the at least one target object. Alternatively, the server can temporarily construct the target object information database by receiving target objects and corresponding target object information from the terminal device.

[0046] This article does not restrict which steps in the method for identifying video objects in a video are performed by the server and which steps are performed by the terminal device. For simplicity, the following detailed explanation will focus on the example where the steps in the method for identifying video objects in a video are performed by the server.

[0047] Figure 2 A flowchart illustrating a method for identifying video objects in a video according to some embodiments of the present disclosure is shown schematically. Figure 2 As shown, according to some embodiments of this disclosure, a method for identifying video objects in a video includes: S110, obtaining at least one video object trace from the video, wherein each video object trace includes multiple video object image frames of the same video object within a segment time of the video, and each of the multiple video object image frames is a part of video frames in the video that include the same video object; S120, clustering the at least one video object trace to obtain each target object cluster corresponding to each video object in the video, each target object cluster including at least one video object trace of the video object corresponding to the target object cluster; S130, determining a target object matching at least a portion of the target object clusters from a target object information database, and identifying the target object information corresponding to the target object as object information of the video object corresponding to the at least a portion of the target object clusters, wherein the target object information database includes at least one target object and target object information corresponding to the at least one target object.

[0048] The steps S110 to S130 described above can all be performed by the terminal device, can all be performed by the server, or some steps are performed by the terminal device and the other steps are performed by the server, that is, are cooperatively performed by both the terminal device and the server.

[0049] In step S110, the server can obtain a video (for example, by uploading a video file to the server via the terminal device), and on this basis, can perform frame analysis on the video, identify a video object in the video and detect the position of the video object, thereby obtaining at least one video object trace, for example, at least one face trace. It can be understood that a certain video object can appear in different time periods of the video, or can not exist in some time periods of the video, so a single video object trace can correspond to a certain time period in the total time length of the video, and the video object trace mentioned herein includes multiple video object image frames of the same video object in a time period of the video. The video object image frame is obtained based on the frame analysis and identification of the video, and the video object image frame is a part of a video frame in which the video object corresponding to the video object image frame is included. For example, the video object image frame can be a face image frame, that is, a part corresponding to a face in a video frame in which the face is included. By step S110, the server can obtain a video object trace corresponding to each video object, and each video object can correspond to at least one video object trace.

[0050] In step S120, the server clusters the at least one video object trace obtained in step S110, and as much as possible, classifies the video object traces pointing to the same video object together. For example, in the case where the video object is a face, different faces can have different features, and correspondingly, different face traces also exhibit different trace features. According to the trace features of each face trace, the same or similar face traces can be classified into a target object cluster, and each face trace included in the target object cluster points to the same face. Therefore, the target object cluster mentioned herein corresponds to a video object, and each target object cluster includes at least one video object trace of the video object to which the target object cluster corresponds. In other words, by step S120, the server clusters the at least one video object trace obtained in step S110 to form at least one set of video object traces, and each set of video object traces is a target object cluster.

[0051] In step S130, the server determines target objects matching at least part of each target object cluster obtained in step S120 from the target object information library, and identifies the target object information of the target objects matching the target object cluster as the object information of the video object corresponding to the target object cluster. On this basis, the server can send the above-mentioned target object information to the terminal device or output or display the target object information in other ways. For example, in the case where the video object is a human face, the above-mentioned target object information can include the name of the person, the image of the person, etc.

[0052] In order to identify the object information of the video object in the video, the conventional solution is a recognition method for a single video frame. This single-frame processing recognition method needs to perform image analysis operations such as video object detection and video object matching on each video frame, and the processing efficiency is not satisfactory, and the processing resources are greatly wasted. The inventors of the present application realize that the same video object has certain space-time relationship in at least part of the video frames, and the video object appearing at the same position in the video frames in succession has a high probability of being the same video object. Therefore, the above-mentioned embodiments of the present disclosure provide a method for identifying a video object in a video by obtaining a video object trace, clustering the video object trace, and identifying the object information in the video based on the target object cluster obtained by clustering the video object trace, which greatly reduces the redundant analysis of each video frame, effectively reduces the access frequency of the target object information library, and saves processing resources. At the same time, based on the target object cluster, the matching between the target object cluster and the target object in the target object information library can effectively reduce the misrecognition caused by single video frame noise, and further improve the accuracy of identifying the object information of the video object in the video.

[0053] Figure 3 An exemplary application scenario of the method for identifying a video object in a video according to the above-mentioned embodiments is schematically shown. Figure 3 A frame of the video is shown on the left side, and in this scenario, the video object is the face of a person. The object information obtained by applying the method for identifying a video object in a video according to the above-mentioned embodiments is shown on the right side, and it can be determined that the video includes two target objects in the target object information library. As shown in Figure 3 It is shown that the industries to which the two target objects belong are singers and teachers respectively, and the names of the two target objects can also be shown.

[0054] According to some embodiments of the present disclosure, as Figure 4As shown, the step S110 of obtaining at least one video object trace from the video comprises: S1101, decoding the video to obtain a plurality of video frames; S1102, identifying a plurality of video frames including video objects from the plurality of video frames. For example, feature extraction can be performed on the video objects in the video frames, and then a convolutional neural network can be used to identify the video frames including video objects. Any appropriate image recognition method can be used to identify the plurality of video frames including video objects from the plurality of video frames, which is not limited herein. S1103, detecting each video object in the plurality of video frames including video objects to obtain a plurality of video object image frame sequences, each of the plurality of video object image frame sequences comprising a plurality of video object image frames that are continuous in time of the video. The detection technology for detecting video objects in the video is not specifically limited herein, for example, the detection technology can include but is not limited to an IoU (Intersection over Union) based object detection technology. In step S1103, each video object image frame sequence obtained by detecting each video object in the video frames including video objects can correspond to a time period in the video, and the video object appears in the video during the time period, so the plurality of video object image frames included in each video object image frame sequence are continuous in time of the video. S1104, and generating the at least one video object trace based on the plurality of video object image frame sequences.

[0055] In some embodiments, the step S1104 can comprise: removing a calibration video object image frame from the plurality of video object image frame sequences to obtain the at least one video object trace, the calibration video object image frame comprising at least one of a first video object image frame with a clarity lower than a clarity threshold and a second video object image frame with a position of the video object unchanged within a predetermined time period in the plurality of video object image frame sequences. Removing the first video object image frame from the plurality of video object image frame sequences can filter out low-quality images with poor clarity, thereby facilitating improvement of the accuracy of the object information of the video objects in the video. In addition, in some scenarios, only the movable video objects of the video can be concerned, and therefore, the second video object image frame with the position of the video object unchanged within the predetermined time period can be regarded as a non-concerned video object. Removing the second video object image frame can reduce unnecessary interference to the object information of the video objects in the video.

[0056] It can be understood that the detection of each video object in the plurality of video frames including video objects also involves feature extraction for the video object. In some embodiments, the detection of each video object in the plurality of video frames including video objects includes video object registration, thereby facilitating the accuracy of feature extraction. For example, in the case of a video object being a human face, the key points of the human face (e.g., eyes, nose) can be registered to enhance the accuracy of feature extraction. Algorithms for video object registration may, for example, include active appearance models and active shape models.

[0057] As shown in FIG. 5A, in another embodiment, the step S1104 described above can include the following steps: S501, ordering the plurality of video object image frames in each video object image frame sequence in time sequence to form a video object image frame set; S502, determining the similarity between each adjacent two video object image frames in the video object image frame set; S503, truncating the video object image frame set between any adjacent two video object image frames in the video object image frame set in response to the similarity between the any adjacent two video object image frames being lower than a first similarity threshold, thereby obtaining at least two video object traces. Examples of the similarity mentioned herein include, but are not limited to, cosine similarity, Euclidean similarity, and Hamming distance similarity, and the same applies to other similarities mentioned below. The specific method of determining the similarity is not specifically limited herein. Figure 5 The steps S501-S503 described above actually order the plurality of video object image frames in each video object image frame sequence in time sequence, and then perform similarity operation on any adjacent two video object image frames in time. If the similarity between the adjacent two video object image frames is low, it indicates that there are probably at least two different video objects in the video object image frame sequence, and therefore the video object image frame sequence can be truncated between the adjacent two video object image frames with low similarity, thereby obtaining at least two video object traces. The steps S501-S503 described above are very beneficial to improving the accuracy of each video object trace obtained.

[0058] In some application scenarios, the shots for the video object in the video often have a jump phenomenon.

[0059] Five consecutive video frames in an exemplary video are schematically shown, the video object of the video includes a human face, and the video object traces obtained by the method described above are shown in FIG. 5B. Figure 6 Figure 6 ​In the example shown, the face of person A is marked with a dashed box in video frames I, II, III and IV. However, immediately after the fourth video frame IV, there is a video frame V in which person A does not exist, but video frame V is still continuous with video frame IV in time, which is caused by a lens jump when the video is being shot. In video frame V, the video object corresponding to the position of person A in video IV is person B, and obviously, video frame V cannot form a video object track together with video frames I, II, III and IV. Application Figure 5 The steps S501-S503 shown can realize truncating the video object image frame set between video frame IV and video frame V, and video frame V belongs to a different video object track from video frames I, II, III and IV, thereby well solving the problem of lens jump in the video and improving the accuracy of each video object track obtained.

[0060] Next, the step S120 described above is specifically explained by way of example. FIGS. 7(a)-7(c) schematically show two different examples for the step S120 described above. As shown in FIG. 7(a), in some embodiments, the clustering of the at least one video object track to obtain each target object cluster corresponding to each video object in the video comprises: S701, clustering the at least one video object track based on the spatiotemporal information corresponding to each video object track in the at least one video object track to obtain an original object cluster corresponding to each video object, the spatiotemporal information comprising at least one of time information and position information of the video object track in the video; and S702', taking the original object cluster corresponding to each video object as the target object cluster. That is, in the example, the original object cluster obtained based on the step S701 is the target object cluster corresponding to the video object.

[0061] Alternatively, as shown in FIG. 7(b), in another example, the clustering of the at least one video object trace to obtain the respective target object cluster corresponding to each video object in the video comprises: S701, clustering the at least one video object trace based on the spatio-temporal information corresponding to each video object trace to obtain a raw object cluster corresponding to each video object, the spatio-temporal information including at least one of time information and location information of the video object trace in the video; S702, determining the similarity between the raw object cluster and a sample object cluster in an object cluster archive, wherein the object cluster archive at least includes sample object cluster features of the sample object cluster; and S703, in response to the similarity between the raw object cluster and the sample object cluster being higher than a second similarity threshold, merging the sample object cluster into the raw object cluster to obtain the respective target object cluster.

[0062] Here, the sample object cluster can be understood as an existing object cluster stored in the server, or an existing object cluster obtained by clustering based on an existing video object trace related to the video, or the sample object cluster can also be an existing object cluster obtained by video object detection and clustering on another video. The object cluster archive can store sample object cluster features of the sample object cluster, and can also store an identification (ID) of each sample object cluster and related information of the video object trace included therein.

[0063] In some embodiments, in step S703, in the case where the similarity between the raw object cluster and the sample object cluster is higher than the second similarity threshold, the sample object cluster is merged into the raw object cluster to obtain the respective target object cluster. In this example, the sample object cluster features of the sample object cluster can be merged into the raw object cluster features of the raw object cluster, and accordingly, the video object traces contained in the sample object cluster will be combined with the video object traces in the raw object cluster. In this way, the possibility of video traces pointing to the same video object being classified into the same object cluster can be improved, and the features of the target object cluster corresponding to the video object can be enriched, further improving the accuracy of identifying the video object in the video.

[0064] With reference back to FIG. 7(c), in some embodiments, the step S702 described above can include: S702a, determining a raw object cluster feature of the raw object cluster, the raw object cluster feature comprising an average feature of a plurality of video object image frames of each video object track in the raw object cluster; S702b, determining a similarity between the raw object cluster feature and a sample object cluster feature, the sample object cluster feature comprising an average feature of a plurality of video object image frames of each video object track in the sample object cluster. The video object feature of each video object image frame can be obtained by any suitable method, including but not limited to convolutional neural network, support vector machine, etc., so that the raw object cluster feature of the raw object cluster, the sample object cluster feature of the sample object cluster, and the track feature of each video object track mentioned below can be obtained by averaging operation.

[0065] Figure 8 FIG. 7(c) illustrates a specific example of the step S701 described above, in which each video object track is clustered based on the spatio-temporal information corresponding to each video object track to obtain a raw object cluster corresponding to each video object. In this exemplary clustering process, a conditional random field model is used. The conditional random field model is a conditional probability distribution model P(Y|X), which represents a Markov random field of another group of output random variables Y given a group of input random variables X, that is, the conditional random field assumes that the output random variables form a Markov random field. In the conditional random field model, the distribution of the random variable Y is a conditional probability, and the given observation is the random variable X. In principle, the graph model layout of the conditional random field can be given arbitrarily, and the commonly used layout is a chain architecture. The chain architecture has high-efficiency algorithms available for training, inference, or decoding.

[0066] As Figure 8The clustering of the at least one video object trace to obtain the original object cluster corresponding to each video object comprises: S801, determining trace features of each video object trace, the trace features comprising average features of multiple video object image frames in the video object trace; S802, determining similarities between the trace features of different video object traces in the at least one video object trace, and matching degrees of spatio-temporal information corresponding to the different video object traces; S803, establishing a same-object probability matrix based on the similarities between the trace features and the matching degrees of the spatio-temporal information, each element in the same-object probability matrix representing a probability that the different video object traces point to a same video object; S804, establishing an adjacency matrix based on the same-object probability matrix, each element in the adjacency matrix representing whether the different video object traces point to a same video object; and S805, obtaining the original object cluster corresponding to each video object based on the adjacency matrix.

[0067] In some embodiments, the spatio-temporal information comprises at least one of time information and position information of a video object trace in a video, the position information referring to a position of a video object in an environment space embodied by the video, which can be obtained from an image acquisition device (e.g., a camera) that acquires images of the video object. In a case where multiple image acquisition devices are involved, the multiple image acquisition devices can be numbered, and different image acquisition devices can acquire images of the video object at different spatial positions, accordingly, in this case, different spatial positions corresponding to the video object can be obtained based on image acquisition devices with different serial numbers. The time information can correspond to a time stamp at which the image acquisition device acquires images of the video object. As mentioned above, a same video object has certain spatio-temporal relationship in at least part of video frames of a video, and a video object appearing at a same position or with a small change in position in temporally consecutive video frames is probably directed to a same video object. Therefore, the matching degree of spatio-temporal information between different video object traces can be determined based on spatio-temporal information of each video object image frame in the video object traces, and the matching degree of spatio-temporal information between different video object traces can indicate a probability that the different video object traces point to a same video object from a spatio-temporal perspective. In step S803, the same-object probability matrix is established by jointly considering the similarities between the trace features and the matching degrees of the spatio-temporal information, so as to obtain the probability that the different video object traces point to a same video object. This probability can also be referred to as a joint probability. In one example, the joint probability G can be represented by the following formula:

[0068]

[0069] where P1represents the probability that different video object traces point to the same video object from the perspective of space-time, and simrepresents the similarity between trace features of different video object traces (e.g., can be the cosine similarity between trace features). Based on sim, the probability that different video object traces point to the same video object can be calculated as , represents the probability that different video object traces point to the same video object given the similarity between trace features. represents the probability that different video object traces point to different video objects given the similarity between trace features. In some embodiments, in the case that N video object traces are obtained based on the video through the aforementioned step S110, the same-object probability matrix can be an N*N matrix, and the element N ij i.e., represents the probability that the i-th video object trace and the j-th video object trace point to the same video object, and N ij The value of P1may be between 0 and 1.

[0070] In step S804, an adjacency matrix is established based on the same-object probability matrix, and each element in the adjacency matrix represents whether the different video object traces point to the same video object. The dimension of the adjacency matrix can be the same as that of the same-object probability matrix, e.g., also an N*N matrix. If each video object trace is regarded as a node, the video object traces can be constructed into a directed graph, and the edges in the directed graph represent whether two video object traces point to the same video object, and then the adjacency matrix can be regarded as the matrix representation of the directed graph. In step S805, a suitable clustering algorithm (e.g., mean shift clustering algorithm, spectral clustering algorithm, density-based clustering algorithm, etc.) can be used to obtain the original object cluster corresponding to each video object based on the adjacency matrix. The applicable clustering algorithm is not specifically limited herein.

[0071] In the above embodiments, the constructed same-object probability matrix combines the similarity of trace features of video object traces and the matching degree of space-time information of video object traces, and thus the clustering of video object traces is realized as clustering based on a conditional random field model, i.e., the clustering problem in the feature space is converted into a conditional random field model in the probability space. The conventional clustering method only calculates the similarity of different features based on the distance in the feature space, and ignores the space-time information. In the probability space, the same-object probability matrix of the whole sample can be constructed by combining the above space-time information (as prior information), which has stronger scalability.

[0072] In some embodiments, video objects with short appearance time in the video can be regarded as non-target objects of no interest, and thus, non-target objects can be removed from each target object cluster obtained from clustering to reduce the interference of non-target objects on the identification of video objects in the video. At this time, the step S130 (determining target objects matching at least part of the target object clusters from the target object information library) can include the following steps: determining the time length of each target object cluster, which is the sum of the duration of video object traces included in each target object cluster; and removing target object clusters with time length less than a first time threshold from the target object clusters to obtain the at least part of the target object clusters.

[0073] In some embodiments, the target object information stored in the target object information library includes at least target object features, and can also include target object images, target object identifiers and other information. As shown in Figure 9 The step S130 (determining target objects matching at least part of the target object clusters from the target object information library) includes: S901, determining the similarity between the target object cluster and the target object based on the target object cluster features of each target object cluster in the at least part of the target object clusters and the target object features of the target objects in the target object information library, the target object cluster features including the average features of the multiple video object image frames included in each video object trace in the target object cluster; and S902, obtaining target objects with similarity to the target object cluster higher than a third similarity threshold from the target object information library as the target objects matching the at least part of the target object clusters. Based on the target objects obtained in step S902, the target object information of the target objects can be output or displayed to indicate the result of identifying video objects in the video in an intuitive way: there is content related to the target objects in the video. For example, as shown in Figure 3 .

[0074] According to another embodiment of the present disclosure, the target object information library includes a first target object information sub-library and a second target object information sub-library, the first target object information sub-library includes at least one first target object and first target object information corresponding to the first target object pre-stored, the second target object information sub-library includes at least one second target object and second target object information corresponding to the second target object, and the second target object information sub-library has a higher priority than the first target object information sub-library. The step S130 (determining the target objects matching at least part of the target object clusters from the target object information library) described above includes: in response to retrieving the first target objects and the second target objects matching the target object clusters from the first target object information sub-library and the second target object information sub-library respectively; and determining the second target objects matching the target object clusters in the second target object information sub-library as the target objects matching at least part of the target object clusters. In some embodiments, the first target object information sub-library described above can be pre-stored in the server, and the second target object information sub-library can be a video object information library constructed by the user himself / herself through the information (e.g., pictures) of the last video object. In this embodiment, the target objects matching at least part of the target object clusters can be determined from the pre-stored first target object information sub-library and the user-defined second target object information sub-library simultaneously, and the target objects matching at least part of the target object clusters obtained based on the user-defined second target object information sub-library are preferentially provided to the user, thereby improving the satisfaction of the user.

[0075] From the perspective of the system, the method of identifying the video objects in the video described in the above embodiments can be summarized as Figure 10 The system is shown. As Figure 10 The system of identifying the object information of the video objects in the video can include a connection layer, a video processing layer and a target object matching layer, and a data storage. The connection layer can be regarded as an interface for receiving the video, and the video is provided to the system through the connection layer. The video obtains the target object clusters through the processing of the video processing layer. The video processing layer includes video object trace detection and video object trace clustering. It can be understood that the video object trace detection and the video object trace clustering can both include a feature extraction operation to obtain trace features of the video object traces and target object cluster features of the target object clusters. In the target object matching layer, the target objects matching at least part of the target object clusters can be determined from a first target object information sub-library (e.g., a built-in general target object information library) and a second target object information sub-library (e.g., a user-defined target object information library) of the system. The data storage can include the storage of the video object traces, the target object clusters, the original object clusters, the target object information, etc.

[0076] To facilitate the user to recognize the object information of the video object in the video by using the method for recognizing the object information of the video object in the video provided by the embodiments of the present disclosure, a corresponding operation interface can be provided for the user on the terminal device. For example, Figure 11 An example of the operation interface on the terminal device is shown. Figure 11 The operation interface shown in FIG. 1 includes a list of videos to be recognized (e.g., video 1, video 2, video 3), video duration, an operation identifier for starting the recognition function (start recognition), and the recognition status for each video, etc.

[0077] According to another aspect of the present disclosure, a device for recognizing object information of a video object in a video is provided, as shown in Figure 12 The device 1000 includes: a video object trace obtaining module 1000a, configured to obtain at least one video object trace from the video, wherein each video object trace in the at least one video object trace includes a plurality of video object image frames of a same video object within a segment time of the video, and each video object image frame in the plurality of video object image frames is a part of a video frame in the video that includes the same video object; a clustering module 1000b, configured to cluster the at least one video object trace to obtain each target object cluster corresponding to each video object in the video, and each target object cluster includes at least one video object trace of the video object corresponding to the target object cluster; and a target object determining module 1000c, configured to determine a target object matching at least part of the target object clusters from a target object information library, and recognize target object information corresponding to the target object as the object information of the video object corresponding to the at least part of the target object clusters, wherein the target object information library includes at least one target object and target object information corresponding to the at least one target object. The device can efficiently recognize the object information of the video object from the video, or in other words, the device can efficiently determine whether the target object is included in the video.

[0078] In yet another aspect of the present disclosure, a computing device is provided, which includes a memory configured to store computer executable instructions, and a processor configured to perform the steps in the method as described in any of the preceding embodiments when the computer executable instructions are executed by the processor.

[0079] In particular, according to the embodiments of the present disclosure, the methods in the methods described above with reference to the flowcharts can be implemented as computer programs. For example, the embodiments of the present disclosure provide a computer program product, which includes a computer program carried on a computer readable medium, and the computer program includes program codes for performing at least one step in the method embodiments of the present disclosure.

[0080] Another embodiment of the present disclosure provides one or more computer-readable storage media having computer-readable instructions stored thereon that, when executed, implement a method of identifying a video object in a video according to some embodiments of the present disclosure. The various steps of the method of identifying a video object in a video can be translated into computer-readable instructions by programming, thereby stored in the computer-readable storage media. When such computer-readable storage media is read or accessed by a computing device or computer, the computer-readable instructions therein are executed by a processor on the computing device or computer to implement the method of identifying a video object in a video.

[0081] Figure 13 An example system 1100 is illustrated that includes an example computing device 1110 representative of one or more systems and / or devices that can implement the techniques described in the various embodiments herein. The computing device 1110 can be, for example, a server of a service provider, a device associated with the server, a system on a chip, and / or any other suitable computing device or computing system. The above description makes reference to the computing device 1110 executing instructions on processors to perform the described operations. However, it will be apparent that any machine that can execute a set of instructions can be utilized. In fact, the term "processor" as used herein can refer to one or more processors that are collectively configured to execute a set of instructions. The set of instructions can include various commands that instruct the machine to perform the methods and processes described herein. The set of instructions can be stored as a whole, or one or more parts of the set of instructions can be distributed over various computers (e.g., computers 1102 and 1104) so that executing the instructions results in a machine carrying out the methods and processes described herein. Figure 12 The apparatus 1000 for identifying object information of a video object in a video described above can take the form of the computing device 1110. Alternatively, the apparatus 1000 for identifying object information of a video object in a video can be implemented as a computer program in the form of the application 1116.

[0082] As Figure 13 The illustrated example computing device 1110 includes a processing system 1111, one or more computer-readable media 1112, and one or more input / output interfaces 1113 that are communicatively coupled with one another. Although not illustrated, the computing device 1110 can also include a system bus or other data and command transfer system that couples the various components with one another. A system bus can include any one or combination of different bus structures, such as a memory bus or memory controller, a peripheral bus, a serial bus, a parallel bus, and / or a local bus using any of a variety of bus architectures by which processors and other digital hardware components can communicate.

[0083] The processing system 1111 is representative of the functionality performed by a processor as software instructions are executed. As such, the processing system 1111 is illustrated as including a hardware element 1114 that can be configured to perform a processor function. For example, the hardware element 1114 can include processing circuitry in the form of semiconductor(s) and / or transistors (e.g., electronic integrated circuits (ICs)). In this context, processor-executable instructions can be electronic executable instructions.

[0084] Computer-readable medium 1112 is illustrated as including memory / storage device 1115. Memory / storage device 1115 represents a memory / storage capacity associated with one or more computer-readable media. Memory / storage device 1115 may include volatile media (such as random access memory (RAM)) and / or non-volatile media (such as read-only memory (ROM), flash memory, optical disk, magnetic disk, etc.). Memory / storage device 1115 may include fixed media (e.g., RAM, ROM, fixed hard disk drive, etc.) and removable media (e.g., flash memory, removable hard disk drive, optical disk, etc.). Computer-readable medium 1112 may be configured in various other ways as further described below.

[0085] One or more I / O interfaces 1113 represent functions that allow users to input commands and information to computing device 1110 using various input devices and optionally also allow information to be presented to the user and / or other components or devices using various output devices. Examples of input devices include keyboards, cursor control devices (e.g., mice), microphones (e.g., for voice input), scanners, touch functionality (e.g., capacitive or other sensors configured to detect physical touch), cameras (e.g., capable of detecting non-touch-related motion as gestures using visible or invisible wavelengths (such as infrared frequencies), etc. Examples of output devices include display devices (e.g., monitors or projectors), speakers, printers, network interface cards, haptic-responsive devices, etc. Therefore, computing device 1110 can be configured to support user interaction in various ways as further described below.

[0086] The computing device 1110 also includes an application 1116. The application 1116 may, for example, refer to... Figure 12 The software instance of the device 1000, which describes object information of video objects in a video, is implemented in combination with other elements in the computing device 1110 to realize the techniques described herein.

[0087] This document describes various technologies within the general context of software and hardware components or program modules. Generally, these modules include routines, programs, objects, elements, components, data structures, etc., that perform specific tasks or implement specific abstract data types. As used herein, the terms "module," "function," and "component" generally refer to software, firmware, hardware, or a combination thereof. The technologies described herein are platform-independent, meaning they can be implemented on a variety of computing platforms with various processors.

[0088] Implementations of the described modules and techniques can be stored or transmitted across some form of computer-readable media. Computer-readable media can include various media that can be accessed by the computing device 1110. By way of example, and not limitation, computer-readable media can include "computer-readable storage media" and "computer-readable signal media."

[0089] In contrast to signal transmission, carrier waves, or signals per se, "computer-readable storage media" refers to media and / or devices that enable persistent storage of information and / or tangible storage of storage devices. Thus, computer-readable storage media refers to non-signal bearing media. Computer-readable storage media includes hardware such as volatile and non-volatile, removable and non-removable media implemented in a method or technology for storage of information such as computer readable instructions, data structures, program modules, logic elements / circuits, or other data. Examples of computer-readable storage media can include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical storage, hard disks, magnetic cassettes, magnetic tapes, magnetic disk storage or other magnetic storage devices, or other storage devices, tangible media, or articles of manufacture that are appropriate for storage of desired information and that can be accessed by a computer.

[0090] "Computer-readable signal media" refers to a signal-bearing medium that is configured to transmit instructions to the hardware of the computing device 1110, such as via network. Computer-readable signal media typically can take the form of modulated data signals, such as coded data signals, transmitted over a carrier wave. Signal media typically can take the form of wave signals, including without limitation, acoustic waves, sound waves, electromagnetic waves, carrier waves, and the like, as previously described. The signals can be transmitted as part of a network or a direct connection, among others.

[0091] As previously described, hardware elements 1114 and computer-readable media 1112 are representative of instructions, modules, programmable device logic and / or fixed device logic implemented in the hardware. Hardware elements can include components of an integrated circuit or the circuitry built into a system on a chip (SoC) as example. Some of these components can be dedicated to a specific function or set of functions, such as a graphics processing unit (GPU), while other components can be dedicated to a more general purpose. A system of treatment devices can include one or more hardware elements configured to carry out a particular activity or activity set. In some embodiments, the hardware elements can each include a computer-readable medium containing instructions that, when executed, complete a set of desired activities. The activities can be those described herein.

[0092] The combinations of the foregoing can also be used to implement various techniques and modules described herein. Thus, software, hardware or program modules and other program modules can be implemented as one or more instructions and / or logic embodied on some form of computer-readable storage media and / or by one or more hardware elements 1114. The computing device 1110 can be configured to implement particular instructions and / or functions corresponding to the software and / or hardware modules. Accordingly, implementation of a module as a software module and / or hardware module, for example, can be

[0093] In various implementations, the computing device 1110 can assume a variety of different configurations. For example, the computing device 1110 can be implemented as a computer-class device comprising a personal computer, desktop computer, multi-screen computer, laptop computer, netbook, etc. The computing device 1110 can also be implemented as a mobile device-class device comprising a mobile phone, portable music player, portable gaming device, tablet computer, multi-screen computer, etc. The computing device 1110 can also be implemented as a television-class device comprising a device having or connected to a generally larger screen in a casual viewing environment. These devices include televisions, set-top boxes, game consoles, etc.

[0094] The techniques described herein can be supported by these various configurations of the computing device 1110 and are not limited to the specific examples of the techniques described herein. Functionality can also be implemented all or in part through the use of distributors, such as platforms 1122 described below, on a "cloud" 1120. The cloud 1120 includes and / or is representative of the platform 1122 for resources 1124. The platform 1122 abstracts underlying functionality of hardware (e.g., servers) and software resources of the cloud 1120. The resources 1124 can include applications and / or data that can be utilized while computer processing is executed on servers that are remote from the computing device 1110. Resources 1124 can also include services provided over the Internet and / or through a subscriber network, such as a cellular or Wi-Fi network.

[0095] The platform 1122 can abstract resources and functions to connect the computing device 1110 with other computing devices. The platform 1122 can also serve to abstract scaling of resources to provide a corresponding level of scale to encountered demand for the resources 1124 that are implemented via the platform 1122. Accordingly, in an interconnected device embodiment, implementation of functionality described herein can be distributed throughout the system 1100. For example, the functionality can be implemented in part on the computing device 1110 as well as via the platform 1122 that abstracts the functionality of the cloud 1120.

[0096] It will be appreciated that, for clarity, embodiments of the disclosure have been described hereinafter with reference to different functional units. It will be apparent, however, that the functional units can be implemented in one single unit or in different units, or as part of other functional units. For example, the functionality of the functional units described as being performed by a single unit can be performed by a plurality of different units. Accordingly, references to specific functional units are only to be seen as references to suitable means for providing the described functionality, rather than indicative of a strict logical or physical structure or organization. The disclosure can be implemented in a single unit, or be physically and functionally distributed between different units and circuitry.

[0097] It will be understood that, although the terms first, second, third, etc. can be used herein to describe various devices, elements, components or parts, these devices, elements, components or parts should not be limited by these terms. These terms are only used to distinguish one device, element, component or part from another device, element, component or part.

[0098] Although the present disclosure has been described in connection with some embodiments, it is not intended to be limited to the specific form set forth herein. Rather, the scope of the present disclosure is limited only by the claims. Additionally, although individual features can be included in different claims, these can possibly depend on others and can be combined, and the inclusion in different claims does not imply that a combination of features is not feasible and / or advantageous. The order in which features are recited in a claim does not imply any specific order in which they must be worked. Additionally, the use of the term "comprising" in a claim does not exclude the presence of other elements or steps than those recited in the claim. Furthermore, the use of the article "a" or "an" preceding an element does not exclude the presence of a plurality of such elements.

Claims

1. A method for identifying video objects in a video, comprising: obtaining at least one video object trace from the video, wherein each video object trace comprises a plurality of video object image frames of a same video object in a time segment of the video, and each video object image frame is at least a part of a video frame in the video comprising the same video object; clustering the at least one video object trace to obtain a plurality of target object clusters corresponding to respective video objects in the video, wherein each target object cluster comprises at least one video object trace of the video object to which the target object cluster corresponds; and determining target objects matching at least part of the target object clusters from a target object information library, and identifying target object information corresponding to the target objects as object information of the video objects to which the at least part of the target object clusters correspond, wherein the target object information library comprises at least one target object and target object information corresponding to the at least one target object; wherein the target object information at least comprises target object features, and wherein determining the target objects matching at least part of the target object clusters from the target object information library comprises: determining similarity between the target object clusters and the target objects based on target object cluster features of each target object cluster in the at least part of the target object clusters and target object features of the target objects in the target object information library, wherein the target object cluster features comprise average features of a plurality of video object image frames comprised by respective video object traces in the target object cluster; and obtaining, from the target object information library, target objects having similarity higher than a third similarity threshold to the target object clusters as the target objects matching the at least part of the target object clusters. 2.The method of claim 1, wherein obtaining at least one video object trace from the video comprises: decoding the video to obtain a plurality of video frames; identifying a plurality of video frames comprising video objects from the plurality of video frames; detecting respective video objects in the plurality of video frames comprising video objects to obtain a plurality of video object image frame sequences, wherein each video object image frame sequence in the plurality of video object image frame sequences comprises a plurality of video object image frames that are continuous in time of the video; and generating the at least one video object trace based on the plurality of video object image frame sequences. 3.The method of claim 2, wherein generating the at least one video object trace based on the plurality of video object image frame sequences comprises: removing calibrated video object image frames from the plurality of video object image frame sequences to obtain the at least one video object trace, wherein the calibrated video object image frames comprise at least one of first video object image frames in the plurality of video object image frame sequences having a clarity lower than a clarity threshold and second video object image frames in which positions of video objects are invariant in a predetermined time period. 4.The method of claim 2, wherein generating the at least one video object track based on the plurality of video object image frame sequences comprises: ordering video object image frames in each of the plurality of video object image frame sequences in a time sequence to form a video object image frame set; determining a similarity between each pair of adjacent video object image frames in the video object image frame set; truncating the video object image frame set between any pair of adjacent video object image frames in the video object image frame set in response to a similarity between the pair of adjacent video object image frames being lower than a first similarity threshold, thereby obtaining at least two video object tracks. 5.The method of claim 1, wherein determining target objects matching at least a portion of the target object clusters from a target object information library comprises: determining a time length of each of the target object clusters, wherein the time length is a sum of durations of video object tracks included in each target object cluster; and removing target object clusters having a time length less than a first time threshold from the target object clusters, thereby obtaining the at least a portion of the target object clusters. 6.The method of claim 1, wherein clustering the at least one video object track to obtain each target object cluster corresponding to each video object in the video comprises: clustering the at least one video object track based on spatio-temporal information corresponding to each of the at least one video object track to obtain an original object cluster corresponding to each video object, wherein the spatio-temporal information comprises at least one of a temporal information and a location information of the video object track in the video; and using the original object cluster corresponding to each video object as the target object cluster. 7.The method of claim 1, wherein clustering the at least one video object track to obtain each target object cluster corresponding to each video object in the video comprises: clustering the at least one video object track based on spatio-temporal information corresponding to each of the at least one video object track to obtain an original object cluster corresponding to each video object, wherein the spatio-temporal information comprises at least one of a temporal information and a location information of the video object track in the video; determining a similarity between the original object cluster and a sample object cluster in an object cluster archive, wherein the object cluster archive comprises at least a sample object cluster feature of the sample object cluster; merging the sample object cluster to the original object cluster in response to the similarity between the original object cluster and the sample object cluster being higher than a second similarity threshold, thereby obtaining the target object cluster. 8.The method of claim 7, wherein determining the similarity between the original object cluster and the sample object cluster in the object cluster archive comprises: determining original object cluster features of the original object clusters, wherein the original object cluster features comprise average features of a plurality of video object image frames of each video object trace in the original object clusters; determining similarities between the original object cluster features and the sample object cluster features, wherein the sample object cluster features comprise average features of a plurality of video object image frames of each video object trace in the sample object clusters. 9.The method of claim 6 or 7, wherein the obtaining of each video object corresponding original object cluster based on clustering of the at least one video object trace according to corresponding spatio-temporal information of each video object trace comprises: determining trace features of each video object trace, wherein the trace features comprise average features of a plurality of video object image frames in the video object trace; determining similarities between the trace features of different video object traces in the at least one video object trace, and matching degrees of the corresponding spatio-temporal information of the different video object traces; establishing a same object probability matrix based on the similarities between the trace features and the matching degrees of the spatio-temporal information, wherein each element in the same object probability matrix represents a probability that the different video object traces point to a same video object; establishing an adjacency matrix based on the same object probability matrix, wherein each element in the adjacency matrix represents whether the different video object traces point to a same video object; and obtaining the original object cluster corresponding to each video object based on the adjacency matrix.

10. The method of claim 1, wherein the target object information library comprises a first target object information sub-library and a second target object information sub-library, the first target object information sub-library comprising at least one first target object pre-stored and first target object information corresponding to the first target object, the second target object information sub-library comprising at least one second target object and second target object information corresponding to the second target object, the second target object information sub-library having a higher priority than the first target object information sub-library. and wherein, the determining of the target objects matching at least part of the target object clusters comprises: in response to retrieving a first target object and a second target object matching the target object cluster from the first target object information sub-library and the second target object information sub-library, respectively; and determining the second target object matching the target object cluster in the second target object information sub-library as the target object matching at least part of the target object clusters. 11.An apparatus for recognizing video objects in a video, comprising: a video object trace obtaining module, configured to obtain at least one video object trace from the video, wherein each video object trace comprises a plurality of video object image frames of a same video object within a segment time of the video, and each video object image frame is at least a part of a video frame in the video comprising the same video object; a clustering module, configured to cluster the at least one video object trace to obtain each target object cluster corresponding to each video object in the video, wherein each target object cluster comprises at least one video object trace of a video object corresponding to the target object cluster; and a target object information library establishing module, configured to establish a target object information library based on the target object clusters. a target object determination module configured to determine target objects matching at least part of the target object clusters from a target object information library, and identify target object information corresponding to the target objects as object information of video objects corresponding to the at least part of the target object clusters, wherein the target object information library comprises at least one target object and target object information corresponding to the at least one target object; wherein the target object information comprises at least target object features, and the target object determination module is further configured to: determine similarity between the target object clusters and the target objects based on target object cluster features of each target object cluster in the at least part of the target object clusters and target object features of the target objects in the target object information library, wherein the target object cluster features comprise average features of a plurality of video object image frames included in each video object trace in the target object cluster; and obtain, from the target object information library, target objects having similarity higher than a third similarity threshold to the target object clusters as the target objects matching the at least part of the target object clusters.

12. A computing device comprising a memory configured to store computer-executable instructions; a processor configured to perform the method of any one of claims 1-10 when the computer-executable instructions are executed by the processor.

13. A computer-readable storage medium storing computer-executable instructions that, when executed, perform the method of any one of claims 1-10.

14. A computer program product comprising a computer program, wherein the computer program, when executed by a processor, implements the steps of the method of any one of claims 1-10.

Citation Information

Patent Citations

  • Object recognition method and device, electronic equipment and storage medium

    CN113283480A

  • Generating labeled images

    US9256807B1