A method and device for determining different subordinate relationships in a video stream

Through the methods of ‘continuous frame guessing’ and ‘interval frame verification’, the problem that traditional image recognition technology cannot recognize always-accompanies animals is solved, and the efficiency and intelligence of computer processors are improved, especially in the fields of road traffic and autonomous driving.

CN116129313BActive Publication Date: 2025-08-12DONGFENG COMML VEHICLE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310052361.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-02
Publication Date
2025-08-12
Estimated Expiration
2043-02-02

AI Technical Summary

Technical Problem

When traditional image recognition technology processes video streams, it is impossible to effectively identify objects that always follow each other as a whole, resulting in inefficiency of computer processors.

Method used

Using the methods of ‘continuous frame guessing’ and ‘interval frame verification’, through the two marking processes, we identify and confirm that the objects that always follow each other are the same target.

Benefits of technology

Reduce the number of targets to focus and improve the efficiency and speed of computer processors when dealing with targets to focus, especially in the fields of road traffic and autonomous driving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116129313B_ABST
    Figure CN116129313B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and device for determining different subordinate relationships in a video stream, relating to the field of artificial intelligence target recognition technology, wherein the method for determining different subordinate relationships in a video stream comprises: identifying candidate regions of all objects in the first frame of the video; based on the identified candidate regions, identifying clustered candidate region groups; based on the candidate region groups in the first frame, screening out candidate region groups that remain unchanged in consecutive preset frames after the first frame, and pre-marking them as the same target object; obtaining video image frames spaced apart after the consecutive preset frames, and based on the obtained spaced video image frames, if the pre-marked candidate region groups are always clustered together in the spaced video frames, then the candidate region groups are finally marked as the same target object. The present invention adopts the method of "continuous frame guessing" and "spaced frame verification", and through two markings, ultimately determines two or more objects that always follow each other as one target, thereby reducing the number of targets of interest, which is conducive to improving the efficiency and speed of the computer processor when processing the targets of interest, making it more efficient and intelligent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence target recognition, and in particular to a method and device for determining different subordinate relationships in a video stream. Background Art

[0002] Traditional image recognition technology can identify every person or object in a video. However, when the video is complex, traditional image recognition technology will identify each person or object in the video as a target, which will lead to an excessive number of targets of interest, resulting in low efficiency and slow speed in the subsequent processing of these targets by the computer processor.

[0003] In the fields of road traffic and autonomous driving, when two or more objects are always moving with each other, they can be judged as a whole, a single traffic target object. However, traditional image recognition technology will identify multiple objects that are always moving with each other as multiple targets, which is obviously not smart enough. Summary of the Invention

[0004] In response to the defects in the prior art, the purpose of the present invention is to provide a method and device for determining different subordinate relationships in a video stream. The method adopts the "continuous frame guessing" and "interval frame verification" methods, and through two markings, two or more objects that always move with each other are finally determined to be one target, which reduces the number of targets of interest and is conducive to improving the efficiency and speed of computer processors when processing targets of interest, making them more efficient and intelligent.

[0005] To achieve the above object, the present invention provides a method for determining different subordination relationships in a video stream, comprising the following steps:

[0006] Identify candidate regions of all objects in the first frame of the video;

[0007] Based on the identified candidate regions, a cluster of candidate regions is identified;

[0008] Based on the candidate region group in the first frame, the candidate region group that remains unchanged in consecutive preset frames after the first frame is screened out and pre-marked as the same target object;

[0009] Get the interval video image frames after the consecutive preset frames, and according to the obtained interval video image frames:

[0010] If the pre-marked candidate region groups are always clustered together in the interval video frames, the candidate region groups are finally marked as the same target object.

[0011] On the basis of the above technical solution, the candidate area in the first frame of the video is identified, wherein the first frame is the video image frame where the object of interest appears for the first time in the video.

[0012] On the basis of the above technical solution, the candidate regions of all objects in the first frame of the video are identified, wherein the candidate regions are regions where all objects in the video picture are identified by the region proposal technology and marked with a box.

[0013] On the basis of the above technical solution, after identifying the candidate areas of all objects in the first frame of the video, the method further includes: setting the duration of the video recognition sample.

[0014] On the basis of the above technical solution, the method of screening out the candidate region groups that remain unchanged in consecutive preset frames after the first frame based on the candidate region groups in the first frame and pre-marking them as the same target object specifically includes the following steps:

[0015] Get all candidate regions in the first frame of the video;

[0016] Based on the obtained candidate regions, a cluster of candidate regions is identified and recorded as a first candidate region group;

[0017] identifying candidate region groups clustered together in each of the consecutive preset frames and recording them as second candidate region groups;

[0018] Based on the identified second candidate region group, candidate region groups that remain unchanged from the first candidate region group are screened out and pre-marked as the same target object.

[0019] On the basis of the above technical solution, the steps of obtaining the interval video image frames and finally marking the same target object if the pre-marked candidate region groups are always clustered together in the interval video frames include:

[0020] Identify candidate region groups clustered together in the video image frames within the sample duration and after the preset frame;

[0021] According to the candidate region group obtained by identification:

[0022] If the pre-labeled candidate regions are always clustered together in the interval frames, they are ultimately labeled as the same object.

[0023] On the basis of the above technical solution, the video image frames spaced after the preset frames are acquired, wherein the number of frames spaced between the spaced frames is the same.

[0024] The present invention also provides a device for determining different subordination relationships in a video stream, comprising:

[0025] A recognition module, which is used to identify candidate regions of all objects in the first frame of the video, and to identify candidate region groups clustered together in the first frame;

[0026] An acquisition module, which is used to acquire a video image frame after a preset frame;

[0027] An execution module is used to screen out candidate region groups that remain unchanged in consecutive preset frames after the first frame based on the candidate region groups in the first frame identified by the recognition module, and pre-mark them as the same target object; according to the interval video image frames obtained by the acquisition module: if the pre-marked candidate region groups are always clustered together in the interval video frames, they are finally marked as the same target object.

[0028] On the basis of the above technical solution, the candidate area in the first frame of the video is identified, wherein the first frame is the video image frame where the object of interest appears for the first time in the video.

[0029] On the basis of the above technical solution, the candidate regions of all objects in the first frame of the video are identified, wherein the candidate regions are regions where all objects in the video picture are identified by the region proposal technology and marked with a box.

[0030] Compared with existing technologies, the present invention offers advantages in that it employs a "continuous frame guessing" and "interval frame verification" approach, using two labeling steps to ultimately identify two or more objects that consistently follow each other as a single target. This reduces the number of targets of interest and improves the efficiency and speed of computer processors when processing these targets. In the fields of road traffic and autonomous driving, when two or more objects consistently follow each other, they are identified as a single entity, a more efficient and intelligent method. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0032] Figure 1 The figure is a flow chart of a method for determining different subordinate relationships in a video stream according to an embodiment of the present invention. DETAILED DESCRIPTION

[0033] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments.

[0034] See also Figure 1 As shown, an embodiment of the present invention provides a method for determining different subordination relationships in a video stream, comprising the following steps:

[0035] S1: Identify candidate regions of all objects in the first frame of the video;

[0036] S2: Based on the identified candidate regions, identify clustered candidate regions;

[0037] S3: Based on the candidate region group in the first frame, the candidate region group that remains unchanged in consecutive preset frames after the first frame is screened out and pre-marked as the same target object;

[0038] S4: acquiring interval video image frames after consecutive preset frames, and based on the acquired interval video image frames: if the pre-marked candidate region groups are always clustered together in the interval video frames, the candidate region groups are finally marked as the same target object.

[0039] After the video is imported into the computer, the computer recognizes all objects in the first frame of the video and marks the recognized objects with boxes. The area marked by the box is the candidate area, and each area contains and only contains a single potential target object; for example, there is a person, a horse and an iron gate in the first frame of the video. The computer can recognize three candidate areas in the first frame, which respectively identify the person, the horse and the iron gate.

[0040] After the computer identifies the candidate areas corresponding to all objects in the first frame of the video, it calculates the Euclidean distance between the centers of each area, uses the Euclidean distance as the distance type, and executes the Kmeans clustering algorithm to identify two or more candidate areas that are clustered together. Two or more candidate areas that are clustered together are called a candidate area group. Among them, clustered candidate areas are two or more adjacent or overlapping candidate areas. For example, in the first frame of the video, a person is riding a horse next to an iron gate. At this time, the candidate areas where the person and the horse are located overlap, and the candidate area where the iron gate is located is adjacent to the candidate area where the person is located. In this case, the computer will identify these three candidate areas as a cluster, that is, identify these three adjacent or overlapping candidate areas as a clustered candidate area group.

[0041] Based on the identified candidate region groups, the computer screens out candidate region groups that remain unchanged in consecutive preset frames after the first frame of the sample video; wherein the unchanged candidate region groups are clustered together in the consecutive preset frames, and their positional relationship with the identified candidate region groups remains unchanged. The screened unchanged candidate region groups are then pre-marked as the same target object. For example, after the computer identifies the candidate regions of the three objects, a person, a horse, and a metal door, as a clustered candidate region group in the first frame of the video, the computer screens out candidate region groups that remain unchanged in consecutive preset frames after the first frame of the sample video. In the consecutive preset frames, the person and the horse move simultaneously, moving with each other and gradually moving away from the immovable metal door. Therefore, the candidate regions of the person and the horse overlap in the consecutive preset frames. The candidate region of the metal door gradually shifts from being adjacent to the candidate region of the person, and at this point, the candidate region of the metal door is no longer clustered with the candidate regions of the person and the horse. The computer screens out the candidate region groups that remain unchanged in the consecutive preset frames, which are formed by the overlap of the candidate regions of the person and the horse, and pre-marks this candidate region group as the same target object.

[0042] After the computer pre-marks the same target object, it obtains the video image frames that are spaced apart after consecutive preset frames in the sample video. Based on the obtained spaced video image frames, it makes a judgment. If the candidate region group pre-marked in the previous step is clustered together in each frame of the spaced frame, then the candidate region group is ultimately marked as the same target object. For example, after the computer has pre-marked the candidate region group containing a person and a horse as the same target object in the previous step, it obtains the video image frames that are spaced apart after consecutive preset frames to determine whether the candidate region group containing the person and the horse is always clustered together in the spaced video frames. If so, the candidate region group containing the person and the horse is ultimately marked as the same target object, that is, the person and the horse are always moving with each other in this video segment, equivalent to being a whole. In the fields of road traffic and autonomous driving, when two or more objects are always moving with each other, it can be determined that they are a whole, a single traffic target object.

[0043] By simply looking at a single frame image, it is impossible to determine the tracking relationship between various objects. The present invention adopts a two-step method of "continuous frame guessing, interval frame verification, and double labeling" to identify groups of objects that move together and determine whether they are a whole object. First, the candidate region groups that are clustered together in the first frame of the video are identified. Then, based on the identified candidate region groups, the candidate region groups that remain unchanged in the consecutive preset frames after the first frame are screened and pre-marked as the same target object. The screened candidate region groups are verified in the interval frames after the preset frame to verify whether the candidate region groups are always clustered together in the interval frames. If the candidate region groups are always clustered together in the interval frames, the candidate region groups are finally marked as the same target object. The present invention takes into account the angle change and occlusion effects caused by the movement of the object relative to the camera. Only continuous video segments less than a specified threshold are used as the basis for identification each time, such as 3s or 5s, depending on the frame rate of the video. In such a short time, the object's angle change and occlusion effects can be approximately considered unchanged.

[0044] In the present invention, the candidate area in the first frame of the video is identified, wherein the first frame is the video image frame where the object of interest appears for the first time in the video.

[0045] That is, the computer recognizes that the first frame referred to by the candidate area in the first frame of the video is the video image frame where the object of interest appears in the video. For example, the initial picture in the video is blank without any object, until an iron door, a person or a horse appears in the video picture. The video image frame at this time is the first frame that the computer needs to process, that is, the first frame of the sample video.

[0046] In the present invention, the candidate regions of all objects in the first frame of the video are identified, wherein the candidate regions are regions where all objects in the video picture are identified by the region proposal technology and marked with a box.

[0047] That is, when the computer recognizes objects in video image frames, it first reuses existing region proposal technologies, such as R-CNN (Region-based Convolutional Neural Networks), to identify the regions of all objects. Each region contains and only contains a single potential target object. For example, in the first frame of the video, a person, a horse, and an iron door are identified, and the regions where these three objects are located are marked with boxes. Each region contains and only contains one object.

[0048] In the present invention, after identifying candidate areas of all objects in the first frame of the video, the method further includes: setting the duration of the video recognition sample.

[0049] That is, after the computer identifies the candidate areas of all objects in the first frame of the video, it also sets the duration of the video recognition sample. In this embodiment, the duration of the video recognition sample is set to 3 seconds.

[0050] In the present invention, based on the candidate region group in the first frame, the candidate region group that remains unchanged in consecutive preset frames after the first frame is screened out and pre-marked as the same target object. The specific steps include:

[0051] S301: Obtain all candidate regions in the first frame of the video;

[0052] S302: Based on the obtained candidate regions, identify a cluster of candidate regions, and record them as a first candidate region group;

[0053] S303: identifying candidate region groups clustered together in each of the consecutive preset frames, and recording them as second candidate region groups;

[0054] S304: Based on the identified second candidate region group, select candidate region groups that remain unchanged from the first candidate region group and pre-mark them as the same target object.

[0055] That is, based on the candidate region group in the first frame, the candidate region group that remains unchanged in consecutive preset frames after the first frame is screened out and pre-marked as the same target object. The specific step is that the computer obtains all candidate regions in the first frame of the video. In this embodiment, the first frame of the video contains a person, a horse, and a metal gate. The computer obtains the candidate regions where the person, horse, and metal gate are located. Based on the obtained candidate regions, the computer identifies clustered candidate region groups. In this embodiment, in the first frame of the video, the person is riding a horse next to the metal gate. At this time, the candidate regions where the person and horse are located overlap, and the candidate region where the metal gate is located is adjacent to the candidate region where the person is located. In this case, the computer will identify these three candidate regions as a cluster, that is, identify these three adjacent or overlapping candidate regions as a clustered candidate region group, and the computer records the candidate region group where the person, horse, and metal gate are located as the first candidate region group. The computer then identifies clustered candidate area groups in each of the consecutive preset frames and records them as second candidate area groups. In this embodiment, the consecutive preset frames are the 19 consecutive frames after the first frame of the video. The computer identifies clustered candidate area groups in each of these 19 frames and records these candidate area groups as second candidate area groups. The computer compares the second candidate area group with the first candidate area group and selects candidate area groups in the second candidate area group that remain unchanged compared to the first candidate area group. In this embodiment, the candidate areas for the person and the horse in the second candidate area group always overlap in the consecutive preset frames, and the candidate area for the iron door gradually moves away from the candidate area for the person and the horse. The candidate area group that remains unchanged selected by the computer is the candidate area group for the person and the horse, and the candidate area group for the person and the horse is pre-marked as the same target object.

[0056] In the present invention, if the pre-marked candidate region groups are always clustered together in the interval video frames obtained, they are finally marked as the same target object. The specific steps include:

[0057] S401: Identify candidate region groups clustered together in video image frames within a sample duration and after a preset frame;

[0058] S402: Based on the identified candidate region groups, if the pre-marked candidate region groups are always clustered together in the interval frames, go to S403;

[0059] S403: The objects are finally marked as the same.

[0060] Specifically, based on the acquired interval video frames, if the pre-marked candidate region groups consistently cluster together in the interval video frames, they are ultimately labeled as the same object. Specifically, the computer first identifies candidate region groups that cluster together in the interval video frames after a preset frame within the sample duration. In this embodiment, the sample duration is 3 seconds, the video frame rate is 60 frames per second, and there are 180 frames in total. The preset frames are the 19 consecutive frames after the first frame. That is, the computer identifies candidate region groups that cluster together in the last 120 frames of the sample video at regular intervals. Based on the identified candidate region groups, the computer then determines if the pre-marked candidate region groups consistently cluster together in the last 120 frames of the video, and ultimately labels them as the same object. In this embodiment, the computer extracts three interval frames from the last 120 frames of the video. If the candidate region groups containing the person and the horse consistently cluster together in these three interval frames, the computer ultimately labels the candidate region groups containing the person and the horse as the same object.

[0061] In the present invention, the video image frames spaced apart after the preset frames are acquired, wherein the number of frames spaced apart between the spaced frames is the same.

[0062] That is, the computer obtains the interval video frames in the last 120 frames of the sample video. The number of frames between the interval frames is the same, and the interval between the interval frames is 39 frames. That is, the computer obtains one frame every 40 frames in the last 120 frames of the sample video.

[0063] In one possible implementation, an embodiment of the present invention further provides a readable storage medium, which is located in a PLC (Programmable Logic Controller) controller. The readable storage medium stores a computer program, which, when executed by a processor, implements the following steps of a method for determining different subordination relationships in a video stream:

[0064] Identify candidate regions of all objects in the first frame of the video;

[0065] Based on the identified candidate regions, a cluster of candidate regions is identified;

[0066] Based on the candidate region group in the first frame, the candidate region group that remains unchanged in consecutive preset frames after the first frame is screened out and pre-marked as the same target object;

[0067] Get the interval video image frames after the consecutive preset frames, and according to the obtained interval video image frames:

[0068] If the pre-marked candidate region groups are always clustered together in the interval video frames, the candidate region groups are finally marked as the same target object.

[0069] The method of screening out candidate region groups that remain unchanged in consecutive preset frames after the first frame based on the candidate region groups in the first frame and pre-marking them as the same target object specifically includes the following steps:

[0070] Get all candidate regions in the first frame of the video;

[0071] Based on the obtained candidate regions, a cluster of candidate regions is identified and recorded as a first candidate region group;

[0072] identifying candidate region groups clustered together in each of the consecutive preset frames and recording them as second candidate region groups;

[0073] Based on the identified second candidate region group, candidate region groups that remain unchanged from the first candidate region group are screened out and pre-marked as the same target object.

[0074] According to the obtained interval video image frames, if the pre-marked candidate region groups are always clustered together in the interval video frames, they are finally marked as the same target object. The specific steps include:

[0075] Identify candidate region groups clustered together in the video image frames within the sample duration and after the preset frame;

[0076] According to the candidate region groups obtained by recognition: if the pre-marked candidate region groups are always clustered together in the interval frames, they are finally marked as the same target object.

[0077] The storage medium can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to: an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination of the above. More specific examples (non-exhaustive list) of computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or device.

[0078] A computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which may transmit, propagate, or transfer a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wire, optical cable, RF, etc., or any suitable combination thereof.

[0079] Computer program code for performing the operations of the present invention can be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0080] The present invention also provides a device for determining different subordination relationships in a video stream, comprising:

[0081] A recognition module, which is used to identify candidate regions of all objects in the first frame of the video, and to identify candidate region groups clustered together in the first frame;

[0082] An acquisition module, which is used to acquire a video image frame after a preset frame;

[0083] An execution module is used to screen out candidate region groups that remain unchanged in consecutive preset frames after the first frame based on the candidate region groups in the first frame identified by the recognition module, and pre-mark them as the same target object; according to the interval video image frames obtained by the acquisition module: if the pre-marked candidate region groups are always clustered together in the interval video frames, they are finally marked as the same target object.

[0084] That is, the device for determining different subordinate relationships in a video stream provided by this embodiment includes a recognition module, an acquisition module, and an execution module. The recognition module is used to identify the candidate areas where all objects in the first frame of the sample video are located, and to identify the candidate area groups clustered together in the first frame of the sample video. In this embodiment, there is a person, a horse, and an iron gate in the first frame of the video. The recognition module identifies the candidate areas where the person, horse, and iron gate are located. The clustered candidate areas are two or more adjacent or overlapping candidate areas. In this embodiment, in the first frame of the video, a person is riding a horse next to the iron gate. At this time, the candidate areas where the person and horse are located are overlapping, and the candidate area where the iron gate is located is adjacent to the candidate area where the person is located. In this case, the recognition module will identify these three candidate areas as a cluster, that is, identify these three adjacent or overlapping candidate areas as a candidate area group clustered together. The acquisition module is used to obtain 120 frames of video image frames after the preset frame of the sample video. The execution module is used to screen out the candidate region groups that remain unchanged in 19 consecutive frames after the first frame of the sample video based on the candidate region groups identified by the recognition module, and pre-mark the candidate region groups as the same target object; and according to the interval video image frames obtained by the acquisition module, determine whether the pre-marked candidate region groups are always clustered together in the interval video image frames. If the candidate region groups are always clustered together in the interval video image frames, the candidate region groups are finally marked as the same target object.

[0085] In the present invention, the candidate area in the first frame of the video is identified, wherein the first frame is the video image frame where the object of interest appears for the first time in the video.

[0086] That is, the recognition module recognizes the candidate area of the first frame of the sample video, where the first frame is the video image frame where the object of interest appears for the first time in the video.

[0087] In the present invention, the candidate regions of all objects in the first frame of the video are identified, wherein the candidate regions are regions where all objects in the video picture are identified by the region proposal technology and marked with a box.

[0088] That is, the recognition module identifies the candidate areas of all objects in the first frame of the sample video. The candidate areas are the practical area suggestion technology that identifies all objects in the first frame and marks the area where each object is located with a box. The area marked with a box is the candidate area.

[0089] The foregoing is merely a list of specific embodiments of the present application, intended to enable those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the broadest scope consistent with the principles and novel features of the present application.

[0090] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems) and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

Claims

1. A method for determining different affiliations in a video stream, characterized in that The following steps are involved: Identify candidate regions of all objects in the first frame of the video; Based on the identified candidate regions, a cluster of candidate regions is identified; Based on the candidate region group in the first frame, the candidate region group that remains unchanged in consecutive preset frames after the first frame is screened out and pre-marked as the same target object; Get the interval video image frames after the consecutive preset frames, and according to the obtained interval video image frames: If the pre-marked candidate region groups are always clustered together in the interval video frames, the candidate region groups are finally marked as the same target object; After the computer identifies the candidate regions corresponding to all objects in the first frame of the video, it calculates the Euclidean distance between the centers of each region and uses the Euclidean distance as the distance type to perform a Kmeans clustering algorithm to identify two or more candidate regions that are clustered together. These two or more candidate regions that are clustered together are called candidate region clusters. The candidate region group includes a plurality of objects.

2. A method for determining different subordination relationships in a video stream according to claim 1, characterized in that: The candidate area in the first frame of the video is identified, wherein the first frame is the video image frame where the object of interest appears for the first time in the video.

3. A method for determining different subordination relationships in a video stream according to claim 2, characterized in that: The candidate regions of all objects in the first frame of the video are identified, wherein the candidate regions are regions where all objects in the video picture are identified by region proposal technology and marked with boxes.

4. The method for determining different subordination relationships in a video stream according to claim 1, wherein: After identifying candidate regions of all objects in the first frame of the video, the method further includes: setting the duration of the video recognition sample.

5. A method for determining different subordination relationships in a video stream according to claim 4, characterized in that: The method of screening out candidate region groups that remain unchanged in consecutive preset frames after the first frame based on the candidate region groups in the first frame and pre-marking them as the same target object specifically includes the following steps: Get all candidate regions in the first frame of the video; Based on the obtained candidate regions, a cluster of candidate regions is identified and recorded as a first candidate region group; identifying candidate region groups clustered together in each of the consecutive preset frames and recording them as second candidate region groups; Based on the identified second candidate region group, candidate region groups that remain unchanged from the first candidate region group are screened out and pre-marked as the same target object.

6. The method for determining different subordination relationships in a video stream according to claim 1, wherein: According to the obtained interval video image frames, if the pre-marked candidate region groups are always clustered together in the interval video frames, they are finally marked as the same target object. The specific steps include: Identify candidate region groups clustered together in the video image frames within the sample duration and after the preset frame; According to the candidate region group obtained by identification: If the pre-labeled candidate regions are always clustered together in the interval frames, they are ultimately labeled as the same object.

7. The method for determining different affiliations in a video stream according to claim 1, wherein: The video image frames spaced apart after the preset frames are acquired, wherein the number of frames spaced apart between the spaced apart frames is the same.

8. A device for determining different affiliations in a video stream, characterized in that include: A recognition module, which is used to identify candidate regions of all objects in the first frame of the video, and to identify candidate region groups clustered together in the first frame; An acquisition module, which is used to acquire a video image frame after a preset frame; an execution module configured to, based on the candidate region groups in the first frame identified by the recognition module, screen out candidate region groups that remain unchanged in consecutive preset frames after the first frame, and pre-mark them as the same target object; and, based on the interval video image frames acquired by the acquisition module, if the pre-marked candidate region groups are always clustered together in the interval video frames, then they are ultimately marked as the same target object; After the computer identifies the candidate regions corresponding to all objects in the first frame of the video, it calculates the Euclidean distance between the centers of each region and uses the Euclidean distance as the distance type to perform a Kmeans clustering algorithm to identify two or more candidate regions that are clustered together. These two or more candidate regions that are clustered together are called candidate region clusters. The candidate region group includes a plurality of objects.

9. The apparatus for determining different subordination relationships in a video stream according to claim 8, wherein: The candidate area in the first frame of the video is identified, wherein the first frame is the video image frame where the object of interest appears for the first time in the video.

10. The apparatus for determining different subordination relationships in a video stream according to claim 8, wherein: The candidate regions of all objects in the first frame of the video are identified, wherein the candidate regions are regions where all objects in the video picture are identified by region proposal technology and marked with boxes.

Citation Information

Patent Citations

  • Target tracking method and system

    CN107832683A

  • Smart city traffic diversion management method, system and device and medium

    CN115271543A