Video image annotation method, device, electronic device and computer readable medium

Through continuous frame image annotation and feature fusion, structured labeling information of video images is generated, which solves the problem that it is difficult to extract continuous features in a single image annotation, and achieves the effect of reducing storage resource usage and improving labeling accuracy.

CN118967427BActive Publication Date: 2025-05-16ADDX (BEIJING) TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411010747.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-26
Publication Date
2025-05-16
Estimated Expiration
2044-07-26

AI Technical Summary

Technical Problem

During the video image annotation process, it is difficult for a single image to extract the continuous features of the image, resulting in the labeling not meeting the standards, requiring repeated labeling, occupying computing resources, and resulting in storage redundancy.

Method used

Through continuous frame image annotation, the video annotation data sequence set is extracted, the anchor image is filtered to remove redundancy, feature fusion and associated feature extraction are performed, the video image structured annotation information is generated, and the video image is stored through a hash table.

Benefits of technology

It reduces the usage of storage resources, improves the accuracy of labeling results, reduces storage redundancy, and optimizes the utilization of computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118967427B_ABST
    Figure CN118967427B_ABST
Patent Text Reader

Abstract

The embodiments of the present disclosure disclose a video image annotation method, device, electronic device and computer-readable medium. A specific implementation of the method includes: obtaining a video to be annotated; annotating continuous frame images in the video to be annotated to obtain a video annotation data sequence set; selecting a group of video images corresponding to each annotation target from the continuous frame video images in the video to be annotated as an anchor image group to obtain an anchor image group set; performing feature fusion on each video annotation data sequence in the video annotation data sequence set to generate fused annotation data to obtain a fused annotation data set; extracting associated features from each fused annotation data in the fused annotation data set to generate a feature association description information set; generating video image structured annotation information. This implementation can reduce the occupation of storage resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of computer technology, and in particular to a video image annotation method, device, electronic device, and computer-readable medium. Background Art

[0002] Video image annotation is a technology that adds text feature information that reflects the content of an image. At present, when annotating video images, the usual method is to use a pre-trained machine learning model to annotate the images in the video, and then store the annotation results and the images.

[0003] However, in practice, it is found that when the above method is used to annotate video images, the following technical problems often occur:

[0004] First, if each annotated image and the corresponding result are stored, more storage resources will be required when there are many annotated images;

[0005] Second, it is difficult to extract continuous features of a single image when annotating it, which results in substandard image annotation. As a result, repeated image annotation is required, which takes up computing resources and causes storage redundancy due to the storage of repeated image annotation information.

[0006] The above information disclosed in this Background section is only for enhancement of understanding of the background of the inventive concept and therefore it may contain information that does not form the prior art that is already known in this country to a person of ordinary skill in the art. Summary of the invention

[0007] The content of this disclosure is used to introduce concepts in a brief form, which will be described in detail in the detailed implementation section below. The content of this disclosure is not intended to identify the key features or essential features of the technical solution claimed for protection, nor is it intended to limit the scope of the technical solution claimed for protection.

[0008] Some embodiments of the present disclosure propose a video image annotation method, an apparatus, an electronic device, and a computer-readable medium to solve one or more of the technical problems mentioned in the above background technology section.

[0009] In a first aspect, some embodiments of the present disclosure provide a video image annotation method, the method comprising: obtaining a video to be annotated; annotating continuous frame images of continuous frame video images in the video to be annotated to obtain a video annotation data sequence set, wherein each video annotation data sequence in the video annotation data sequence set corresponds to a labeled target in the video image, and each video annotation data sequence represents the time coordinates, image position coordinates, object attributes and object state of the labeled target in the continuous frame video images included in the video to be annotated; based on the video annotation data sequence set, selecting a group of video images corresponding to each labeled target from the continuous frame video images in the video to be annotated as an anchor image group, to obtain an anchor image group. Image group set; based on the above anchor image group set, perform feature fusion on each video annotation data sequence in the above video annotation data sequence set to generate fused annotation data, and obtain a fused annotation data set; based on each anchor image in the above anchor image group set, perform associated feature extraction on each fused annotation data in the above fused annotation data set to generate a feature associated description information set; generate video image structured annotation information by using the above feature associated description information set, the above anchor image group set and the video annotation data in the above video annotation data sequence set, including temporal coordinates, image position coordinates and object states, and store the video image structured annotation information with mapping relationship through a preset hash table.

[0010] In a second aspect, some embodiments of the present disclosure provide a video image annotation device, which includes: an acquisition unit, configured to acquire a video to be annotated; an image annotation unit, configured to perform continuous frame image annotation on continuous frame video images in the above-mentioned video to be annotated, and obtain a video annotation data sequence set, wherein each video annotation data sequence in the above-mentioned video annotation data sequence set corresponds to a labeled target in the video image, and each video annotation data sequence represents the time coordinates, image position coordinates, object attributes and object state of the labeled target in the continuous frame video images included in the video to be annotated; a selection unit, configured to select a group of video images corresponding to each labeled target from the continuous frame video images in the above-mentioned video to be annotated as anchor points based on the above-mentioned video annotation data sequence set. The image group is obtained by the anchor image group set; the feature fusion unit is configured to perform feature fusion on each video annotation data sequence in the video annotation data sequence set to obtain a fused annotation data set; the feature association unit is configured to extract associated features of each fused annotation data in the fused annotation data set based on each anchor image in the anchor image group set to generate a feature association description information set; the generation and storage unit is configured to generate video image structured annotation information by using the feature association description information set, the anchor image group set and the time coordinates, image position coordinates and object state included in the video annotation data in the video annotation data sequence set, and store the video image structured annotation information with a mapping relationship through a preset hash table.

[0011] In a third aspect, some embodiments of the present disclosure provide an electronic device comprising: one or more processors; a storage device on which one or more programs are stored, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation manner of the above-mentioned first aspect.

[0012] In a fourth aspect, some embodiments of the present disclosure provide a computer-readable medium having a computer program stored thereon, wherein when the program is executed by a processor, the method described in any implementation manner of the above-mentioned first aspect is implemented.

[0013] The above-mentioned embodiments of the present disclosure have the following beneficial effects: through the video image annotation method of some embodiments of the present disclosure, the storage resource occupation can be reduced when the image annotation results are stored. Specifically, the reason for occupying more storage resources is that if each annotated image and the corresponding result are stored. Based on this, the video image annotation method of some embodiments of the present disclosure, first, by annotating continuous frame images, a video annotation data sequence set can be extracted. Then, considering that the storage image and the corresponding annotation results need to occupy more memory, and there are more redundant images in the continuous frame video images. Therefore, by screening the anchor point image, the redundant images can be greatly removed. At the same time, the annotation results corresponding to the continuous frame video images are fused through feature fusion. Thus, the redundant data is further removed. Afterwards, the associated features of the fused annotation data can be extracted to extract the associated features in the continuous frame images. In this way, the accuracy of the annotation results is improved. Finally, the annotation results are structured through the information such as the time coordinates, image position coordinates and object status in the annotation results to generate video image structured annotation information. Compared with the common method of storing a single image and its corresponding result, the continuous frame video images and their corresponding annotation results in the video to be annotated are structured for easy storage. Therefore, the video image structured annotation information with mapping relationship is stored in a preset hash table, which can be used to reduce the occupation of storage resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that components and elements are not necessarily drawn to scale.

[0015] Figure 1 is a flow chart of some embodiments of the video image annotation method according to the present disclosure;

[0016] Figure 2 is a schematic structural diagram of some embodiments of the video image annotation device according to the present disclosure;

[0017] Figure 3 It is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION

[0018] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments set forth herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.

[0019] It should also be noted that, for ease of description, only the parts related to the invention are shown in the drawings. In the absence of conflict, the embodiments and features in the embodiments of the present disclosure can be combined with each other.

[0020] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0021] It should be noted that the modifications of "one" and "plurality" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".

[0022] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0023] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.

[0024] Figure 1 The process 100 of some embodiments of the video image annotation method according to the present disclosure is shown. The video image annotation method comprises the following steps:

[0025] Step 101: Obtain a video to be labeled.

[0026] In some embodiments, the execution subject of the video image annotation method can obtain the video to be annotated by wired or wireless means. The video to be annotated can be a video of any scene shot in advance. For example, it can include but is not limited to at least one of the following: road scene, construction scene, shopping mall scene, etc.

[0027] It should be noted that the above-mentioned wireless connection methods may include but are not limited to 3G / 4G connection, WiFi connection, Bluetooth connection, WiMAX connection, Zigbee connection, UWB (ultra wideband) connection, and other wireless connection methods currently known or to be developed in the future.

[0028] Step 102 , annotate consecutive frames of video images in the video to be annotated to obtain a video annotation data sequence set.

[0029] In some embodiments, the execution subject may annotate the continuous frames of the video images in the video to be annotated to obtain a video annotation data sequence set. Each video annotation data sequence in the video annotation data sequence set corresponds to a labeled target in the video image. Each video annotation data sequence represents the time coordinates, image position coordinates, object attributes and object state of the labeled target in the continuous frames of the video images included in the video to be annotated. Here, the labeled target may be an object or a thing in the image. For example, it may include people, cars, animals, plants, ground, sky, buildings, etc. The time coordinates may be the coordinates of the video image in the video to be annotated. The frame sequence value may be used as the ordinate of the time coordinate on the time axis. Then the abscissa of the time coordinate may be 0. The object attribute may be the attribute of the labeled target. The object state may be the state of the labeled target in the image. For example, the object state may include: (person or animal) walking, (person or animal) standing, (person or animal) running, (person or animal) resting, (vehicle) stopping, (vehicle) moving, (sky) clear, (plant) green, etc.

[0030] In some optional implementations of some embodiments, the execution subject annotates the continuous frame images of the continuous frames of the video to be annotated to obtain a video annotation data sequence set, which may include the following steps:

[0031] The first step is to detect the first frame of video image in the video to be labeled to generate a labeled target group. The first frame of video image in the video to be labeled can be detected by a preset network model to generate a labeled target group.

[0032] As an example, the preset network model may include but is not limited to at least one of the following: YOLO-v3 (You Only Look Once-Version3) algorithm, Faster Rcnn (faster regional proposal networks) fast regional convolutional network DENG.

[0033] Here, for each annotated target in the above annotated target group, the following consecutive frame detection steps are performed:

[0034] Step 1: Starting from the video image where the above-mentioned marked target is located, the continuous frame video images with the same marked target are detected along the frame sequence to obtain a continuous frame image sequence of the marked target. Here, the continuous frame video images with the above-mentioned marked target can be detected starting from the first frame video image through the above-mentioned network model. Finally, the detected continuous frame video images can be determined as a continuous frame image sequence of the marked target. For example, if the marked target appears in the video images of frames 1 to 10, then the video images of frames 1 to 10 are the continuous frame images of the marked target.

[0035] Step 2: The frame number corresponding to each continuous frame image of the marked target in the continuous frame image sequence of the marked target is determined as the time coordinate of the marked target. The time coordinate can be a coordinate on the time axis. Therefore, 0 can be used as the horizontal coordinate value of the time coordinate, and the scale on the time axis, that is, the number of frames, can be used as the vertical coordinate value to obtain the time coordinate. In addition, the vertical coordinate value of the time coordinate can also be set to other values, without specific limitation.

[0036] Step 3: Determine the position of the annotated object in each continuous frame image of the annotated object as the image position coordinates of the annotated object. The image position coordinates are the coordinates of the center point of the annotated object in the image, or the coordinates of the center point of the detection frame of the annotated object in the image, or the coordinates of the border of the detection frame of the annotated object. Here, the detection frame can be a rectangular frame or an irregular detection frame that fits the boundary of the annotated object.

[0037] Step 4: Detect the object attributes and object status of the annotated target in each consecutive frame image of the annotated target. The object attributes and object status of the annotated target in each consecutive frame image of the annotated target can be detected by a preset semantic detection algorithm.

[0038] As an example, the semantic detection algorithm may include but is not limited to at least one of the following: Resnet (Residual Network, residual neural network) model, VGG (Visual Geometry Group Network, convolutional neural network) model and GoogLeNet (deep neural network) model, etc.

[0039] Step 5: The time sequence coordinates, image position coordinates, object attributes and / or object states of the above-mentioned marked target in each continuous frame image of the marked target are determined as video marking data to obtain a video marking data sequence. And it is determined that the detection of the above-mentioned marked target is completed, so as to perform the continuous frame detection step for the next marked target in the above-mentioned marked target group. If the continuous frame images of each marked target corresponding to the marked target have been detected, it is confirmed that the detection of the above-mentioned marked target is completed.

[0040] In the second step, in response to determining that the detection of each labeled target in the first frame of video image is completed, the next frame of video image in the video to be labeled is detected to generate a next frame of labeled target group. And for the labeled targets in the above-mentioned labeled target group that have no association relationship with the labeled targets corresponding to the previous frame of video image, the above-mentioned continuous frame detection step is performed to generate a video annotation data sequence. Among them, the completion of the detection of each labeled target in the first frame of video image means that all areas in the first frame of video image have completed detection, so it is possible to switch to the next frame of video image for detection. In practice, the labeled targets in the labeled target group that have no association relationship with the labeled targets corresponding to the previous frame of video image can indicate that the next frame of video image includes labeled targets that are not in the previous frame of image, so they can be detected again through the above-mentioned continuous frame detection step to generate a video annotation data sequence.

[0041] In the third step, in response to determining that all frames of video images in the video to be labeled have been detected, each video labeling data sequence is combined into a video labeling data sequence set.

[0042] Step 103 : Based on the video annotation data sequence set, a group of video images corresponding to each annotation target is selected from the continuous frame video images in the video to be annotated as an anchor image group to obtain an anchor image group set.

[0043] In some embodiments, the execution subject may select a group of video images corresponding to each annotation target from the continuous frame video images in the video to be annotated as an anchor image group in various ways based on the video annotation data sequence set to obtain an anchor image group set. The anchor image may be an image that meets the conditions in the continuous frame image sequence of the annotation target corresponding to the annotation target. Here, the conditions may include but are not limited to at least one of the following: image size meets, the position of the annotation target in the image meets, image clarity meets, image frame number meets, and other conditions.

[0044] In some optional implementations of some embodiments, the execution subject selects a group of video images corresponding to each annotation target from the continuous frame video images in the video to be annotated as the anchor image group based on the video annotation data sequence set to obtain the anchor image group set, which may include the following steps:

[0045] For each video annotation data sequence corresponding to the annotation target, perform the following steps:

[0046] In the first step, according to the video annotation data sequence corresponding to the above-mentioned annotated target, image masks are established for each frame of video image corresponding to the above-mentioned annotated target to obtain an image mask group. Among them, each image mask uses the area where the annotated target is located in the video image as the area of ​​interest, and other areas as the occlusion area, and the image mask has the same size as the video image. For each corresponding video annotation data and the corresponding continuous frame image of the annotated target with the same frame number, the border coordinates of the detection box in the video annotation data can be used as the mask dividing line to distinguish the annotated target area from other areas. Here, the mask color space of each coordinate in the annotated target area can be set to 0. The mask color of each coordinate in other areas can be set to any color. For example, black or white. In this way, an image mask that conforms to the position of the annotated target in video images with different frame numbers can be obtained.

[0047] In the second step, the image mask group is used to process each frame of the video image corresponding to the annotated target to generate a sequence of processed video images. The image mask can be superimposed on the coordinates of the video image to obtain a processed video image. In this case, due to the mask superposition, only the area where the annotated target is located can be displayed in the processed video image.

[0048] The third step is to determine the clarity and image proportion of each processed video image in the above processed video image sequence to obtain a clarity sequence and an image proportion sequence. The clarity of the processed video image can be determined by an image clarity calculation function. The ratio of the number of coordinates of the area where the marked target is located in the processed video image to the total number of coordinates in the image can be determined as the image proportion. Here, the clarity can be a floating point value between 0 and 1. 1 point clarity can represent the preset highest clarity.

[0049] In the fourth step, the clarity and the image proportion corresponding to the same image in the clarity sequence and the image proportion sequence are weighted to generate an image anchor feature value, thereby obtaining an image anchor feature value sequence. The clarity and the image proportion corresponding to the same image can be weighted to generate an image anchor feature value according to a preset weight value. Here, the image proportion can be a floating point value between 0 and 1.

[0050] As an example, the clarity is 0.6 and the image ratio is 0.3. The corresponding weights are 1 and 2 respectively. Then the anchor feature value can be 1.2.

[0051] In the fifth step, the video image corresponding to the largest image anchor feature value in the above image anchor feature value sequence and the first and last two frames of video images in the continuous frame video images corresponding to the above labeled target are used as anchor images to obtain an anchor image group.

[0052] Here, the video image corresponding to the largest image anchor feature value in the image anchor feature value sequence can represent the image with the most complete features of the above-mentioned labeled target, and thus serve as the anchor image. In addition, the first and last two frames of video images can be used as anchor images to compare the start and end of the labeled target. Or only the video image corresponding to the largest image anchor feature value in the image anchor feature value sequence can be determined as the anchor image.

[0053] Step 104 , based on the anchor image group set, feature fusion is performed on each video annotation data sequence in the video annotation data sequence set to generate fused annotation data, thereby obtaining a fused annotation data set.

[0054] In some embodiments, the execution subject may perform feature fusion on each video annotation data sequence in the video annotation data sequence set based on the anchor image group set in various ways to generate fused annotation data and obtain a fused annotation data set.

[0055] In some optional implementations of some embodiments, the execution subject performs feature fusion on each video annotation data sequence in the video annotation data sequence set based on the anchor image group set to generate fused annotation data, which may include the following steps:

[0056] The first step is to fuse the time coordinates included in each video annotation data in the above video annotation data sequence according to the above anchor image group set to obtain the fused time coordinates. Among them, for the anchor images in the anchor image group corresponding to the above video annotation data sequence, the frame number of the anchor image before each anchor image can be determined as the horizontal coordinate value of the time coordinate in the order of frame number, and the current frame number of the anchor image can be determined as the vertical coordinate value. In this way, the fused time coordinates corresponding to each anchor image are obtained. Therefore, each fused time coordinate can not only carry the time position of the anchor image in the video to be annotated, but also carry the frame number interval characteristics between two adjacent anchor images.

[0057] In addition, the frame number of the first anchor image in each anchor image group can be used as the horizontal coordinate value, and the frame number of the last anchor image in each anchor image group can be used as the vertical coordinate value to obtain the fused temporal coordinates. Thus, the fused temporal coordinates can represent the temporal positioning of an anchor image coordinate group in the video to be annotated.

[0058] In the second step, the image position coordinates included in each video annotation data in the video annotation data sequence correspond to the image position coordinates of the anchor image and are determined as the fused image position coordinates.

[0059] In the third step, the object attributes and object states included in the video annotation data sequence are selected as the fused object attributes and object states. Among them, the number of each object attribute and object state included in the video annotation data can be determined. Here, the object attributes and object states with the largest number can represent the highest confidence.

[0060] The fourth step is to determine the fused time series coordinates, fused image position coordinates, fused object attributes and fused object states corresponding to the above video annotation data sequence as fused annotation data.

[0061] Step 105 , based on each anchor image in the anchor image group set, extract associated features from each fused annotated data in the fused annotated data set to generate a feature associated description information set.

[0062] In some embodiments, the execution entity may extract associated features from each fused annotated data in the fused annotated data set based on each anchor image in the anchor image group set to generate a feature associated description information set.

[0063] In some optional implementations of some embodiments, the execution subject extracts associated features from each fused annotation data in the fused annotation data set based on each anchor image in the anchor image group set to generate a feature association description information set, which may include the following steps:

[0064] For each anchor image and the corresponding annotated object, perform the following association steps:

[0065] In the first step, the fused image position coordinates included in each fused annotation data in the fused annotation data set are used to determine the intersection target of the annotated target in the video image corresponding to the anchor image, and obtain an intersection target group. Among them, each intersection target in the above intersection target group is at a pixel adjacent position with the above annotation target in the image.

[0066] The second step is to determine the association relationship information between the above-mentioned marked target and at least one intersecting target in the above-mentioned intersecting target group based on the image position coordinates, object attributes and object states corresponding to the above-mentioned marked target and at least one intersecting target, so as to generate an association relationship information group. The image position coordinates, object attributes and object states of the above-mentioned marked target and the image position coordinates, object attributes and object states corresponding to at least one intersecting target can be input into a preset scene graph relationship detection algorithm to generate the association relationship information between the above-mentioned marked target and at least one intersecting target in the above-mentioned intersecting target group. The association relationship information can characterize the scene relationship between the marked target and the intersecting target.

[0067] For example, there is a state relationship of "riding" between a person and a car. It can also include structural relationships between annotated objects and intersecting objects, such as a person "on the car". There are relationships such as "reading" and "taking" between a person and a book.

[0068] As an example, the scene graph relationship detection algorithm may include but is not limited to at least one of the following: a knowledge graph embedded Translate model, a Scene Graph Generation model, a Relation Transformer Network relationship transformer model, a Road Scene Graph model, a Human-Object Interaction model, etc.

[0069] The third step is to perform natural language processing on the above association relationship information group to generate feature association description information. The feature association description information represents the image scene relationship between the marked object and at least one intersecting object. Secondly, the feature association description information can be generated by a preset natural language processing algorithm.

[0070] As an example, if the attribute of the labeled target is "person", and the attribute of the intersecting target is "book", and the associated relationship information is "read", then the generated feature association description information may be: "a person is reading a book". For another example, the attribute of the labeled target is "person". The attribute of the intersecting target is "chair". The associated relationship information may be "sit". Then the generated feature association description information may be "a person sitting on a chair" or "a person on chair". In addition, the feature association description of the three labeled targets may be "a person sitting on a chair reading a book".

[0071] As an example, a natural language processing algorithm may include, but is not limited to, at least one of the following: OpenNLP (Open Neuro-Linguistic Programming) natural language generation algorithm, a language model based on HMM (Hidden Markov Model), HMM-LM (Language Model) hidden Markov model, a recurrent neural network (RNN), a CLM (Chinese Language Model) Chinese language model, NLG (Natural Language Generation) natural language generation, etc.

[0072] Step 106, using the time coordinates, image position coordinates and object status included in the video annotation data in the feature association description information set, the anchor image group set and the video annotation data sequence set, generate video image structured annotation information, and store the video image structured annotation information with a mapping relationship through a preset hash table.

[0073] In some embodiments, the execution subject may generate video image structured annotation information using the feature association description information set, the anchor image group set, and the time coordinates, image position coordinates, and object status included in the video annotation data sequence set, and store the video image structured annotation information with mapping relationships through a preset hash table. The hash table may be a hash table that can be used to store mapping relationships between data and nodes.

[0074] In some optional implementations of some embodiments, the execution subject generates video image structured annotation information using the feature association description information set, the anchor image group set, and the temporal coordinates, image position coordinates, and object states included in the video annotation data in the video annotation data sequence set, which may include the following steps:

[0075] The first step is to establish image nodes corresponding to each anchor image group in the above anchor image group set to obtain an image node set. Each image node in the above image node set may include an index of each anchor image in an anchor image group, feature association description information corresponding to each anchor image, and video annotation data of the corresponding annotation target including time series coordinates, image position coordinates, object attributes, and object status.

[0076] The second step is to determine the feature index relationship between the image nodes in the above-mentioned image node set using the associated annotation targets represented by each feature association description information in the above-mentioned feature association description information set, and determine each image node with a mapping relationship as the structured annotation information of the video image. Among them, the feature index relationship between each image node represents the scene association relationship of the annotation target in the image. Here, for each image node, each image node associated with the feature association description information of the node can be determined as a feature index relationship. Thus, the image node can include not only the identifiers of several image nodes with feature association relationships, but also an index pointer pointing to the associated image node. Thus, each image node has a mapping relationship.

[0077] The above steps 102-106 and their related contents, as an inventive point of an embodiment of the present disclosure, solve the second technical problem mentioned in the background technology: "It is difficult to extract the continuous features of the image when annotating a single image, which leads to substandard image annotation. Therefore, it is necessary to repeat image annotation, which takes up computing resources, and also causes storage redundancy due to the storage of repeated image annotation information." The factors that lead to storage redundancy are often as follows: It is difficult to extract the continuous features of the image when annotating a single image, which leads to substandard image annotation. Therefore, it is necessary to repeat image annotation, which takes up computing resources, and also causes storage redundancy due to the storage of repeated image annotation information. In order to achieve the effect of reducing storage redundancy, first, by determining the time coordinates of each annotated target and the continuous frame video images with the annotated target, the features of different images in the video to be annotated can be integrated. Then, by establishing an image mask, it can not only be used to reduce the influence of other regional features on the detection target during image annotation, but also select a suitable anchor image through the processed video image. Therefore, the anchor image can be used as the anchor point of multiple frames of video images, which is convenient for backtracking and storage. At the same time, according to the anchor image, the features corresponding to the same annotation target are fused to reduce feature redundancy. In addition, in order to facilitate structured storage and display of annotation results to users, the association relationship between different annotation targets is established, and the description information that is visualized and easy for users to understand is formed through natural language processing. Finally, by establishing image nodes and determining the feature index relationship between image nodes, the feature association between different annotation targets and different frames of video images is realized, and the purpose of structuring the annotation results is achieved. Therefore, by storing the structured annotation information of video images with mapping relationships in a preset hash table, not only can the storage resource occupation be reduced, but also the storage of repeated image annotation information and images can be reduced, and storage redundancy can be reduced.

[0078] Optionally, the execution subject may also extract corresponding feature association description information from the hash table and send it to the calling end in response to receiving the image annotation call information, wherein the image annotation call information may be a query code for retrieving the anchor image and the image annotation result.

[0079] The above-mentioned embodiments of the present disclosure have the following beneficial effects: through the video image annotation method of some embodiments of the present disclosure, the storage resource occupation can be reduced when the image annotation results are stored. Specifically, the reason for occupying more storage resources is that if each annotated image and the corresponding result are stored. Based on this, the video image annotation method of some embodiments of the present disclosure, first, by annotating continuous frame images, a video annotation data sequence set can be extracted. Then, considering that the storage image and the corresponding annotation results need to occupy more memory, and there are more redundant images in the continuous frame video images. Therefore, by screening the anchor point image, the redundant images can be greatly removed. At the same time, the annotation results corresponding to the continuous frame video images are fused through feature fusion. Thus, the redundant data is further removed. Afterwards, the associated features of the fused annotation data can be extracted to extract the associated features in the continuous frame images. In this way, the accuracy of the annotation results is improved. Finally, the annotation results are structured through the information such as the time coordinates, image position coordinates and object status in the annotation results to generate video image structured annotation information. Compared with the common method of storing a single image and its corresponding result, the continuous frame video images and their corresponding annotation results in the video to be annotated are structured for easy storage. Therefore, the video image structured annotation information with mapping relationship is stored in a preset hash table, which can be used to reduce the occupation of storage resources.

[0080] Further references Figure 2 As an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of a video image annotation device. These device embodiments are similar to Figure 1 Corresponding to the method embodiments shown, the device can be specifically applied to various electronic devices.

[0081] like Figure 2As shown, the video image annotation device 200 of some embodiments includes: an acquisition unit 201, an image annotation unit 202, a selection unit 203, a feature fusion unit 204, a feature association unit 205 and a generation and storage unit 206. The acquisition unit is configured to acquire the video to be annotated; the image annotation unit is configured to annotate the continuous frame video images in the above-mentioned video to be annotated to obtain a video annotation data sequence set, wherein each video annotation data sequence in the above-mentioned video annotation data sequence set corresponds to a labeled target in the video image, and each video annotation data sequence represents the time coordinates, image position coordinates, object attributes and object state of the labeled target in the continuous frame video images included in the video to be annotated; the selection unit is configured to select a group of video images corresponding to each labeled target from the continuous frame video images in the above-mentioned video to be annotated as an anchor image group based on the above-mentioned video annotation data sequence set to obtain an anchor image group set; the feature The fusion unit is configured to perform feature fusion on each video annotation data sequence in the above-mentioned video annotation data sequence set to obtain a fused annotation data set; the feature association unit is configured to extract associated features on each fused annotation data in the above-mentioned fused annotation data set based on each anchor image in the above-mentioned anchor image group set to generate a feature association description information set; the generation and storage unit is configured to generate video image structured annotation information using the above-mentioned feature association description information set, the above-mentioned anchor image group set and the time coordinates, image position coordinates and object status included in the video annotation data in the above-mentioned video annotation data sequence set, and store the video image structured annotation information with a mapping relationship through a preset hash table.

[0082] It is understood that the units described in the device 200 are similar to those described in the reference Figure 1 Therefore, the operations, features and beneficial effects described above for the method are also applicable to the device 200 and the units included therein, and will not be described in detail here.

[0083] Reference below Figure 3 , which shows a structural schematic diagram of an electronic device (eg, a computing device) 300 suitable for implementing some embodiments of the present disclosure. Figure 3 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0084] like Figure 3As shown, the electronic device 300 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory 302 or a program loaded from a storage device 308 into a random access memory 303. Various programs and data required for the operation of the electronic device 300 are also stored in the random access memory 303. The processing device 301, the read-only memory 302, and the random access memory 303 are connected to each other via a bus 304. An input / output interface 305 is also connected to the bus 304.

[0085] Typically, the following devices may be connected to the I / O interface 305: input devices 306 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 308 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 309. The communication devices 309 may allow the electronic device 300 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 3 The electronic device 300 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead. Figure 3 Each block shown in the figure may represent one device, or may represent multiple devices as required.

[0086] In particular, according to some embodiments of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, some embodiments of the present disclosure include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart. In some such embodiments, the computer program can be downloaded and installed from the network through the communication device 309, or installed from the storage device 308, or installed from the read-only memory 302. When the computer program is executed by the processing device 301, the above-mentioned functions defined in the method of some embodiments of the present disclosure are executed.

[0087] It should be noted that the computer-readable medium recorded in some embodiments of the present disclosure may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In some embodiments of the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0088] In some embodiments, the client and the server may communicate using any currently known or future developed network protocol such as HTTP (Hyper Text Transfer Protocol), and may be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.

[0089] The above-mentioned computer-readable medium may be included in the above-mentioned electronic device; or it may exist independently without being assembled into the electronic device. The above-mentioned computer-readable medium carries one or more programs. When the above-mentioned one or more programs are executed by the electronic device, the electronic device: obtains the video to be annotated; annotates the continuous frame images of the continuous frame video images in the above-mentioned video to be annotated to obtain a video annotation data sequence set, wherein each video annotation data sequence in the above-mentioned video annotation data sequence set corresponds to a labeled target in the video image, and each video annotation data sequence represents the time coordinates, image position coordinates, object attributes and object state of the labeled target in the continuous frame video images included in the video to be annotated; based on the above-mentioned video annotation data sequence set, a group of video images corresponding to each labeled target are selected from the continuous frame video images in the above-mentioned video to be annotated as anchor points image group, obtain an anchor image group set; based on the above anchor image group set, perform feature fusion on each video annotation data sequence in the above video annotation data sequence set to generate fused annotation data, and obtain a fused annotation data set; based on each anchor image in the above anchor image group set, perform associated feature extraction on each fused annotation data in the above fused annotation data set to generate a feature associated description information set; generate video image structured annotation information by using the above feature associated description information set, the above anchor image group set and the time series coordinates, image position coordinates and object state included in the video annotation data in the video annotation data sequence set, and store the video image structured annotation information with a mapping relationship through a preset hash table.

[0090] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0091] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0092] The units described in some embodiments of the present disclosure may be implemented by software or by hardware. The units described may also be provided in a processor, for example, may be described as: a processor including an acquisition unit, an image annotation unit, a selection unit, a feature fusion unit, a feature association unit, and a generation and storage unit. The names of these units do not, in some cases, constitute limitations on the units themselves, for example, the acquisition unit may also be described as a "unit for acquiring a video to be annotated".

[0093] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.

[0094] The above descriptions are only some preferred embodiments of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by a specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the above-mentioned inventive concept. For example, the above-mentioned features are replaced with the technical features with similar functions disclosed in the embodiments of the present disclosure (but not limited to) and the technical solutions formed.

Claims

1. A video image annotation method, comprising: Get the video to be annotated; Performing continuous frame image annotation on continuous frame video images in the video to be annotated to obtain a video annotation data sequence set, wherein each video annotation data sequence in the video annotation data sequence set corresponds to a labeled target in the video image, and each video annotation data sequence represents the time sequence coordinates, image position coordinates, object attributes and object state of the labeled target in the continuous frame video images included in the video to be annotated; Based on the video annotation data sequence set, a group of video images corresponding to each annotation target is selected from the continuous frame video images in the video to be annotated as an anchor image group to obtain an anchor image group set; Based on the anchor image group set, feature fusion is performed on each video annotation data sequence in the video annotation data sequence set to generate fused annotation data, thereby obtaining a fused annotation data set; Based on each anchor image in the anchor image group set, extract associated features from each fused annotated data in the fused annotated data set to generate a feature associated description information set; Generate video image structured annotation information by using the feature association description information set, the anchor image group set and the time coordinates, image position coordinates and object status included in the video annotation data sequence set, and store the video image structured annotation information with a mapping relationship through a preset hash table; Wherein, based on each anchor image in the anchor image group set, extracting associated features from each fused annotated data in the fused annotated data set to generate a feature associated description information set, including: For each anchor image and the corresponding annotated object, perform the following association steps: Using the fused image position coordinates included in each fused annotation data in the fused annotation data set, the intersection target of the annotated target in the video image corresponding to the anchor point image is determined to obtain an intersection target group; Determine association relationship information between the marked object and at least one intersecting object in the intersecting object group based on image position coordinates, object attributes, and object states corresponding to the marked object and at least one intersecting object, so as to generate an association relationship information group; The association relationship information group is subjected to natural language processing to generate feature association description information, wherein the feature association description information represents an image scene relationship between the marked object and at least one intersecting object.

2. The method according to claim 1, wherein: The method further comprises: In response to receiving the image annotation call information, the corresponding feature association description information is extracted from the hash table and sent to the caller.

3. The method according to claim 1, wherein: The step of annotating the continuous frames of video images in the video to be annotated to obtain a video annotation data sequence set includes: Detect the first frame of the video image in the video to be labeled to generate a labeled object group, and for each labeled object in the labeled object group, perform the following continuous frame detection steps: Starting from the video image where the marked target is located, continuous frame video images with the same marked target are detected along the frame sequence to obtain a sequence of continuous frame images of the marked target; Determine the frame number corresponding to each of the continuous frame images of the marked target in the continuous frame image sequence of the marked target as the temporal coordinate of the marked target; Determine the position of the marked object in each continuous frame image of the marked object as the image position coordinates of the marked object, wherein the image position coordinates are the center point coordinates of the marked object in the image, or the center point coordinates of the detection frame of the marked object in the image, or the frame coordinates of the detection frame of the marked object; Detecting the object attributes and object states of the marked object in each continuous frame image of the marked object; Determine the time sequence coordinates, image position coordinates, object attributes and / or object states of the marked object corresponding to each continuous frame image of the marked object as video marking data, obtain a video marking data sequence, and determine that the marked object detection is completed, so as to perform a continuous frame detection step on the next marked object in the marked object group; In response to determining that the detection of each labeled target in the first frame of video image is completed, detecting the next frame of video image in the video to be labeled to generate a next frame of labeled target group, and for a labeled target in the labeled target group that has no association relationship with a labeled target corresponding to the previous frame of video image, performing the continuous frame detection step to generate a video labeling data sequence; In response to determining that all frames of video images in the video to be labeled have been detected, each video labeling data sequence is combined into a video labeling data sequence set.

4. The method according to claim 1, wherein: The step of selecting a group of video images corresponding to each annotation target from the continuous frame video images in the video to be annotated as an anchor image group based on the video annotation data sequence set to obtain an anchor image group set includes: For each video annotation data sequence corresponding to the annotation target, perform the following steps: According to the video annotation data sequence corresponding to the annotated target, image masks are respectively established for each frame of the video image corresponding to the annotated target to obtain an image mask group, wherein each image mask uses the area where the annotated target is located in the video image as the area of ​​interest and other areas as the occlusion area, and the image mask has the same size as the video image; Using the image mask group, performing image processing on each frame of video image corresponding to the marked target to generate a processed video image sequence; Determine the definition and image proportion of each processed video image in the processed video image sequence to obtain a definition sequence and an image proportion sequence; Performing weighted processing on the clarity and image occupancy ratio corresponding to the same image in the clarity sequence and the image occupancy ratio sequence to generate image anchor point feature values, thereby obtaining an image anchor point feature value sequence; The video image corresponding to the largest image anchor feature value in the image anchor feature value sequence and the first and last two frames of video images in the continuous frames of video images corresponding to the marked target are taken as anchor images to obtain an anchor image group.

5. A video image annotation device, comprising: An acquisition unit, configured to acquire a video to be annotated; An image annotation unit is configured to perform continuous frame image annotation on continuous frame video images in the video to be annotated, and obtain a video annotation data sequence set, wherein each video annotation data sequence in the video annotation data sequence set corresponds to a labeled target in the video image, and each video annotation data sequence represents the time sequence coordinates, image position coordinates, object attributes and object state of the labeled target in the continuous frame video images included in the video to be annotated; A selection unit is configured to select a group of video images corresponding to each annotation target from the continuous frame video images in the video to be annotated as an anchor image group based on the video annotation data sequence set, to obtain an anchor image group set; A feature fusion unit is configured to perform feature fusion on each video annotation data sequence in the video annotation data sequence set to obtain a fused annotation data set; A feature association unit is configured to extract associated features from each fused annotation data in the fused annotation data set based on each anchor image in the anchor image group set, so as to generate a feature association description information set; A generating and storing unit is configured to generate video image structured annotation information by using the feature association description information set, the anchor image group set and the time sequence coordinates, image position coordinates and object state included in the video annotation data in the video annotation data sequence set, and store the video image structured annotation information with a mapping relationship through a preset hash table; Wherein, based on each anchor image in the anchor image group set, extracting associated features from each fused annotated data in the fused annotated data set to generate a feature associated description information set, including: For each anchor image and the corresponding annotated object, perform the following association steps: Using the fused image position coordinates included in each fused annotation data in the fused annotation data set, the intersection target of the annotated target in the video image corresponding to the anchor point image is determined to obtain an intersection target group; Determine association relationship information between the marked object and at least one intersecting object in the intersecting object group based on image position coordinates, object attributes, and object states corresponding to the marked object and at least one intersecting object, so as to generate an association relationship information group; The association relationship information group is subjected to natural language processing to generate feature association description information, wherein the feature association description information represents an image scene relationship between the marked object and at least one intersecting object.

6. An electronic device comprising: one or more processors; a storage device having one or more programs stored thereon, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 4.

7. A computer readable medium having a computer program stored thereon, wherein: When the computer program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.

Citation Information

Patent Citations

  • Real-time continuous frame embedded information recognition system

    CN108833964A

  • Video stream processing method and device, computer equipment and storage medium

    CN111898416A