Multimedia data annotation model generation method and multimedia data annotation method

Through the joint training of the target recognition and segmentation sub-models of the multimedia data annotation model, the problems of low efficiency in finding target information and poor segmentation accuracy in videos are solved, and efficient and accurate target object recognition and segmentation are achieved, reducing training costs and search time.

CN120673413APending Publication Date: 2025-09-19SINOVATION (BEIJING) MEDICAL TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510767922.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing technologies are inefficient in finding target information and segmenting target objects in videos, with poor segmentation accuracy. In addition, target recognition and target segmentation models do not fully utilize correlations, resulting in limited model performance.

Method used

By generating a multimedia data annotation model, including joint training of the target recognition sub-model and the target segmentation sub-model, local images are extracted using the target object annotation box for training, combined with pre-trained DETR and SAM models for fine-tuning, and the loss function is optimized to improve model performance.

Benefits of technology

It improves the recognition accuracy and pixel-level segmentation and positioning accuracy of target objects, reduces the cost of model training, speeds up model convergence, and saves time in finding target objects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673413A_ABST
    Figure CN120673413A_ABST
Patent Text Reader

Abstract

The invention provides a multimedia data annotation model generation method and a multimedia data annotation method, and the method comprises the steps: annotating target object pixels for target scene multimedia data, and generating a target object annotation box; training according to the target scene multimedia data and the corresponding target object labeling box to obtain a target recognition sub-model; and extracting a local picture according to the target object labeling box, and training a target segmentation sub-model in combination with target object pixels in the local picture. According to the method, the target identification sub-model and the target segmentation sub-model jointly form the multimedia data annotation model, the multimedia data annotation model can identify the target object from the multimedia data, and the approximate position of the target object is given through the annotation box; and then the target object pixels are further segmented from the local picture corresponding to the extracted labeling box, the segmentation accuracy is improved through the two sub-models, and the defect that the segmentation accuracy is not high when pixel-level segmentation is directly carried out on the whole multimedia data is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of video processing technology, and in particular to a multimedia data annotation model generation method and a multimedia data annotation method. Background Art

[0002] Videos contain a large amount of information, and users can learn, create, etc. from them. For example, users can search for video frames containing a certain film or television character from existing videos, and then cut out the film or television character from the video frames for secondary creation; for example, doctors can search for video frames containing a certain type of surgical tool from surgical videos, and learn to use specific surgical tools to complete specific surgical operations.

[0003] However, these videos are often quite long. For example, a neurosurgery video can last approximately three hours. Learning how to operate a specific surgical tool by watching a completed surgical video is a time-consuming task. For example, a single operation of a pair of scissors might only take 10 seconds, but a neurosurgery video lasts less than 20 minutes. Watching a three-hour video for just 20 minutes of content is clearly a waste of time.

[0004] In addition, users may have further needs, such as extracting film and television characters from video frames for secondary creation, or extracting tools from surgical videos to facilitate further production of teaching videos, or extracting tools from real-time surgical video streams to assist in performing some automatic operations, such as early warning, visual focus, etc.

[0005] Deep learning-based object recognition and segmentation models have made some progress, but they still have some limitations. For example, the accuracy of object recognition models in complex scenarios needs to be improved, while the accuracy of object boundary segmentation in object segmentation models also needs to be further optimized. Direct pixel-level segmentation of entire images / videos is not very accurate, and "improving the accuracy of object segmentation models by expanding and improving the quality of training samples" requires significant training costs. Furthermore, existing methods often treat object recognition and segmentation as two separate tasks, training machine learning models for each task separately. This fails to fully exploit the correlation between the two, resulting in limited model performance.

[0006] In order to solve or at least partially solve the defects of the prior art in finding target information and segmenting target objects in videos, which are low efficiency and poor segmentation accuracy, the present invention provides a multimedia data annotation model generation method and a multimedia data annotation method. Summary of the Invention

[0007] The present invention provides a multimedia data annotation model generation method and a multimedia data annotation method, which are used to at least partially solve the defects of the prior art in finding target information and segmenting target objects in videos, such as low efficiency and poor segmentation accuracy.

[0008] The present invention provides a method for generating a multimedia data annotation model, comprising:

[0009] Mark the target object pixels in the target scene multimedia data and generate the target object annotation box;

[0010] The target recognition sub-model is obtained by training the target scene multimedia data and the corresponding target object annotation box;

[0011] A local image is extracted based on the target object annotation box, and the target segmentation sub-model is trained based on the target object pixels in the local image.

[0012] Optionally, the target scene multimedia data includes real scene multimedia data, and the step of marking target object pixels in the target scene multimedia data to generate a target object marking frame includes:

[0013] Marking target object pixels in the picture of the real scene multimedia data;

[0014] A labeling box encompassing the target object is generated according to the range of the target object in the image.

[0015] Furthermore, the real scene multimedia data includes real video data, and before marking target object pixels on a picture in the real scene multimedia data, the steps include:

[0016] Images are extracted from the real video data according to a preset or user-input interval.

[0017] Optionally, the target scene multimedia data includes synthetic data, and marking target object pixels in the target scene multimedia data to generate a target object annotation box includes:

[0018] With green as the background, shoot videos for each type of target object;

[0019] Extract images from the captured video and automatically segment the target objects into corresponding categories;

[0020] Based on the segmented target objects and the background image that matches the target scene, a synthetic image of the target scene is generated;

[0021] Determine the labeling data of the composite image according to the positions of various target objects in the composite image.

[0022] Furthermore, generating a composite image of the target scene based on the segmented target objects and a background image that matches the target scene also includes:

[0023] The composite image is subjected to boundary blurring.

[0024] Optionally, generating a composite image of the target scene based on the segmented target objects and a background image that matches the target scene includes:

[0025] The background image is placed at the bottom layer, and a number of different types of target objects are superimposed on the background image at different levels to generate multiple composite images.

[0026] Furthermore, determining the annotation data of the synthetic image according to the positions of various target objects in the synthetic image includes:

[0027] For each pixel in the composite image, the target object type located at the top layer is used as the label corresponding to the pixel.

[0028] Optionally, before the target recognition sub-model is obtained by training based on the target scene multimedia data and the corresponding target object annotation box, it also includes: preprocessing the target scene multimedia data, the annotated target object pixels, and the generated target object annotation box, and adjusting them to a first preset size.

[0029] Optionally, the target recognition sub-model is obtained by training the target scene multimedia data and the corresponding target object annotation box, including:

[0030] Generate a first training sample set based on the target scene multimedia data and the corresponding target object annotation box;

[0031] The pre-trained object recognition model is trained and fine-tuned using the first training sample set to obtain the target recognition sub-model.

[0032] Furthermore, the pre-trained object recognition model is a DETR pre-trained model.

[0033] Optionally, extracting a local image according to the target object annotation box and training a target segmentation sub-model based on the target object pixels in the local image includes:

[0034] intercepting a partial picture from the target scene multimedia data according to the target object annotation frame, and generating a second training sample set by combining the target object pixel annotation in the partial picture;

[0035] The second training sample set is used to perform training and fine-tuning on the pre-trained object segmentation model to obtain a target segmentation sub-model.

[0036] Furthermore, the pre-trained object segmentation model is a SAM pre-trained model.

[0037] The SAM pre-training model segments the local images corresponding to the target object annotation boxes in a picture respectively to obtain the segmentation results of each type of target object, and fuses the segmentation results of each type of target object to obtain the segmentation result of the picture.

[0038] Optionally, the local image is extracted according to the target object annotation box, and the target segmentation sub-model is trained in combination with the target object pixels in the local image, and the intercepted local image and the target object pixel standard in the local image are also adjusted to a second preset size.

[0039] Optionally, during model training:

[0040] First, perform preliminary training on the target recognition sub-model;

[0041] The target recognition sub-model is then jointly trained with the target segmentation sub-model to jointly optimize the loss functions of the two sub-models.

[0042] The present invention also provides a multimedia data annotation method, comprising:

[0043] Extracting a picture to be processed from multimedia data of a target scene to be processed;

[0044] Inputting the image to be processed into the target recognition sub-model obtained according to any of the aforementioned multimedia data annotation model generation methods, and outputting the target object annotation box in the image to be processed;

[0045] Extracting a local image from the image to be processed according to the segmented target object annotation frame;

[0046] The local image is input into the target segmentation sub-model obtained according to any of the aforementioned multimedia data annotation model generation methods, and the target object pixels in the image to be processed are output.

[0047] Furthermore, the method further comprises: searching for a video frame in which the target object to be searched appears in the target scene multimedia data according to the target object to be searched input by the user, and presenting the video frame to the user.

[0048] The present invention also provides a computer program product, which includes computer-executable instructions, characterized in that when the instructions are executed, they are used to implement the steps of the multimedia data annotation model generation method as described in any of the above items, or when the instructions are executed, they are used to implement the steps of the multimedia data annotation method as described in any of the above items.

[0049] The present invention provides a method for generating a multimedia data annotation model and a method for annotating multimedia data, which have at least the following beneficial effects:

[0050] 1. The target recognition sub-model identifies the category and approximate location of target objects in multimedia data. On this basis, a local image containing the target object is extracted and input into the target segmentation sub-model. This can accurately locate the target objects and target object pixels appearing in the video, comprehensively improving the recognition accuracy of target objects and the pixel-level segmentation and positioning accuracy.

[0051] 2. By synthesizing multimedia data of the target scene, the training samples are greatly expanded, the model training cost is reduced, and the model performance is improved; by adjusting the level and angle of the target object, the target scene data is enriched.

[0052] 3. By first preliminarily training the target recognition sub-model and then jointly training the target recognition sub-model with the target segmentation sub-model, the convergence speed of the joint model (multimedia data labeling model) and the segmentation accuracy of the model are accelerated.

[0053] 4. For offline videos, users can enter the target object to be searched and obtain the time period in which all objects of this type appear in the video, saving a lot of search time. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0055] Figure 1 This is a flowchart of a method for generating a multimedia data annotation model provided by the present invention;

[0056] Figure 2 This is an example of a picture taken of a real surgical scene;

[0057] Figure 3 yes Figure 2 An example of marking the pixels of the target object in the picture;

[0058] Figure 4 is a video frame in a video shot of a surgical tool;

[0059] Figure 5 yes Figure 4 The annotation corresponding to the video frame;

[0060] Figure 6 This is one of the examples of synthesized images and their annotations;

[0061] Figure 7 This is the second example of a composite image and its annotations;

[0062] Figure 8 It is a flowchart of a multimedia data annotation method provided by the present invention. DETAILED DESCRIPTION

[0063] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0064] The following combination Figures 1-8 Describe a multimedia data annotation model generation method and a multimedia data annotation method of the present invention, Figure 1 This is a flow chart of a method for generating a multimedia data annotation model provided by the present invention. Figure 1 As shown, the method includes:

[0065] S11, marking target object pixels in the target scene multimedia data and generating a target object marking frame;

[0066] S12, training a target recognition sub-model based on the target scene multimedia data and the corresponding target object annotation box;

[0067] S13. Extract a local image according to the target object annotation box, and train a target segmentation sub-model based on the target object pixels in the local image.

[0068] Specifically, the target scene is picture / video data taken of the scene that the user is interested in. For example, if the user wants to observe and learn specific surgical operations, the target scene multimedia data can be pictures / videos taken of the surgical scene; for example, if the user wants to edit the video of character A, the target scene multimedia data includes the video data of character A.

[0069] Step S11 obtains the pixel-level annotation of the target object in the multimedia data, and generates an annotation box based on the pixel-level annotation of the target object. The annotation box, such as a rectangular box, a polygonal box, or a circular box, can circle the target object in the picture with a reasonable size. Step S12 first uses the target scene multimedia data and the corresponding target object annotation box to train to obtain a "target recognition sub-model that can select the target object from the picture and video data frame". The target recognition sub-model can be a deep learning model using architectures such as Faster R-CNN, R-FCN, and YOLO. Step S13 trains the local picture within the standard frame and the pixel annotation of the target object in the local picture to obtain a "target segmentation sub-model that can segment the target object pixels from the local picture within the standard frame". The target segmentation sub-model can be a deep learning model using architectures such as FCN, Mask R-CNN, and U-Net.

[0070] It can be understood that the above-mentioned target recognition sub-model and target segmentation sub-model together constitute a multimedia data annotation model. The multimedia data annotation model obtained by this method can first identify the target object in the multimedia data, give the approximate position of the target object through the annotation box, and then further segment the target object pixels from the local image corresponding to the extracted annotation box, thereby improving the segmentation accuracy and avoiding the defect of low segmentation accuracy when directly performing pixel-level segmentation on the entire media data.

[0071] In addition, the above-mentioned recognition and segmentation of target objects may be recognition and segmentation of target objects of a single category, or recognition and segmentation of target objects of two or more categories.

[0072] In a specific example, the target recognition sub-model of the present invention has an accuracy rate of 99% in identifying targets, and the target segmentation sub-model has an accuracy rate of 96% in segmenting target pixels.

[0073] Based on any embodiment, in one embodiment, the target scene multimedia data includes real scene multimedia data, and S11 includes:

[0074] S112, marking target object pixels in the picture in the real scene multimedia data;

[0075] S113 : Generate a labeling box encompassing the target object according to the range of the target object in the image.

[0076] Specifically, taking the surgical scene as an example, the real scene multimedia data is the pictures and video data taken of the real surgical scene. Step S112 gives the pixel-level annotation of the target object to the real data. Figure 2 、 Figure 3 , Figure 2 Shows pictures taken of real surgical scenes. Figure 3Two surgical tools (two types of target objects) are marked in the picture. Step S113 automatically generates a marking frame according to the range of target object pixels, without manually providing a marking frame.

[0077] Taking the rectangular annotation box as an example, the rectangular annotation box can be [X min ,X max ,Y min ,Y max , category label] description, where X min Indicates the minimum X coordinate value of the pixel of this type of target object, X max Indicates the maximum X coordinate of the pixel of this type of target object, Y min Indicates the minimum Y coordinate value of the pixel of this type of target object, Y max Indicates the maximum Y coordinate of the pixel of the target object of this class. For the category label, if only one type of target object is recognized, "1" can be used as the label of the target object of this class, and "0" can represent the background. Similarly, if multiple types of target objects are recognized, "1", "2", "3", etc. can be used to represent the target objects of different classes respectively.

[0078] Based on the previous embodiment, in one embodiment, the real scene multimedia data includes real video data, and the steps before S112 include:

[0079] S111 . Extracting pictures from the real video data according to a preset interval or an interval input by the user.

[0080] Specifically, the interval here is, for example, extracting 1 frame every 10 frames, or extracting 1 frame of image per second. Images are extracted from real videos for processing according to preset or user-entered intervals, avoiding large data processing volume, reducing similar redundant data, and avoiding overfitting of the model caused by too much identical training data.

[0081] Based on any of the embodiments, in one embodiment, the target scene multimedia data includes synthesized data, and S11 includes:

[0082] S115, taking a green background and shooting a video of each type of target object;

[0083] S116, extracting images from the captured video and automatically segmenting target objects of corresponding categories;

[0084] S117, generating a composite image of the target scene based on the segmented target objects and a background image that matches the target scene;

[0085] S118 . Determine the annotation data of the synthetic image according to the positions of the various target objects in the synthetic image.

[0086] Specifically, Figure 4 The figure shows a video frame (picture) in a video shot of a surgical tool. In step S115, the background is green, which makes the target object and the background most distinct, making it easier to cut out the image in step S116. It can be understood that step S115 is to shoot a video for each type of target object separately, and step S116 is to cut out the image for each type of target object separately. Figure 5 Shows the Figure 4 The surgical tool area (white area) is extracted from the video frame through the automatic detection algorithm, and the pixels in this area are the surgical tool pixels that are automatically marked. Step S117 synthesizes a series of pictures. Taking the surgical scene as an example, the background picture in step S117 can be an image of the surgical area tissue. It can be understood that in the process of synthesizing pictures, the surgical area tissue image is superimposed on the bottom layer, and the target object is placed on the upper layer. Only one type of target object may appear in the synthesized picture, or multiple types of target objects may appear to simulate complex scenes. S118 can directly determine the annotation data corresponding to the synthesized picture based on the various types of target objects placed, Figure 6 、 Figure 7 An example of two composite images and corresponding annotations is shown.

[0087] This embodiment can greatly reduce the difficulty of obtaining labeled data by synthesizing data, expand training data, and reduce model training costs.

[0088] Based on the previous embodiment, in one embodiment, the S117 further includes:

[0089] The composite image is subjected to boundary blurring.

[0090] Specifically, the image can be blurred using blurring algorithms such as Gaussian blur, Laplace pyramid, Alpha blur, etc. Blurring the boundaries can make the transition between the target object boundary and the background natural, avoid the segmentation task being too simple, and improve the generalization ability of the model.

[0091] Based on any embodiment, in one embodiment, the S117 includes:

[0092] The background image is placed at the bottom layer, and a number of different types of target objects are superimposed on the background image at different levels to generate multiple composite images.

[0093] Specifically, there is no limit on the number of categories of target objects that appear on the background image. For example, in a composite image, only one type of target object is placed on the background image; for another example, in another composite image, three types of target objects are placed on the background image. In addition, when two or more types of target objects are placed on the background image, the stacking order of each type of target object can be flexibly adjusted to generate multiple composite images, thereby improving the diversity of training samples. In addition, the composite image can also be adjusted in other ways, for example, by flipping or translating a target object, or by flipping the background image, or by adding blood drop data, random spots, or motion blur to the composite image, or by adjusting the brightness of the composite image.

[0094] Based on any embodiment, in one embodiment, the S118 includes:

[0095] For each pixel in the composite image, the target object type located at the top layer is used as the label corresponding to the pixel.

[0096] For example, in the process of synthesizing an image, the first type of tool M and the second type of tool N are superimposed on the background image in sequence. For the pixels where the two types of tools overlap, the corresponding standard is the second type of tool N; for the non-overlapping pixels of the tools, the corresponding label is the corresponding tool type; for the pixels where the tools are not covered, the corresponding label is background.

[0097] Based on any embodiment, in one embodiment, before S12, the step further includes: pre-processing the target scene multimedia data, the marked target object pixels, and the generated target object marking frame, and adjusting them to a first preset size.

[0098] Specifically, the multimedia data and the corresponding annotation data are adjusted to a first preset size to facilitate processing by the target recognition sub-model. The first preset size is set according to demand, for example, 1024×1024 pixels. During the adjustment process, operations such as cutting, scaling, rotating, and filling can be performed. For example, if the image is too large, the image can be cut so that the target object appears in the cut image with an appropriate size ratio. For another example, after the image is scaled and adjusted, the size of the image in a certain direction is smaller, and the missing part in that direction can be filled with a pure black background to the preset size.

[0099] Based on any embodiment, in one embodiment, the S12 includes:

[0100] Generate a first training sample set based on the target scene multimedia data and the corresponding target object annotation box;

[0101] The pre-trained object recognition model is trained and fine-tuned using the first training sample set to obtain the target recognition sub-model.

[0102] Specifically, the pre-trained object recognition model can accelerate the convergence of the model in the subsequent training process, reduce the number of data sets required for training and the training time. Pre-trained object recognition models, such as the DETR pre-trained model, DETR (Detection Transformer), is a visual version of the Transformer proposed by researchers at Facebook AI and can be used for target detection. The pre-trained object recognition model is trained and fine-tuned using the first training sample set, so that the trained target recognition sub-model has the ability to "identify target objects in other image processing" (i.e., give a label box to the target object in the image).

[0103] Based on any embodiment, in one embodiment, the S13 includes:

[0104] intercepting a partial picture from the target scene multimedia data according to the target object annotation frame, and generating a second training sample set by combining the target object pixel annotation in the partial picture;

[0105] The second training sample set is used to perform training and fine-tuning on the pre-trained object segmentation model to obtain a target segmentation sub-model.

[0106] Specifically, the pre-trained object segmentation model can speed up the convergence of the model in the subsequent training process, reduce the number of data sets required for training and the training time. The pre-trained object segmentation model is, for example, the SAM pre-training model. The SAM model (Segment Anything Model) is a general artificial intelligence model released by Meta, which focuses on image segmentation tasks. The SAM model adopts an encoder-decoder design to achieve efficient segmentation of any object by fusing image features with prompt information. The model is trained based on a data set containing 11 million images and 1.1 billion annotations. It has zero-sample generalization capabilities and can handle untrained scenes and blurred objects. The pre-trained object segmentation model is trained and fine-tuned using the second training sample set, so that the trained target segmentation sub-model has the ability to "segment out target object pixels from other local images".

[0107] Based on the previous embodiment, in one embodiment, the pre-trained object segmentation model is a SAM pre-trained model.

[0108] The SAM pre-training model segments the local images corresponding to the target object annotation boxes in a picture respectively to obtain the segmentation results of each type of target object, and fuses the segmentation results of each type of target object to obtain the segmentation result of the picture.

[0109] Specifically, during the fusion process, the segmentation results of each target object annotation box are restored to the image space, that is, restored to the position of the local image, and the pixels in other areas are marked as background. When there are multiple annotation boxes in the image, for each pixel, its pixel label is determined based on the segmentation results of the local image in each annotation box. Pixels without labels are classified as background.

[0110] Based on any embodiment, in one embodiment, in S13, the captured partial image and the pixel standard of the target object in the partial image are further adjusted to a second preset size.

[0111] Specifically, the local image and the corresponding annotation data are adjusted to a second preset size to facilitate processing by the target segmentation sub-model. The second preset size is set according to demand, for example, 512×512 pixels. During the adjustment process, operations such as scaling, rotation, and padding can be performed. For example, if the local image is too large or too small, the image can be scaled to reach the appropriate size. For another example, after the image is scaled and adjusted, the size of the image in a certain direction is smaller, and the missing part in that direction can be filled with a pure black background to the preset size.

[0112] Based on any embodiment, in one embodiment, during the model training process:

[0113] First, perform preliminary training on the target recognition sub-model;

[0114] The target recognition sub-model is then jointly trained with the target segmentation sub-model to jointly optimize the loss functions of the two sub-models.

[0115] Specifically, this embodiment first trains the target recognition sub-model separately, then jointly trains it with the target segmentation sub-model. The parameters of the two sub-models are adjusted to minimize the joint loss function of the two models, resulting in a multimedia data annotation model. This embodiment can accelerate the convergence of the multimedia data annotation model and improve its segmentation accuracy. Through the coordinated optimization of the two sub-models, the correlation between target recognition and target segmentation can be fully utilized, improving the overall performance of the model.

[0116] A preferred embodiment is described below. The method for generating a multimedia data annotation model provided in this embodiment includes the following steps:

[0117] Step 1: Preprocess the target scene multimedia data:

[0118] Step 101: extract pictures from real video data according to the interval input by the user (e.g., every 20 frames);

[0119] Step 102: manually label the target object pixels in the extracted image to generate a target object labeling box. The coordinate range of the labeling box is the minimum circumscribed rectangle of the target object. If there are multiple types of target objects in the image, labeling boxes are given separately.

[0120] Step 103: Adjust the real scene multimedia data, the marked target object pixels, and the generated target object annotation box to a first preset size, for example, adjusting the image resolution to 800×600 pixels;

[0121] Step 2: Preprocess the synthetic data:

[0122] Step 201: Using a green background, capture a video of each target object (e.g., various surgical tools) with a resolution of 1920×1080 pixels.

[0123] Step 202: extracting images from the captured video and automatically segmenting the target objects of the corresponding categories using an image segmentation algorithm;

[0124] Step 203: Generate a composite image of the target scene based on the segmented target objects and a background image that matches the target scene (such as an image of the surgical area tissue). During the synthesis process, the background image is placed at the bottom layer, and several different types of target objects are superimposed on the background image at different levels to generate multiple composite images to simulate complex scenes. This also includes performing boundary blurring on the composite image, using a Gaussian blur algorithm to blur the boundaries of the target objects.

[0125] Step 204: Determine the annotation data of the composite image based on the positions of each type of target object in the composite image. For each pixel in the composite image, use the target object type at the top layer as the annotation corresponding to the pixel. On this basis, use the minimum enclosing rectangle of each type of target object pixel as the annotation box of the type of target object. If there are multiple types of target objects in the composite image, give annotation boxes for each type.

[0126] Step 205 : Adjust the synthesized image data, the marked target object pixels, and the generated target object marking frame to a first preset size, for example, adjust the image resolution to 800×600 pixels.

[0127] It is understandable that there is no order in which steps 1 and 2 are performed.

[0128] Step 3: Generate a first training sample set based on the preprocessed target scene multimedia data and the corresponding target object annotation box, and use the first training sample set to train and fine-tune the pre-trained object recognition model (such as the detr pre-training model) to obtain a target recognition sub-model;

[0129] Step 4: Extract a partial image from the target scene multimedia data based on the target object annotation frame, and adjust the extracted partial image and the target object pixel annotation within the partial image to a second preset size (e.g., 256×256 pixels) in combination with the target object pixel annotation within the partial image as the second training sample set. The pre-trained object segmentation model (e.g., the SAM pre-trained model) is trained and fine-tuned using the first and second training sample sets to obtain a target segmentation sub-model.

[0130] Step 5: First, perform preliminary training on the target recognition sub-model, and then jointly train the target recognition sub-model and the target segmentation sub-model to jointly optimize the loss functions of the two sub-models. During the joint training process, the input of the target recognition sub-model is the target scene multimedia data, and the output is the annotation box given by the target recognition sub-model. According to the annotation box, refer to step 4 to intercept the local image and input it into the target segmentation sub-model, and output its pixel-level annotation. Then, the annotation box given by the target recognition sub-model is compared with the standard annotation box of the training data, and the pixel-level annotation output by the target segmentation sub-model is compared with the standard annotation of the training data. The loss function is jointly calculated and the loss function of the two sub-models is jointly optimized.

[0131] A multimedia data annotation method provided by the present invention is described below. The multimedia data annotation method described below and the multimedia data annotation model generation method described above can be referenced to each other.

[0132] Reference Figure 8 The present invention provides a multimedia data tagging method, comprising:

[0133] S21, extracting a picture to be processed from the target scene multimedia data to be processed;

[0134] S22: inputting the image to be processed into the target recognition sub-model obtained according to any of the aforementioned multimedia data annotation model generation methods, and outputting the target object annotation box in the image to be processed;

[0135] S23, extracting a local image from the image to be processed according to the segmented target object annotation frame;

[0136] S24. Input the local image into the target segmentation sub-model obtained according to any of the aforementioned multimedia data annotation model generation methods, and output the target object pixels in the image to be processed.

[0137] Specifically, step S21 selects the current image to be processed from a number of target scene images to be processed, or extracts a video frame to be processed from a target scene video. Step S22 first uses the target recognition sub-model trained using the aforementioned method to give a target object annotation box in the image to be processed, first identifying the category and approximate location of the target object. Then, steps S23-S24 extract a local image from the image based on the target object annotation box and input it into the target segmentation sub-model trained using the aforementioned method to identify the target object pixels.

[0138] The marked target object pixels can also be used to support other functions and assist in implementing some automatic operations, such as tool operation tracking, tool tip trajectory analysis, safety warnings, guiding equipment autofocus, etc.

[0139] It can be understood that each local image corresponds to a category of target objects. The type of target object in the local image has been determined. The target recognition sub-model only needs to further determine whether each pixel in the local image belongs to this category of target object. That is to say, the method of the present invention not only transforms the task of overall image processing into the task of local image processing, reducing the interference of irrelevant pixel information, but also transforms the multi-classification task into a binary classification task, greatly improving the segmentation accuracy.

[0140] After segmenting each local image, the (binary classification) segmentation results of each local image can be restored to the original image space and fused with the types of each local image identified by the target recognition sub-model to obtain a fused multi-classification result. It is understood that this step is not required. For example, only the pixels corresponding to the target type of object of interest can be restored to the original image space.

[0141] This method first identifies the category and approximate location of target objects in multimedia data through the target recognition sub-model, and then extracts the local image that includes the target object and inputs it into the target segmentation model. Through the coordinated optimization of the two sub-models, it can fully utilize the correlation between target recognition and target segmentation, and can accurately locate the target objects and target object pixels that appear in the video, comprehensively improving the recognition accuracy of target objects and the pixel-level segmentation and positioning accuracy, avoiding the defect of low segmentation accuracy when directly performing pixel-level segmentation on the entire media data, and improving the overall performance of the model.

[0142] Based on any embodiment, in one embodiment, the method further includes:

[0143] S25 . According to the target object to be searched input by the user, a video frame in which the target object to be searched appears is searched in the target scene multimedia data, and the video frame is presented to the user.

[0144] Specifically, the target scene video is processed using a multimedia data annotation method. For example, every 15 frames are processed to identify the type of target object appearing in the video frame. Users can enter the target object they want to view, such as "tweezers." This method can then identify the video frames where "tweezers" appear. As an equivalent variation, this method can also identify the video time periods where "tweezers" appear, making it easier for users to observe and learn.

[0145] More embodiments of a multimedia data annotation method can be compared with the above multimedia data annotation model generation method, which will not be repeated here.

[0146] The multimedia data tagging device provided by the present invention is described below. The multimedia data tagging device described below and the multimedia data tagging method described above can be referenced to each other.

[0147] The present invention provides a multimedia data annotation device, comprising:

[0148] An acquisition module is used to extract the image to be processed from the multimedia data of the target scene to be processed;

[0149] a recognition module, configured to input the image to be processed into a target recognition sub-model obtained according to any of the aforementioned multimedia data annotation model generation methods, and output a target object annotation box in the image to be processed;

[0150] An extraction module is used to extract a local image from the image to be processed according to the segmented target object annotation box;

[0151] A segmentation module is used to input the local image into the target segmentation sub-model obtained according to any of the aforementioned multimedia data annotation model generation methods, and output the target object pixels in the image to be processed.

[0152] This device uses a recognition module to identify the category and approximate location of target objects in multimedia data. Based on this, an extraction module extracts a partial image encompassing the target object, which is then fed into a segmentation module to accurately segment the target object's pixels. By leveraging the collaborative optimization of two sub-models, this device leverages the correlation between target recognition and segmentation, accurately locating target objects and their pixels within the video. This improves both target recognition accuracy and pixel-level segmentation accuracy, avoiding the inherent inaccuracy of direct pixel-level segmentation of the entire media data.

[0153] The technical solution of the present invention, or the portion that contributes to the prior art, or the portion of the technical solution, may be embodied in the form of a software product. This computer software product is stored in a storage medium and includes instructions for causing a computer device (which may be a personal computer, server, or network device, etc.) to execute the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0154] On the other hand, the present invention also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the steps of the multimedia data annotation model generation method provided above.

[0155] On the other hand, the present invention also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the steps of the multimedia data labeling method provided above.

[0156] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0157] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.

[0158] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A method for generating a multimedia data annotation model, characterized in that: include: S11, marking target object pixels in the target scene multimedia data and generating a target object marking frame; S12, training a target recognition sub-model based on the target scene multimedia data and the corresponding target object annotation box; S13. Extract a local image according to the target object annotation box, and train a target segmentation sub-model based on the target object pixels in the local image.

2. The method for generating a multimedia data annotation model according to claim 1, wherein: The target scene multimedia data includes real scene multimedia data, and S11 includes: S112, marking target object pixels in the picture in the real scene multimedia data; S113 : Generate a labeling box encompassing the target object according to the range of the target object in the image.

3. The method for generating a multimedia data annotation model according to claim 1, wherein: The target scene multimedia data includes synthetic data, and S11 includes: S115, taking a green background and shooting a video of each type of target object; S116, extracting images from the captured video and automatically segmenting target objects of corresponding categories; S117, generating a composite image of the target scene based on the segmented target objects and a background image that matches the target scene; S118 . Determine the annotation data of the synthetic image according to the positions of the various target objects in the synthetic image.

4. The method for generating a multimedia data annotation model according to claim 3, wherein: The S117 includes: The background image is placed at the bottom layer, and a number of different types of target objects are superimposed on the background image at different levels to generate multiple composite images.

5. The method for generating a multimedia data annotation model according to claim 4, wherein: The S118 includes: For each pixel in the composite image, the target object type located at the top layer is used as the label corresponding to the pixel.

6. The method for generating a multimedia data annotation model according to claim 1, wherein: Before S12, the method further includes: pre-processing the target scene multimedia data, the marked target object pixels, and the generated target object marking frame, and adjusting them to a first preset size.

7. The method for generating a multimedia data annotation model according to claim 1, wherein: The S12 includes: Generate a first training sample set based on the target scene multimedia data and the corresponding target object annotation box; The pre-trained object recognition model is trained and fine-tuned using the first training sample set to obtain the target recognition sub-model.

8. The method for generating a multimedia data annotation model according to claim 1, wherein: The S13 includes: intercepting a partial picture from the target scene multimedia data according to the target object annotation frame, and generating a second training sample set by combining the target object pixel annotation in the partial picture; The second training sample set is used to perform training and fine-tuning on the pre-trained object segmentation model to obtain a target segmentation sub-model.

9. The method for generating a multimedia data annotation model according to claim 8, wherein: The pre-trained object segmentation model is a SAM pre-trained model. The SAM pre-training model segments the local images corresponding to the target object annotation boxes in a picture respectively to obtain the segmentation results of each type of target object, and fuses the segmentation results of each type of target object to obtain the segmentation result of the picture.

10. The method for generating a multimedia data annotation model according to claim 1, wherein: During model training: First, perform preliminary training on the target recognition sub-model; The target recognition sub-model is then jointly trained with the target segmentation sub-model to jointly optimize the loss functions of the two sub-models.

11. A multimedia data annotation method, characterized in that: include: S21, extracting a picture to be processed from the target scene multimedia data to be processed; S22: inputting the image to be processed into a target recognition sub-model obtained by the multimedia data annotation model generation method according to any one of claims 1 to 10, and outputting a target object annotation box in the image to be processed; S23, extracting a local image from the image to be processed according to the segmented target object annotation frame; S24. Input the local image into the target segmentation sub-model obtained according to any one of the multimedia data annotation model generation methods according to claims 1-10, and output the target object pixels in the image to be processed.

12. A multimedia data annotation method according to claim 11, characterized in that: The method also includes: S25 . According to the target object to be searched input by the user, a video frame in which the target object to be searched appears is searched in the target scene multimedia data, and the video frame is presented to the user.

13. A computer program product comprising computer-executable instructions, characterized in that: When executed, the instruction is used to implement the steps of the multimedia data annotation model generation method as described in any one of claims 1 to 10, or, when executed, the instruction is used to implement the steps of the multimedia data annotation method as described in any one of claims 11 or 12.