A content annotation method, device, computer equipment and readable storage medium based on human body skill operation video

By preprocessing and feature extraction of human limb skills operation videos, combining matching recognition technology and pre-training content label library, high-precision and efficient content annotation are achieved, solving the problems of low labeling accuracy and insufficient efficiency in the existing technology.

CN118781516BActive Publication Date: 2025-05-06DMAI (GUANGZHOU) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410785783.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-18
Publication Date
2025-05-06
Estimated Expiration
2044-06-18

AI Technical Summary

Technical Problem

When the prior art automatically labels human limb skills operation videos, the labeling accuracy is not high, the key points are not accurate, the efficiency is low, and it is susceptible to human factors.

Method used

By acquiring multiple video streaming media data, preprocessing is performed to extract video and audio features, evaluation result data is generated using matching recognition technology, operation key points and video clip positioning of operation units, and labeling is performed in combination with pre-trained content label library.

Benefits of technology

It improves the labeling accuracy and efficiency, accurately identify and mark the operation key points and operating units in the human limb skills operation video, and provides accurate evaluation support for human limb skills operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118781516B_ABST
    Figure CN118781516B_ABST
Patent Text Reader

Abstract

The present invention discloses a content annotation method, device, computer equipment and readable storage medium based on human body skill operation video, including: first, obtaining and preprocessing multi-channel video streaming media data, extracting video key frames, visual and acoustic features. Subsequently, video and audio evaluation result data are generated through matching recognition operations, and the video segment positioning of operation points and operation units is determined accordingly. Label annotation is performed using a pre-trained content label library to generate pending label annotation results. After passing the review, the result will be used as the final target annotation result. Such a design improves the annotation accuracy and efficiency, and provides strong support for the accurate evaluation of human body skill operations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a content annotation method, device, computer equipment and readable storage medium based on a human body skill operation video. Background Art

[0002] In the prior art, the content annotation of human body skill operation videos usually relies on manual work, which is not only inefficient, but also the annotation quality is easily affected by human factors. With the development of multimedia technology, automatic annotation technology has gradually become a research hotspot. However, when dealing with complex human body skill operations, the existing automatic annotation methods often have problems such as low annotation accuracy and inaccurate identification of operation points. Therefore, there is an urgent need for a method that can accurately annotate the content of human body skill operation videos. Summary of the invention

[0003] The object of the present invention is to provide a content annotation method, device, computer equipment and readable storage medium based on human body skill operation video.

[0004] In a first aspect, an embodiment of the present invention provides a content annotation method based on a human body skill operation video, comprising:

[0005] Acquire multi-channel video streaming data containing human body skill operations;

[0006] Preprocessing the multi-channel video streaming media data to obtain preprocessed video data and audio data;

[0007] Extracting features from the video data and the audio data respectively to obtain video key frames and visual feature data corresponding to the video data, and acoustic feature data corresponding to the audio data;

[0008] Performing matching and recognition operations on the video key frame and the visual feature data, as well as the acoustic feature data, respectively, to obtain video evaluation result data and audio evaluation result data;

[0009] Determine the operation points of the human body skill operation and the video segment location corresponding to the operation unit according to the video evaluation result data and the audio evaluation result data;

[0010] Based on the video clip positioning, the operation points and the operation units are labeled in combination with a pre-trained content label library to obtain a pending label labeling result;

[0011] In the case that the pending label marking result passes the review, the pending label marking result is used as the target marking result for the human body skill operation.

[0012] In the embodiment of the present invention, the preprocessing of the multi-channel video streaming media data to obtain the preprocessed video data and audio data includes:

[0013] The synchronous processor synchronously processes the multi-channel video streaming media data, and decodes the synchronously processed multi-channel video streaming media data to obtain initial video data and initial audio data;

[0014] De-noising and image enhancement are performed on the initial video data to obtain the pre-processed video data;

[0015] The initial audio data is denoised and filtered to obtain the preprocessed audio data.

[0016] In the embodiment of the present invention, the feature extraction of the video data and the audio data is performed respectively to obtain the video key frame and visual feature data corresponding to the video data, and the acoustic feature data corresponding to the audio data, including:

[0017] Sparsely sampling the video data at fixed intervals to obtain a plurality of original video key frames;

[0018] Performing deep feature extraction on the multiple original video key frames to obtain original video feature data corresponding to each of the original video key frames;

[0019] Performing clustering processing on the plurality of original video feature data, and determining the video key frames and the visual feature data corresponding to the plurality of original video feature data from the plurality of original video key frames;

[0020] The acoustic feature data corresponding to the audio data is extracted using the MFCC algorithm.

[0021] In the embodiment of the present invention, the matching and identifying operations are performed on the video key frame and the visual feature data, and the acoustic feature data, respectively, to obtain the video evaluation result data and the audio evaluation result data, including:

[0022] Inputting the video key frames into a pre-trained object recognition model to obtain key frame categories;

[0023] Input the visual feature data into a pre-trained behavior recognition model to obtain the target object and visual related attributes corresponding to the video key frame, wherein the target object includes body parts, clothing and operating tools;

[0024] Performing matching and identification operations on the target object and visual related attributes to obtain the video evaluation result data;

[0025] Using automatic speech recognition technology to convert the acoustic feature data to obtain speech text content;

[0026] The speech text content is input into a pre-trained speech matching model, and a matching and recognition operation is performed with the standard speech text to obtain the audio evaluation result data.

[0027] In the embodiment of the present invention, determining the operation points of the human body skill operation and the video segment location corresponding to the operation unit according to the video evaluation result data and the audio evaluation result data includes:

[0028] Determining target evaluation result data from the video evaluation result data and the audio evaluation result data, wherein the target evaluation result data includes key frames, key frame categories, target objects of key frames, visual related attributes, matched voice text segments, matching scores, evaluation results, and error cause analysis;

[0029] According to preset rules and in combination with the start time frame and the end time frame corresponding to the target evaluation result data, the operation points of the human body skill operation and the video segment location corresponding to the operation unit are determined.

[0030] In the embodiment of the present invention, the positioning of the video clip and labeling the operation points and the operation units in combination with the pre-trained content label library to obtain the pending label labeling result include:

[0031] Determine the start and end times corresponding to the operation points and the operation units respectively based on the video segment positioning;

[0032] The operation key points and the operation units are labeled according to the start and end times in combination with a pre-trained content label library to obtain the pending label labeling result.

[0033] In an embodiment of the present invention, the method further includes:

[0034] Obtaining the offset time data of the operation unit and the start and end time of the operation key point determined during the review process of the pending labeling result;

[0035] Acquire comparison time data between the operation unit and the operation point according to a preset interval distance;

[0036] Determine a time point offset feature according to the offset time data and the comparison time data;

[0037] The time point offset feature is used as an optimization parameter to optimize the video segment positioning operation.

[0038] In a second aspect, an embodiment of the present invention provides a content annotation device based on a human body skill operation video, comprising:

[0039] An acquisition module is used to acquire multiple channels of video streaming media data containing human body skill operations; preprocess the multiple channels of video streaming media data to obtain preprocessed video data and audio data;

[0040] An execution module is used to perform feature extraction on the video data and the audio data respectively to obtain video key frames and visual feature data corresponding to the video data, and acoustic feature data corresponding to the audio data; perform matching and recognition operations on the video key frames and the visual feature data, and the acoustic feature data respectively to obtain video evaluation result data and audio evaluation result data; and determine the operation points of the human body skill operation and the video segment location corresponding to the operation unit according to the video evaluation result data and the audio evaluation result data;

[0041] The labeling module is used to label the operation points and the operation units based on the positioning of the video clip and in combination with the pre-trained content label library to obtain a pending labeling result; if the pending labeling result passes the review, the pending labeling result is used as the target labeling result for the human body skill operation.

[0042] In a third aspect, an embodiment of the present invention provides a computer device, comprising a processor and a non-volatile memory storing computer instructions, wherein when the computer instructions are executed by the processor, the computer device executes the method described in at least one possible implementation of the first aspect.

[0043] In a fourth aspect, an embodiment of the present invention provides a readable storage medium, wherein the readable storage medium includes a computer program, and when the computer program is executed, the computer device where the readable storage medium is located is controlled to execute the method described in at least one possible implementation manner of the first aspect.

[0044] Compared with the prior art, the beneficial effects provided by the present invention include: using a content annotation method, device, computer equipment and readable storage medium based on human body skill operation video disclosed by the present invention, including: first acquiring and preprocessing multi-channel video streaming data, extracting video key frames, visual and acoustic features. Subsequently, video and audio evaluation result data are generated by matching recognition operations, and the video clip positioning of operation points and operation units is determined accordingly. Label annotation is performed using a pre-trained content label library to generate pending label annotation results. After passing the review, the result will be used as the final target annotation result. Such a design improves the annotation accuracy and efficiency, and provides strong support for the accurate evaluation of human body skill operations. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only illustrate certain embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can also be obtained based on these drawings without creative work.

[0046] Figure 1 A schematic diagram of the steps of a content annotation method based on a human body skill operation video provided by an embodiment of the present invention;

[0047] Figure 2 A schematic block diagram of the structure of a content annotation device based on a human body skill operation video provided by an embodiment of the present invention;

[0048] Figure 3 A schematic block diagram of the structure of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0049] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.

[0050] The specific implementation modes of the present invention are described in detail below in conjunction with the accompanying drawings.

[0051] In order to solve the technical problems in the aforementioned background technology, Figure 1 This is a flow chart of a method for content annotation based on a video of human body skill operation provided by an embodiment of the present disclosure. The method for content annotation based on a video of human body skill operation is introduced in detail below.

[0052] Step S201, obtaining multi-channel video streaming media data containing human body skill operations;

[0053] Step S202, preprocessing the multi-channel video streaming media data to obtain preprocessed video data and audio data;

[0054] Step S203, performing feature extraction on the video data and the audio data respectively to obtain video key frames and visual feature data corresponding to the video data, and acoustic feature data corresponding to the audio data;

[0055] Step S204, performing matching and recognition operations on the video key frame and the visual feature data, as well as the acoustic feature data, to obtain video evaluation result data and audio evaluation result data;

[0056] Step S205, determining the operation points of the human body skill operation and the video segment location corresponding to the operation unit according to the video evaluation result data and the audio evaluation result data;

[0057] Step S206, labeling the operation points and the operation units based on the video segment positioning and in combination with the pre-trained content label library to obtain a pending label labeling result;

[0058] Step S207, when the pending label marking result passes the review, the pending label marking result is used as the target marking result for the human body skill operation.

[0059] In an embodiment of the present invention, exemplarily, the server obtains multi-channel video streaming data containing human body skill operations from various surveillance cameras or videos recorded by operators themselves. For example, in a data center, in order to monitor the server maintenance operations of technicians, the server captures the real-time operation videos of technicians from different monitoring angles. The server preprocesses the acquired raw video streaming data, including steps such as denoising, enhancement, and format conversion, to improve the accuracy of subsequent feature extraction and recognition. For example, for the captured server maintenance operation video, the server performs image stabilization processing, contrast enhancement, and audio noise reduction to obtain clearer video data and purer audio data. The server uses advanced algorithms to extract key frames and analyze visual features of the preprocessed video data, and extracts acoustic features of the audio data. Taking the server maintenance operation as an example, the server may recognize visual features such as the technician's hand movements, facial expressions, and tool usage, and extract acoustic features such as voice commands and equipment sounds. The server matches and identifies the extracted visual features and acoustic features with predefined standard operations to evaluate the skill level of the operator. For example, in server maintenance operations, the server will compare the actual operation of the technician with the standard maintenance process, identify whether each step in the operation is performed correctly, and analyze whether the voice instructions conform to the standard script. Based on the evaluation results of the previous step, the server can accurately locate the key operation points and operation units in the video. In the case of server maintenance, this means that the server can identify when the technician starts to check the hardware, when to debug the software, and other key operations, and accurately locate the video clips of these operations. The server uses the pre-trained label library to automatically annotate the located key operation video clips. For example, in the video of server maintenance, a clip is automatically annotated as "checking the status of the hardware device" and another clip is annotated as "performing software update operations." Finally, the server will send the results of the automatic annotation to the professionals for review. If the review is passed, these annotation results will be regarded as the final target annotation results for subsequent skill assessment and training feedback. In the example of server maintenance, once the review confirms that the annotation is accurate, the company can conduct a detailed evaluation of the technician's operations based on these annotation results, thereby providing them with more personalized training and development plans.

[0060] In the embodiment of the present invention, the aforementioned step S202 may be implemented through the following example execution.

[0061] The synchronous processor synchronously processes the multi-channel video streaming media data, and decodes the synchronously processed multi-channel video streaming media data to obtain initial video data and initial audio data;

[0062] De-noising and image enhancement are performed on the initial video data to obtain the pre-processed video data;

[0063] The initial audio data is denoised and filtered to obtain the preprocessed audio data.

[0064] In an embodiment of the present invention, illustratively, in a data center, multiple cameras simultaneously monitor the server maintenance operations of technicians. These cameras capture the operation process from different angles and generate multiple video streaming media data. The server synchronizes the video streams from different cameras in time through a built-in synchronization processor to ensure that they show the operation scenes at the same time point. After synchronization, the server decodes these video streams and converts them into initial video data and initial audio data for subsequent processing. The initial video data obtained after decoding may contain some noise, such as the electronic noise of the camera itself, noise caused by changes in ambient light, etc. In order to improve the video quality, the server will perform denoising on these initial video data, such as reducing the noise in the image by methods such as median filtering or Gaussian filtering. At the same time, the server will also enhance the video picture, such as adjusting the contrast, brightness and color balance to make the picture clearer, which is convenient for subsequent feature extraction and recognition. Similar to video data, initial audio data may also contain interference factors such as background noise and current sound. The server will perform denoising on these audio data, such as using spectral subtraction or Wiener filtering to reduce the noise level. In addition, the server will filter the audio data to eliminate unnecessary frequency components and highlight the voice signal or key sounds, thereby improving the accuracy of subsequent acoustic feature extraction. Through these preprocessing steps, the server can significantly improve the quality of video and audio data, providing a more accurate data basis for subsequent human body skill operation evaluation and labeling.

[0065] In the embodiment of the present invention, the aforementioned step S203 can be implemented through the following example execution.

[0066] Sparsely sampling the video data at fixed intervals to obtain a plurality of original video key frames;

[0067] Performing deep feature extraction on the multiple original video key frames to obtain original video feature data corresponding to each of the original video key frames;

[0068] Performing clustering processing on the plurality of original video feature data, and determining the video key frames and the visual feature data corresponding to the plurality of original video feature data from the plurality of original video key frames;

[0069] The acoustic feature data corresponding to the audio data is extracted using the MFCC algorithm.

[0070] In an embodiment of the present invention, exemplarily, in a data center, a server processes a technician's server maintenance operation video. In order to reduce the processing complexity while retaining the key information in the video, the server performs sparse sampling of the video data at a fixed time interval (such as one frame per second). In this way, the server extracts a series of representative original video key frames from a continuous video stream. The server then uses a deep learning model (such as a convolutional neural network) to perform deep feature extraction on each original video key frame. These features may include the technician's action posture, tool usage, facial expressions, etc. Through deep feature extraction, the server obtains the original video feature data corresponding to each original video key frame, which represents the visual information in the key frame in the form of a high-dimensional vector. In order to further streamline the data and highlight the key information, the server performs clustering processing on the extracted original video feature data. Through a clustering algorithm (such as K-means or hierarchical clustering), the server classifies similar feature data into one category and selects the most representative feature data from each category as the visual feature data of the category. At the same time, the original video key frames corresponding to these visual feature data are determined as the final video key frames. In this way, the server obtains a set of concise and representative video key frames and visual feature data. For audio data, the server uses the Mel Frequency Cepstral Coefficient (MFCC) algorithm to extract acoustic features. MFCC is an acoustic feature representation method widely used in fields such as speech recognition and music information retrieval. Through the MFCC algorithm, the server can extract features such as rhythm, timbre, and intensity from audio data, which represent the acoustic properties of audio signals in the form of a set of coefficients. These acoustic feature data will be used for subsequent operation recognition and evaluation. Through these steps, the server can extract key frames and visual feature data from video data, and extract acoustic feature data from audio data, providing effective feature support for subsequent skill operation evaluation and label annotation.

[0071] In the embodiment of the present invention, the aforementioned step S204 may be implemented through the following example execution.

[0072] Inputting the video key frames into a pre-trained object recognition model to obtain key frame categories;

[0073] Input the visual feature data into a pre-trained behavior recognition model to obtain the target object and visual related attributes corresponding to the video key frame, wherein the target object includes body parts, clothing and operating tools;

[0074] Performing matching and identification operations on the target object and visual related attributes to obtain the video evaluation result data;

[0075] Using automatic speech recognition technology to convert the acoustic feature data to obtain speech text content;

[0076] The speech text content is input into a pre-trained speech matching model, and a matching and recognition operation is performed with the standard speech text to obtain the audio evaluation result data.

[0077] In an embodiment of the present invention, for example, in the monitoring room of a data center, a server is processing a video of a technician performing server maintenance. The server first inputs the extracted video key frame into a pre-trained target recognition model. This model can identify the main content in the key frame. For example, it can determine whether the key frame shows the technician's hand operation, facial expression, or a close-up of the equipment, and classify this information into different key frame categories. Next, the server inputs the visual feature data into a pre-trained behavior recognition model. This model can analyze the feature data and identify the target objects in the video key frame, such as the technician's limbs (hands, feet, head, etc.), clothing (such as gloves, helmets, etc.) and the operating tools they are using (such as screwdrivers, cables, etc.). At the same time, the model can also extract the visual related attributes of these objects, such as color, shape, size, etc. The server performs matching and recognition operations based on the identified target objects and their visual attributes. For example, it will check whether the technician is wearing the prescribed safety equipment, whether the hand movements comply with the standard operating procedures, and whether the tools used are correct. These matching results will be integrated into video evaluation result data to evaluate the technical staff's operational standardization and safety. When processing audio data, the server first converts the extracted acoustic feature data using automatic speech recognition (ASR) technology. ASR technology can convert voice signals into text content, allowing the server to understand and analyze the voice instructions or communication content issued by the technician during the operation. Finally, the server inputs the converted voice text content into a pre-trained voice matching model. This model compares the voice text with the standard speech text to check whether the technician's voice instructions comply with the standardized operation speech. For example, it will determine whether the technician used the correct instructions when starting the equipment, or whether the problem was reported in accordance with the standard process. The matching results will be recorded as audio evaluation result data to evaluate the standardization of the technician's voice operation.

[0078] In the embodiment of the present invention, the aforementioned step S205 can be implemented through the following example execution.

[0079] Determining target evaluation result data from the video evaluation result data and the audio evaluation result data, wherein the target evaluation result data includes key frames, key frame categories, target objects of key frames, visual related attributes, matched voice text segments, matching scores, evaluation results, and error cause analysis;

[0080] According to preset rules and in combination with the start time frame and the end time frame corresponding to the target evaluation result data, the operation points of the human body skill operation and the video segment location corresponding to the operation unit are determined.

[0081] In an embodiment of the present invention, exemplarily, in the monitoring room of the data center, the server has completed the processing of the technician's operation video and audio, and obtained the video evaluation result data and the audio evaluation result data. Now, the server needs to filter out key information from these data, that is, the target evaluation result data. These data include key frames (images showing important operation moments), key frame categories (such as hand operations, tool use, etc.), target objects in key frames (such as hands, tools, etc.), visual related attributes of these objects (such as color, shape, etc.), voice text fragments matching key frames (voice instructions or descriptions of technicians during operation), matching scores (indicating the degree of association between voice text and key frame content), evaluation results (whether the operation is standardized, whether there is an error judgment) and error cause analysis (if the operation is wrong, analyze the possible reasons). For example, the server may filter out a key frame showing that the technician is using a screwdriver to tighten the screws on the server. The category of this key frame is "tool use", the target objects are "hands" and "screwdriver", and the visual related attributes include "hand action" is tightening, and "screwdriver color" is silver. At the same time, the matching voice text segment is "I am tightening this screw now", and the matching score is high, indicating that the voice text is highly correlated with the key frame content. The evaluation results show that this operation is standardized and there is no error cause analysis. The server will then determine the operation points of the technician's physical skill operation and the specific location of the operation unit in the video according to the preset rules and the start time frame and end time frame in the target evaluation result data. For example, the server may set the rule as "take 5 seconds before and after the key frame as the time range of the operation point", then according to this rule, the server can determine the precise positioning of the operation point of "tightening the screw" in the video, that is, starting from the first 5 seconds of the key frame and ending at the last 5 seconds of the key frame. In this way, the server successfully locates the video clip of the technician performing the operation point of "tightening the screw". Through the above steps, the server can accurately extract key operation information from a large amount of video and audio data, and accurately locate the specific location in the video, providing strong support for subsequent skill training, operation analysis and optimization.

[0082] In the embodiment of the present invention, the aforementioned step S206 can be implemented through the following example execution.

[0083] Determine the start and end times corresponding to the operation points and the operation units respectively based on the video segment positioning;

[0084] The operation key points and the operation units are labeled according to the start and end times in combination with a pre-trained content label library to obtain the pending label labeling result.

[0085] In an embodiment of the present invention, illustratively, on the server in the data center, the key frames and operation points in the technician's operation video have been determined through a series of processing, and the time positions of these operation points in the video have been accurately located. Now, the server needs to determine the specific start and end time of each operation point and operation unit based on these positioning information. For example, the server previously located the video clip of the operation point of "tightening the screws", and now it will record the start time and end time of this clip, such as from the 30th second to the 35th second of the video. Similarly, for other operation units, such as "checking the device status", "connecting cables", etc., the server will also determine their start and end times respectively. Next, the server will use a pre-trained content tag library to label these located operation points and operation units. This tag library contains various possible operation tags, such as "screwdriver operation", "cable connection", "equipment inspection", etc. Taking "tightening the screws" as an example, the server will find the corresponding clip in the video according to the start and end time of this operation point, and select the tag that best matches the content of this clip from the tag library for labeling. In this example, the server may label this clip with tags such as "screwdriver operation" and "tightening screws". Similarly, for other operation units, the server will select appropriate tags from the tag library for labeling based on their start and end time and video content. In the end, the server will obtain a pending labeling result containing all operation points and operation units. Through the above steps, the server can automatically accurately label each operation point and operation unit in the technician's operation video, which will help the subsequent evaluation and improvement of the technician's operation skills.

[0086] In the embodiments of the present invention, the following implementation modes are also provided.

[0087] Obtaining the offset time data of the operation unit and the start and end time of the operation key point determined during the review process of the pending labeling result;

[0088] Acquire comparison time data between the operation unit and the operation point according to a preset interval distance;

[0089] Determine a time point offset feature according to the offset time data and the comparison time data;

[0090] The time point offset feature is used as an optimization parameter to optimize the video segment positioning operation.

[0091] In an embodiment of the present invention, exemplarily, in the audit process of the data center, the server will collect feedback from the auditor on the pending labeling results. These feedbacks include the start and end times of the actual operation units and operation points that the auditor believes, which may deviate from the time initially located by the server. For example, for the operation point of "tightening the screws", the start and end time initially located by the server is from the 30th second to the 35th second of the video, while the auditor may believe that the actual operation time is from the 29th second to the 34th second. This time deviation, i.e., offset time data, is valuable optimization information for the server. In order to improve the accuracy of video clip positioning, the server will also subdivide the time of the operation units and operation points according to a preset interval distance (such as every 1 second or 0.5 seconds), and generate a series of comparative time data. These data are used for comparative analysis with the actual time feedback from the auditor to find out the regularity of the positioning deviation. The server will compare the offset time data feedback from the auditor with the comparative time data generated at a preset interval. By analyzing these data, the server can determine the characteristics of the time point offset. For example, the server may find that certain types of operations (such as fast hand movements) are prone to cause positioning time to be biased forward or backward, or that certain operations at specific locations in the video (such as the edge of the screen) are prone to inaccurate positioning, etc. Once the time point offset characteristics are determined, the server will use these characteristics as optimization parameters to improve its video segment positioning algorithm. For example, if the server finds that fast hand movements are prone to cause positioning time to be biased backward, it can add a corresponding compensation mechanism to the algorithm to automatically adjust the positioning time when similar movements are detected. Similarly, if operations at specific locations are prone to inaccurate positioning, the server can optimize its image processing algorithm to better identify and locate these operations. Through the above steps, the server can continuously learn and improve the accuracy of its video segment positioning, thereby providing more accurate annotation results of operation points and operation units. This is of great significance for subsequent skill assessment, training, and operation optimization.

[0092] In order to more clearly describe the solution provided by the embodiments of the present application, a relatively complete implementation scheme is provided below.

[0093] S1, stores the multi-channel video streaming data of an employee's skill training in real time and synchronously to the database.

[0094] S2, obtains multi-channel video streaming data of a skill training from the database and inputs it into the preprocessing module. The subdivision steps of S2 are as follows:

[0095] S2-1, synchronizes the multi-channel audio and video through the synchronization processor to ensure the synchronization of sound and picture and the alignment of the time points of multiple videos.

[0096] S2-2, use a video encoder to decode the video stream and separate the video frame data and the audio data.

[0097] S2-3, perform necessary preprocessing operations on the extracted video frame data, such as video denoising and image enhancement, and perform noise reduction and filtering on the audio data to improve the accuracy of subsequent feature extraction and recognition.

[0098] S3, the pre-processed video data and audio data are input into the feature extraction module. The subdivision steps of S3 are as follows:

[0099] S3-1, obtain the first set of video key frames, video features, and timestamp data through feature extraction and video sampling. Perform the following operations:

[0100] The first step is to perform sparse sampling based on a fixed interval to obtain a batch of original video key frames. The video sampling algorithm formula is: selected_frame_index = start_index + n*interval`, where `selected_frame_index` is the index of the selected frame, 'start_index` is the index of the starting frame, 'n` is an increasing integer representing the sequence number of the extracted frame, and 'interval` is the fixed interval between frames.

[0101] The second step is to extract video features through deep learning algorithms based on the first step.

[0102] The third step is to perform cluster analysis on the feature data of the original batch of video key frames based on the second step, and classify the original batch of video key frames; by calculating the differences between the key frames in the clusters, the representative key frames of each cluster are extracted as the first group of video key frames, video features, and timestamp data.

[0103] The clustering algorithm used for feature extraction is K-means. The K-means algorithm is a very popular clustering algorithm, which is usually used to divide data points into K non-overlapping subsets (or "clusters"). In the context of video key frame extraction, K-means can be used to group frames with similar visual features into the same cluster and select a representative frame from each cluster as a key frame. The basic steps of the K-means algorithm are as follows:

[0104] (1) Initialization: Select K points as the initial cluster centers (also called centroids). These points can be random points in the dataset.

[0105] (2) Assignment step: For each point in the data set, assign it to the cluster where the nearest centroid is located based on its distance to the K centroids. The Euclidean distance is usually used to calculate the distance between a point and a centroid.

[0106] (3) Update step: For each cluster, calculate the average of all points assigned to the cluster and set the average as the new centroid.

[0107] (4) Iteration: Repeat steps 2 and 3 until the centroid no longer changes significantly or the predetermined number of iterations is reached.

[0108] (5) Output: At the end of the algorithm, each data point is assigned to a cluster, and each cluster has a centroid. In the context of keyframe extraction, one frame from each cluster (e.g., the frame closest to the centroid) can be selected as the keyframe.

[0109] The K-means algorithm itself does not have a single mathematical formula, in which the similarity between data points is mainly measured by the calculation formula of the Euclidean distance. For two points p and q in the data set, the Euclidean distance between them can be calculated by the following formula:

[0110] d(p,q)=sqrt((p1-q1)^2+(p2-q2)^2+...+(pn-qn)^2)

[0111] Among them, p1, p2, ..., pn and q1, q2, ..., qn are the coordinate values ​​of points p and q in each dimension respectively, and n is the dimension of the point data (in video key frame extraction, this usually corresponds to the length of the feature vector).

[0112] S3-2, extract the first group of audio acoustic feature data through the MFCC (Mel Frequency Cepstrum Coefficient) algorithm.

[0113] S4, inputting the first group of video key frames and feature data and the first group of audio acoustic feature data into the model matching and recognition module respectively, and obtaining video feature evaluation result data and audio feature evaluation result data respectively.

[0114] S4-1, input the first group of video key frames and feature data into the model matching and recognition module, and first perform target detection and then perform pre-trained model inference through model scheduling of the pre-trained content tag library.

[0115] S4-2, the first group of video key frames and feature data, uses machine learning algorithms to perform pre-trained model inference, and performs target recognition and behavior recognition to obtain matching key frame categories (there are hundreds of operation points, here referring to target categories and behavior categories, etc.), key frame target objects and attributes (there are three types of target objects identified in key frames, namely limb parts, clothing, operating tools, etc.) and other evaluation results and other data.

[0116] S4-3, converting the first group of audio acoustic feature data into speech text content in real time through speech recognition technology (ASR).

[0117] S4-4, a speech matching model based on a pre-trained content tag library, compares the speech text content with the standard speech text, and obtains audio evaluation result data through text matching, natural language processing (NLP), semantic analysis and other technologies.

[0118] S5, the evaluation result data returned by the model matching and recognition module is input into the video positioning analysis module to obtain the video segment positioning of the operation points and the video segment positioning of the operation units.

[0119] S5-1, the evaluation result data returned by the model matching and recognition module is input into the video positioning analysis module and associated with the video frame timestamp.

[0120] Among them, the evaluation result data includes: key frames, key frame categories (such as target categories, behavior categories, etc.), target objects and attributes of key frames (target objects such as limbs, clothing, operating tools, etc.), matching voice text fragments, matching scores, evaluation results (recognition results include: complete match, partial match, mismatch), error cause analysis (such as: causes of incorrect use of tools, causes of incorrect clothing, speech content that has been spoken / omitted in the voice text), etc.

[0121] S5-2, combining the above S4 and S5-1, through the operation point positioning analysis module, determine the start and end time of the video clip of the operation point. The following two cases are processed:

[0122] Case 1 (focusing on target objects and operation behavior, mainly based on video feature processing)

[0123] 1. Taking an operation point as an example, when a "completely matching" target or behavior (one-to-one corresponding to the operation point) is identified through the results returned by the S4 process, the process is as follows:

[0124] (1) The neighboring keyframes of the matched keyframe are sampled based on the neighboring sampling algorithm to obtain the second set of video keyframes, video features, and timestamp data. The neighboring sampling algorithm is usually used to further sample near the selected keyframe to obtain more relevant or continuous information. When processing video data, neighboring sampling can help capture subtle changes in key actions or events.

[0125] (2) The second group of video key frames and video feature data are input into the S4 model matching and recognition module to obtain the recognition evaluation results.

[0126] (3) Based on the principle of identical content matching, key frames with a high matching degree are extracted in a targeted manner to obtain the third group of video key frames and video feature data.

[0127] (4) Repeat steps (1) to (3) above until the content is found to be irrelevant (i.e., the collected keyframe does not belong to the target category or behavior category), then terminate this loop process.

[0128] (5) The third group of video key frames are sorted according to timestamps, and the threshold range of the time interval between key frames is considered as a condition, and the key frames returned by the above process link (4) are grouped. The earliest key frame in each group of video key frames is the start time of the operation point, and the latest key frame is the end time of the operation point.

[0129] 2. Taking an operation point as an example, when the results returned by the S4 process cannot identify a completely matching target or behavior (partial match), but can find a target or behavior with a high degree of match, the processing method is as above.

[0130] The matching degree of an operation point is determined by multiple categories of target objects (body parts, clothing, operating tools, etc.) and their related distance relationships: when the matching degree is within a certain threshold range, it is regarded as valuable content (i.e., the employee made an operation error, resulting in partial correctness and partial errors), and needs to be processed according to the above "1" method; when the matching degree is not within the threshold range, it is regarded as worthless content (i.e., the employee did not operate or missed an operation).

[0131] Case 2 (for standard speech operation points, mainly based on audio feature processing):

[0132] 1. Taking an operation point as an example, when a "completely matched" target or behavior (one-to-one corresponding to the operation point) is identified through the results returned by the S4 process, the start and end time of the "matched voice text segment" obtained in the S5-1 link is used as the start and end time of the video segment of the operation point.

[0133] 2. Taking an operation point as an example, when the result returned by the S4 process cannot identify a completely matching voice text segment (partial match, partial non-match), the start and end time of the "matched voice text segment" obtained in the S5-1 link (the text segment here is complete, including voice text segments that completely match the standard speech text segment, and also including text segments that do not completely match the standard speech text segment) is used as the start and end time of the video segment of the operation point.

[0134] S5-3, based on step S5-2, the start and end time of the video clip of the operation point is obtained as input, and the start and end time of the video clip of the operation unit is obtained by combining the association between the predefined operation unit and the operation point. The operation unit contains 1-n operation points. In one operation unit, the earliest start time of the operation point contained in it is used as the start time of the operation unit; the latest end time of the operation point contained in it is used as the end time of the operation unit.

[0135] S6, based on the start and end time of each operation unit returned by S5, combined with the content tags in the pre-trained content tag library (where the operation units correspond to the content tag names one by one), realizes the function of intelligent video tagging: by accurately locating the start and end time of each operation unit, and labeling the video clip content with the corresponding operation unit name.

[0136] S7 is the review module: based on the operation unit video positioning, operation unit content labeling, and operation point positioning (including operation point labeling, the implementation method is the same as S6) information returned by S6, the front-end display is based on the video content automatic labeling platform for manual review by content labeling personnel: During the review process, the start and end time of the video of the operation unit and the operation point can be located through hierarchical content labels; the video can be played through the player to confirm whether the corresponding content is the operation unit and the operation point; the control can be captured through the start and end time range in the editing state to adjust the video clip positioning of the operation unit and the operation point.

[0137] S8, based on the offset data of each operation unit and the start and end time of the operation key points collected by the S7 audit module, input it into the intelligent correction analysis module for deviation feature analysis to obtain the offset parameter D:

[0138] The first step is to combine the start and end time points of the operation units and operation points video clips output by the original process link S5-2 with the start and end time points of the operation units and operation points obtained in the S7 link after the review link adjustment.

[0139] The second step is to calculate the interval distance d of each pair of operation units and operation key points to supplement the video key frames forward or backward. The definition of a pair of data is to form a data pair based on the start and end time output by S5-2 and the start and end time output by S7.

[0140] The third step is to extract the start and end time point offset feature D (D is the fixed interval between video frames) through a deep learning algorithm.

[0141] In the fourth step, the offset parameter D is input into the S5-2 link as an offset parameter to adjust the parameters of adjacent sampling, thereby realizing a more intelligent and precise positioning model for video intelligent labeling.

[0142] The complete system process of S1-S8 above is the full process of the method for labeling video content based on human body skill operation.

[0143] Please refer to Figure 2 , Figure 2 A content annotation device 110 based on a human body skill operation video provided in an embodiment of the present invention includes:

[0144] The acquisition module 1101 is used to acquire multiple channels of video streaming media data containing human body skill operations; pre-process the multiple channels of video streaming media data to obtain pre-processed video data and audio data;

[0145] The execution module 1102 is used to perform feature extraction on the video data and the audio data respectively to obtain video key frames and visual feature data corresponding to the video data, and acoustic feature data corresponding to the audio data; perform matching and recognition operations on the video key frames and the visual feature data, and the acoustic feature data respectively to obtain video evaluation result data and audio evaluation result data; determine the operation points of the human body skill operation and the video segment location corresponding to the operation unit according to the video evaluation result data and the audio evaluation result data;

[0146] The labeling module 1103 is used to label the operation points and the operation units based on the positioning of the video clip and in combination with the pre-trained content label library to obtain a pending labeling result; if the pending labeling result passes the review, the pending labeling result is used as the target labeling result for the human body skill operation.

[0147] It should be noted that the implementation principle of the aforementioned content annotation device 110 based on the human body skill operation video can refer to the implementation principle of the aforementioned content annotation method based on the human body skill operation video, which will not be repeated here. It should be understood that the division of the various modules of the above device is only a division of logical functions. In actual implementation, they can be fully or partially integrated into a physical entity, or physically separated. And these modules can all be implemented in the form of software called by processing elements; they can also be all implemented in the form of hardware; some modules can also be implemented in the form of software called by processing elements, and some modules can be implemented in the form of hardware. For example, the content annotation device 110 based on the human body skill operation video can be a separately established processing element, or it can be integrated in a chip of the above device. In addition, it can also be stored in the memory of the above device in the form of program code, and called and executed by a processing element of the above device. The function of the content annotation device 110 based on the human body skill operation video. The implementation of other modules is similar. In addition, all or part of these modules can be integrated together, or they can be implemented independently. The processing element described here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each module above can be completed by an integrated logic circuit of hardware in a processor element or an instruction in software form.

[0148] For example, the above modules may be one or more integrated circuits configured to implement the above methods, such as one or more application specific integrated circuits (ASIC), or one or more microprocessors (digital signal processors, DSP), or one or more field programmable gate arrays (FPGA), etc. For another example, when a module above is implemented in the form of a processing element scheduling program code, the processing element may be a general-purpose processor, such as a central processing unit (CPU) or other processor that can call program code. For another example, these modules may be integrated together and implemented in the form of a system-on-a-chip (SOC).

[0149] The embodiment of the present invention provides a computer device 100, which includes a processor and a non-volatile memory storing computer instructions. When the computer instructions are executed by the processor, the computer device 100 executes the aforementioned content annotation device 110 based on the human body skill operation video. Figure 3 As shown, Figure 3The computer device 100 provided in the embodiment of the present invention is a structural block diagram. The computer device 100 includes a content annotation device 110 based on a human body skill operation video, a memory 111, a processor 112 and a communication unit 113.

[0150] To achieve data transmission or interaction, the memory 111, the processor 112 and the communication unit 113 are electrically connected to each other directly or indirectly. For example, these components can be electrically connected to each other through one or more communication buses or signal lines. The content annotation device 110 based on the human body skill operation video includes at least one software function module that can be stored in the memory 111 in the form of software or firmware or solidified in the operating system (OS) of the computer device 100. The processor 112 is used to execute the content annotation device 110 based on the human body skill operation video stored in the memory 111, such as the software function modules and computer programs included in the content annotation device 110 based on the human body skill operation video.

[0151] An embodiment of the present invention provides a readable storage medium, which includes a computer program. When the computer program is running, it controls the computer device where the readable storage medium is located to execute the aforementioned content labeling method based on human body skill operation video.

[0152] For illustrative purposes, the foregoing description is made with reference to specific embodiments. However, the above illustrative discussion is not intended to be exhaustive or to limit the present disclosure to the precise form disclosed. Numerous modifications and variations are possible in accordance with the above teachings. These embodiments are selected and described in order to best illustrate the principles of the present disclosure and its practical application, so that those skilled in the art can best utilize the present disclosure and utilize various embodiments with different modifications to suit the intended specific application.

Claims

1. A content annotation method based on human body skill operation video, characterized in that: include: Acquire multi-channel video streaming data containing human body skill operations; Preprocessing the multi-channel video streaming media data to obtain preprocessed video data and audio data; Extracting features from the video data and the audio data respectively to obtain video key frames and visual feature data corresponding to the video data, and acoustic feature data corresponding to the audio data; Performing matching and recognition operations on the video key frame and the visual feature data, as well as the acoustic feature data, respectively, to obtain video evaluation result data and audio evaluation result data; Determine the operation points of the human body skill operation and the video segment location corresponding to the operation unit according to the video evaluation result data and the audio evaluation result data; Based on the video clip positioning, the operation points and the operation units are labeled in combination with a pre-trained content label library to obtain a pending label labeling result; If the pending label marking result passes the review, the pending label marking result is used as the target marking result for the human body skill operation; The step of determining the operation points of the human body skill operation and the video segment location corresponding to the operation unit according to the video evaluation result data and the audio evaluation result data includes: Determine target evaluation result data from the video evaluation result data and the audio evaluation result data, the target evaluation result data including key frames, key frame categories, target objects of key frames, visual related attributes, matched voice text segments, matching scores, evaluation results and error cause analysis; the key frames are images showing important operation moments, the key frame categories include hand operations and tool use, the target objects in the key frames include hands and tools, the visual related attributes include color and shape, the matched voice text segments are matched with the key frames, the matched voice text segments are voice instructions or descriptions of operations, the matching scores are the degree of association between the voice text and the key frame content, the evaluation results include the judgment results of whether the operation is standardized and whether there are errors, and the error cause analysis includes analyzing the reasons for the incorrect operation; According to preset rules and in combination with the start time frame and the end time frame corresponding to the target evaluation result data, the operation points of the human body skill operation and the video segment location corresponding to the operation unit are determined.

2. The method according to claim 1, characterized in that The preprocessing of the multi-channel video streaming media data to obtain preprocessed video data and audio data includes: The synchronous processor synchronously processes the multi-channel video streaming media data, and decodes the synchronously processed multi-channel video streaming media data to obtain initial video data and initial audio data; De-noising and image enhancement are performed on the initial video data to obtain the pre-processed video data; The initial audio data is denoised and filtered to obtain the preprocessed audio data.

3. The method according to claim 1, characterized in that The extracting features of the video data and the audio data respectively to obtain video key frames and visual feature data corresponding to the video data, and acoustic feature data corresponding to the audio data, includes: Sparsely sampling the video data at fixed intervals to obtain a plurality of original video key frames; Performing deep feature extraction on the multiple original video key frames to obtain original video feature data corresponding to each of the original video key frames; Performing clustering processing on the plurality of original video feature data, and determining the video key frames and the visual feature data corresponding to the plurality of original video feature data from the plurality of original video key frames; The acoustic feature data corresponding to the audio data is extracted using the MFCC algorithm.

4. The method according to claim 1, characterized in that: The matching and identifying operations are performed on the video key frame, the visual feature data, and the acoustic feature data to obtain video evaluation result data and audio evaluation result data, including: Inputting the video key frames into a pre-trained object recognition model to obtain key frame categories; Input the visual feature data into a pre-trained behavior recognition model to obtain the target object and visual related attributes corresponding to the video key frame, wherein the target object includes body parts, clothing and operating tools; Performing matching and identification operations on the target object and visual related attributes to obtain the video evaluation result data; The acoustic feature data is converted using automatic speech recognition technology to obtain speech text content; The speech text content is input into a pre-trained speech matching model, and a matching and recognition operation is performed with the standard speech text to obtain the audio evaluation result data.

5. The method according to claim 1, characterized in that: The step of labeling the operation points and the operation units based on the video segment positioning and combining the pre-trained content label library to obtain a pending label labeling result includes: Determine the start and end times corresponding to the operation points and the operation units respectively based on the video segment positioning; The operation key points and the operation units are labeled according to the start and end times in combination with a pre-trained content label library to obtain the pending label labeling result.

6. The method according to claim 1, characterized in that The method further comprises: Obtaining the offset time data of the start and end time of the operation unit and the operation key point determined during the review process of the pending labeling result; Acquire comparison time data between the operation unit and the operation point according to a preset interval distance; Determine a time point offset feature according to the offset time data and the comparison time data; The time point offset feature is used as an optimization parameter to optimize the video segment positioning operation.

7. A content annotation device based on a human body skill operation video, characterized in that: include: An acquisition module, used for acquiring multi-channel video streaming media data containing human body skill operations; Preprocessing the multi-channel video streaming media data to obtain preprocessed video data and audio data; An execution module, configured to perform feature extraction on the video data and the audio data respectively, to obtain video key frames and visual feature data corresponding to the video data, and acoustic feature data corresponding to the audio data; Perform matching and recognition operations on the video key frame and the visual feature data, as well as the acoustic feature data, to obtain video evaluation result data and audio evaluation result data; determine the operation points of the human body skill operation and the video segment location corresponding to the operation unit according to the video evaluation result data and the audio evaluation result data; A labeling module is used to label the operation points and the operation units based on the positioning of the video clip and in combination with a pre-trained content label library to obtain a pending labeling result; if the pending labeling result passes the review, the pending labeling result is used as a target labeling result for the human body skill operation; The execution module is specifically used for: Determining target evaluation result data from the video evaluation result data and the audio evaluation result data, wherein the target evaluation result data includes key frames, key frame categories, target objects of key frames, visual related attributes, matched voice text segments, matching scores, evaluation results, and error cause analysis; The key frame is an image showing an important operation moment, the key frame category includes hand operation and tool use, the target objects in the key frame include hands and tools, the visual related attributes include color and shape, the matched voice text segment is matched with the key frame, the matched voice text segment is a voice instruction or description of the operation, the matching score is the degree of association between the voice text and the key frame content, the evaluation result includes a judgment result on whether the operation is standardized and whether there is an error, and the error cause analysis includes analyzing the cause of the incorrect operation; according to the preset rules, combined with the start time frame and the end time frame corresponding to the target evaluation result data, the operation points of the human body skill operation and the video segment positioning corresponding to the operation unit are determined.

8. A computer device, characterized in that: The computer device comprises a processor and a non-volatile memory storing computer instructions. When the computer instructions are executed by the processor, the computer device executes the method according to any one of claims 1 to 6.

9. A readable storage medium, characterized in that: The readable storage medium includes a computer program, and when the computer program is executed, the computer device where the readable storage medium is located is controlled to execute the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Voice and video positioning model and construction method, device and application thereof

    CN115359398A

  • Video labeling method and device, computer equipment and computer readable storage medium

    CN117710845A