Video processing methods, apparatus, electronic devices and storage media

By identifying moments of hand-raising and audio segments in teaching videos, and combining image and audio data to assess teaching quality, the problem of assessment errors in existing technologies has been solved, achieving a more accurate and objective assessment of teaching quality.

CN117173615BActive Publication Date: 2026-04-03BEIJING WISDOM RONGSHENG TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-27
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

In existing technologies, the use of image detection to assess teaching quality is prone to errors due to chance.

Method used

By acquiring multiple frames of images and audio data from teaching videos, the system identifies the moment a hand is raised and the audio segment, and combines this with video clips of the hand-raising response to assess teaching quality.

Benefits of technology

It improves the accuracy and effectiveness of teaching assessment, achieves objective evaluation of the teaching process, and helps to improve the quality of education and teaching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117173615B_ABST
    Figure CN117173615B_ABST
Patent Text Reader

Abstract

This invention discloses a video processing method, apparatus, electronic device, and storage medium. The method includes: extracting multiple frames of images to be recognized and audio data from a teaching video; for each of the multiple frames, if the recognition result of the image is a preset action, determining the hand-raising moment corresponding to the image; performing recognition processing on the audio data to obtain an audio segment of a target object corresponding to a preset role, and determining the starting audio moment of the audio segment; based on the hand-raising moment and the starting audio moment, determining a hand-raising response video segment from the teaching video, and evaluating the teaching quality of the teaching video based on the hand-raising response video segment. This solves the problem of evaluation errors caused by image detection in existing technologies, improves the accuracy of determining the hand-raising response video segment, and thus improves the effectiveness of teaching evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer processing technology, and in particular to a video processing method, apparatus, electronic device, and storage medium. Background Technology

[0002] With the development of information technology in education, monitoring teaching quality has become increasingly important. It is necessary to monitor and analyze students' classroom status during class, and then use the analysis results to standardize teaching.

[0003] In existing technologies, teaching monitoring typically involves capturing videos of students attending class, then detecting and analyzing the raised hands in these videos to assess student engagement and determine the effectiveness of the lesson. However, various unpredictable factors exist during class, and image-based detection methods can lead to errors in teaching assessment. Summary of the Invention

[0004] This invention provides a video processing method, apparatus, electronic device, and storage medium to improve the accuracy of determining hand-raising responses via video, thereby enhancing the accuracy and effectiveness of teaching assessment.

[0005] According to one aspect of the present invention, a video processing method is provided, the method comprising:

[0006] The teaching video to be extracted is obtained, and multiple frames of images and audio data to be recognized are extracted from the teaching video.

[0007] For the multiple frames of images to be identified, if the identification result of the image to be identified is a preset action, then the hand-raising time corresponding to the image to be identified is determined;

[0008] The voice data is processed to obtain a voice segment of the target object corresponding to a preset role, and the start time of the voice segment is determined.

[0009] Based on the time of raising hands and the time of starting the speech, video segments of raising hands and answering are determined from the teaching video to evaluate the teaching quality of the teaching video based on the video segments of raising hands and answering.

[0010] According to another aspect of the present invention, a video processing apparatus is provided, the apparatus comprising:

[0011] The extraction module is used to acquire the teaching video to be extracted, and to extract multiple frames of images to be recognized and audio data from the teaching video;

[0012] The image recognition module is used to determine the hand-raising moment corresponding to the image to be recognized if the recognition result of the image to be recognized is a preset action for the multiple frames of images to be recognized.

[0013] The speech recognition module is used to recognize and process the speech data to obtain a speech segment of the target object corresponding to a preset role, and to determine the start time of the speech segment.

[0014] The video segment determination module is used to determine the hand-raising and answering video segments from the teaching video based on the hand-raising time and the start voice time, so as to evaluate the teaching quality of the teaching video based on the hand-raising and answering video segments.

[0015] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0016] At least one processor; and

[0017] A memory communicatively connected to the at least one processor; wherein,

[0018] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the video processing method according to any embodiment of the present invention.

[0019] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the video processing method according to any embodiment of the present invention.

[0020] The technical solution of this invention involves acquiring a teaching video to be extracted and extracting multiple frames of images to be recognized and audio data from the video. For each frame of images to be recognized, if the recognition result is a preset action, the hand-raising moment corresponding to the image is determined. The audio data is processed to obtain audio segments of the target object corresponding to a preset role, and the starting audio moment of the audio segments is determined. Based on the hand-raising moment and the starting audio moment, hand-raising and answering video segments are determined from the teaching video. The teaching quality of the teaching video is evaluated based on these hand-raising and answering video segments. This solves the problem of evaluation errors caused by image detection in existing technologies. By performing action recognition on multiple frames of images to be recognized in the teaching video and audio recognition on each role in the video, and combining the hand-raising action in the image with the target object's response in the audio, hand-raising and answering video segments that show both hand-raising action and response are comprehensively determined. The teaching video is then evaluated using these hand-raising and answering video segments, making the evaluation of the teaching process more effective and objective, and thus improving the quality of education and teaching.

[0021] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 This is a flowchart of a video processing method provided according to Embodiment 1 of the present invention;

[0024] Figure 2 This is a flowchart of a video processing method provided according to Embodiment 2 of the present invention;

[0025] Figure 3 This is a schematic diagram of the structure of module C3 provided in Embodiment 2 of the present invention;

[0026] Figure 4 This is a schematic diagram of the backbone network structure provided in Embodiment 2 of the present invention.

[0027] Figure 5 This is a schematic diagram of the structure of a video processing device according to Embodiment 3 of the present invention;

[0028] Figure 6 This is a schematic diagram of the structure of an electronic device that implements the video processing method of the present invention. Detailed Implementation

[0029] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0030] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0031] Example 1

[0032] Figure 1 This is a flowchart of a video processing method according to Embodiment 1 of the present invention. This embodiment is applicable to teaching analysis. The method can be executed by a video processing device, which can be implemented in hardware and / or software and can be configured in a computing device. Figure 1 As shown, the method includes:

[0033] S110. Obtain the teaching video to be extracted, and extract multiple frames of images and audio data to be recognized from the teaching video.

[0034] The teaching videos can be student class videos. It should be noted that in practical applications, the videos are played frame by frame, and each video frame is an image. Multiple frames can be extracted from the teaching videos as images to be recognized.

[0035] In this embodiment, each video frame in the teaching video can be used as the image to be recognized. Alternatively, a frame extraction rule can be preset to extract multiple frames of students during class from the teaching video as images to be recognized. Optionally, the frame extraction rule can be to extract frames at intervals according to a preset extraction frequency, such as five frames per second, or collecting video frames every three video frames as images to be recognized, or selecting images to be recognized based on a timeline. Staff can determine how to collect the images to be recognized based on actual needs. The audio in the teaching video is used as audio data to identify segments of speech by different characters.

[0036] S120. For multiple frames of images to be recognized, if the recognition result of the images to be recognized is a preset action, then determine the hand-raising time corresponding to the images to be recognized.

[0037] The preset actions can be set according to the needs of teaching assessment. For example, they can be actions such as raising a hand, looking down, looking up, or sleeping, to assess classroom teaching by recognizing students' actions. It should be noted that the processing method for each frame of the image to be recognized is the same, and the explanation can be given by taking the processing of any one of the images to be recognized as an example.

[0038] In this embodiment, image recognition technology can be used to perform action recognition on the image to be recognized, identifying whether a preset action exists in the image. If it exists, the moment the image to be recognized appears in the video can be recorded as the hand-raising moment; if it does not exist, the image to be recognized is skipped, and the next frame of the image to be recognized is proceeded. Accordingly, the hand-raising moments of all images to be recognized that exhibit the preset action in the teaching video can be determined. For example, frames are extracted from a teaching video, five frames per second, and hand-raising detection is performed on the extracted images to be recognized, identifying the images with hand-raising actions and determining the time point of the image in the teaching video, i.e., the hand-raising moment.

[0039] It should be noted that in practical applications, there may be situations where adjacent video frames have high similarity. To improve data processing efficiency, the images to be identified in consecutive frames where the recognition result is a preset action can be filtered out, and the hand-raising moment of the filtered images can be determined. For example, if the recognition result of three consecutive frames of images to be identified is a preset action, the first frame of the image to be identified can be selected, and the hand-raising moment of that frame can be determined. This hand-raising moment can be considered as the starting voice moment of raising the hand. The hand-raising moments of the other two frames of the image to be identified can be ignored to reduce the amount of data processing.

[0040] S130. The speech data is processed to obtain the speech segment of the target object corresponding to the preset role, and the start time of the speech segment is determined.

[0041] The preset roles include those who raise their hands to answer; correspondingly, the target audience for each preset role can be different students.

[0042] In practical applications, speech recognition technology can be used to process and separate different roles within the speech data, such as the person raising their hand to answer and the speaker. For example, the speaker's voice can be pre-set, and then compared with the speech data in the teaching video. Speech data matching the speaker's voice is identified as speech information for non-preset roles, while speech data not matching the speaker's voice is identified as speech information for preset roles, thus obtaining speech segments corresponding to the target object. Furthermore, the starting time of each speech segment is determined, reflecting the start time of the target object's speech in that speech segment.

[0043] To improve the versatility of video analysis, role separation technology can be used to separate different users in the voice data, such as user A, user B, and user C, to determine the voice information corresponding to different users. By comparing the voice information of different users, it can be determined which users are raising their hands to answer questions and which users are lecturing, so as to obtain the voice segments of the target objects corresponding to the preset roles.

[0044] In this embodiment, the voice data is subjected to role separation processing to obtain the voice segments of the target objects corresponding to preset roles. This includes: separating the voice data based on a role separation algorithm to determine the total voice duration corresponding to different objects; determining the target objects corresponding to preset roles from different objects based on the total voice duration; and determining the voice segments of the target objects.

[0045] Specifically, role separation algorithms can be used to identify different objects in the audio data of teaching videos, such as students or teachers. The total audio duration of each object is calculated, and based on this total duration, the object can be determined as either a student or a teacher. The object with the longest total audio duration is identified as the teacher, and the rest as students. Students are the target objects corresponding to the preset roles, and their audio segments can be identified separately. Alternatively, the audio duration of each object in each audio segment can be calculated, and the object with the longest single audio segment is identified as the teacher, and the rest as students.

[0046] It should be noted that S120 to S130 can be executed sequentially or in parallel. The specific execution order is not limited. The above order is only the order in which the technical solutions in each step are explained, not the execution order of each step.

[0047] S140. Based on the time of raising hands and the time of starting speech, identify video segments of raising hands and answering from the teaching video to evaluate the teaching quality of the teaching video.

[0048] It should be noted that when the target subject raises their hand, it may be in different situations, such as answering a question, going to the restroom, or simply raising their hand without any further action. Therefore, the image to be identified, which contains images of pre-defined actions, may include hand-raising in different situations. To accurately determine the video segment where the target subject is answering a question, the moment the hand is raised in the teaching video can be combined with the start time of the target subject's speech in each audio segment to extract the video segment of the hand-raising and answering situation in the teaching video.

[0049] In this embodiment, determining the hand-raising and answering video segment from the teaching video based on the hand-raising time and the start voice time includes: if a matching start voice time exists within a preset time after the hand-raising time, then the hand-raising and answering video segment is determined based on the hand-raising time and the end voice time of the voice segment to which the start voice time belongs.

[0050] The preset duration can be determined based on the actual work situation, such as 5 seconds or 6 seconds, and there is no limit to the comparison.

[0051] Specifically, for each hand-raising moment, a preset time interval is used to check if a starting voice moment exists. If it exists, it means a student spoke after the hand-raising action appeared in the teaching video; if it doesn't exist, it means no student spoke after the hand-raising action appeared in the teaching video. For example, if the hand-raising moment is 1.1 seconds, meaning the hand-raising action occurs at 1.1 seconds, and a starting voice moment of 3.1 seconds is found within 5 seconds after 1.1 seconds, it means a student spoke. Given that a matching starting voice moment exists within the preset time interval after the hand-raising moment, this hand-raising moment can be used as the start time of the hand-raising answer video segment, and the end time of the voice segment to which the starting voice moment belongs can be used as the end time of the hand-raising answer video segment. The video segment from this start time to this end time is then extracted from the teaching video as the hand-raising answer video segment. By integrating all the hand-raising answer video segments in the teaching video, the final student answer segment can be obtained.

[0052] Based on the above scheme, the teaching videos can also be evaluated based on the final student response segments to assess the quality of the teaching videos and the teaching situation, so as to objectively evaluate the teacher's teaching process, obtain evaluation results, and improve teaching quality through the evaluation results.

[0053] The technical solution of this embodiment acquires a teaching video to be extracted and extracts multiple frames of images to be recognized and audio data from the teaching video. For multiple frames of images to be recognized, if the recognition result of the image to be recognized is a preset action, the hand-raising moment corresponding to the image to be recognized is determined. The audio data is processed to obtain the audio segment of the target object corresponding to the preset role, and the starting audio moment of the audio segment is determined. Based on the hand-raising moment and the starting audio moment, the hand-raising and answering video segment is determined from the teaching video. The teaching quality of the teaching video is evaluated based on the hand-raising and answering video segment. This solves the problem of evaluation errors caused by image detection in the prior art. It realizes that by performing action recognition on multiple frames of images to be recognized in the teaching video, and performing audio recognition on each role in the teaching video, the hand-raising action in the image and the answer of the target object in the audio are combined to comprehensively determine the hand-raising and answering video segment where the hand-raising action occurs and there is an answer. Then, the teaching video is evaluated based on the hand-raising and answering video segment, making the evaluation of the teaching process more effective and objective, which is conducive to improving the quality of education and teaching.

[0054] Example 2

[0055] Figure 2 This is a flowchart of a video processing method according to Embodiment 2 of the present invention. Based on the foregoing embodiments, a target hand-raising detection model can be pre-trained by acquiring training samples, so that the target hand-raising detection model can be used to process the image to be recognized, thereby obtaining the recognition result of the image to be recognized. Specific implementation methods can be found in the technical solution of this embodiment. Technical terms that are the same as or corresponding to those in the above embodiments will not be repeated here.

[0056] like Figure 2 As shown, the method specifically includes the following steps:

[0057] S210. Obtain multiple training samples.

[0058] The training samples include images to be trained and corresponding annotations. The images to be trained can be pictures of students in class, and the annotations can be pre-annotated on the images to be trained, for example, annotating the images to be trained to identify students raising their hands.

[0059] To improve the accuracy of the model, as many training samples as possible can be obtained. Based on a large number of training samples, the hand-raising features can be extracted, the model can be trained, and a target hand-raising detection model can be obtained.

[0060] S220. A target hand-raising detection model is obtained by training multiple training samples.

[0061] In this embodiment, the method for training a target hand-raising detection model based on multiple training samples can be as follows: input the images to be trained from multiple training samples into the hand-raising detection model to be trained; process each group of images to be trained based on the data augmentation model in the hand-raising detection model to obtain a stitched image corresponding to each group of images to be trained; for each stitched image, process the stitched image based on the backbone network in the hand-raising detection model to obtain an output label; determine the loss value based on the annotation results in the stitched image and the output label; correct the model parameters in the hand-raising detection model to be trained based on the loss value; and take the convergence of the loss function of the hand-raising detection model to be trained as the training objective to obtain the target hand-raising detection model.

[0062] The training image set includes a preset number of training images; the number of annotations in the stitched image corresponds to the preset number. For example, the preset number could be 4. The loss function could be the Mish function.

[0063] In this embodiment, all training images from the training samples can be input into the hand-raising detection model. The data augmentation model in the hand-raising detection model is used as the input. The data augmentation model can randomly select a preset number of training images as a group of training images using mosaic data augmentation. Random deformations (such as scaling, translation, rotation, symmetry, cropping, etc.) are applied to the images within each group. The deformed images within the group are then stitched together to obtain the stitched image corresponding to each group of training images. For example, four training images are randomly selected, randomly deformed, and then randomly stitched together to obtain a stitched image. The advantages of using mosaic data augmentation are: it can add many small targets, enriching the dataset and improving the robustness of the model; at the same time, the random stitching method allows one image to calculate the data of four images, reducing the number of images per batch and reducing GPU usage; and by cropping the target to be identified, the model can identify the target based on local features, which helps in the detection of occluded targets and improves the model's detection capability. Furthermore, after obtaining the stitched images, they can be input into the backbone network of the hand-raising detection model to be trained. Hand-raising recognition detection is then performed on the stitched images to obtain the hand-raising recognition results, which serve as the output labels. Further, the loss function of the hand-raising detection model to be trained can be used to process the labeling results of the stitched images and the output labels, thereby calculating the loss value between them. For example, the loss can be calculated based on the difference between the predicted bounding boxes and the ground truth bounding boxes in the stitched images, and the model parameters in the model can be corrected based on the loss value. Alternatively, the calculated loss values ​​corresponding to all stitched images can be fused to obtain a fused loss value, which can be used to correct the model parameters in the model. The convergence of the loss function is used as the training objective. When the loss function is determined to converge, it indicates that this model can be used as the target hand-raising detection model. For example, the training error of the loss function, i.e., the loss parameter, can be used as a condition to detect whether the loss function has reached convergence, such as whether the training error is less than a preset error, whether the error change trend is stable, or whether the current number of iterations is equal to a preset number. If the detection meets the convergence criteria, such as the training error of the loss function being less than a preset error or the error change stabilizing, it indicates that the hand-raising detection model to be trained has completed training, and iterative training can be stopped at this point. If the convergence criteria have not yet been met, further training samples can be obtained to continue training the hand-raising detection model until the training error of the loss function is within a preset range. When the training error of the loss function converges, the hand-raising detection model to be trained can be used as the target hand-raising detection model, so that subsequent processing of the image to be recognized can be based on the target hand-raising detection model to obtain the recognition result of the image to be recognized.

[0064] In this embodiment, the backbone network includes a first module and a second module; the first module consists of multiple convolutional layers, multiple C3 modules, and an SPPF module. The stitched image is processed based on the backbone network of the hand-raising detection model to obtain output labels, including: processing the stitched image based on the first module of the backbone network to obtain a first image to be processed output by the second C3 module, a second image to be processed output by the third C3 module, and a third image to be processed output by the SPPF module; and processing the first image to be processed, the second image to be processed, and the third image to be processed based on the second module to obtain a target connected image and its output label.

[0065] The convolutional layer (conv) can consist of convolution, batch normalization, and activation layers. Batch normalization helps prevent overfitting and accelerates convergence. The activation layer can use either the SiLu or ReLU activation function. The C3 module can contain three standard convolutional layers (Conv and Bottleneck), and its structural diagram can be found in [reference needed]. Figure 3 The SPPF (Spatial Pyramid Pooling Fast) module uses three 5×5 max pooling operations and multiple small-sized pooling kernels cascaded together. This can improve the running speed while effectively fusing feature maps from different receptive fields and enriching the expressive power of the feature maps.

[0066] See Figure 4 The first module in the backbone network can consist of convolutional layers, C3 modules, and an SPPF module. Based on this, a stitched image can be input into the first module for processing. The second C3 module outputs a first image to be processed corresponding to the stitched image, the third C3 module outputs a second image to be processed corresponding to the stitched image, and the SPPF module outputs a third image to be processed corresponding to the stitched image. Further, the first image to be processed, the second image to be processed, and the third image to be processed can be input into different processing layers in the second module for further processing to obtain the processed target connection image and the recognition result of the target connection image, i.e., the output label.

[0067] Optionally, the second module includes a first connection layer, a second connection layer, a third convolutional layer, and an upsampling layer. Based on the second module, the first image to be processed, the second image to be processed, and the first image to be processed are processed to obtain a target connected image and its output label. This includes: processing the third image to be processed using the third convolutional layer in the second module to obtain a first image to be upsampled; processing the first image to be upsampled using the upsampling layer in the second module to obtain a first image to be connected; processing the second image to be processed and the first image to be connected using the second connection layer in the second module to obtain a second image to be upsampled; processing the second image to be upsampled using the upsampling layer in the second module to obtain a second image to be connected; processing the first image to be processed and the second image to be connected using the first connection layer in the second module to obtain an output image; and determining the target connected image and its output label based on the first image to be upsampled, the second image to be upsampled, and the output image.

[0068] See also Figure 4 The third convolutional layer contains three connected convolutional layers, the second connection layer contains a concat connection layer and three connected convolutional layers, and the first connection layer contains a concat connection layer and three connected convolutional layers. In practical applications, the first image to be processed can be input into the first connection layer of the second module, the second image to be processed can be input into the second connection layer of the second module, and the third image to be processed can be input into the third convolutional layer of the second module. The third convolutional layer outputs the first upsampled image after processing the third image. The first upsampled image is then input into the upsampling layer of the second module for upsampling, outputting the first image to be connected. The first image to be connected is then input into the second connection layer, concatenated with the second image to be processed, and then processed by three convolutional layers to output the second upsampled image. The second upsampled image is then input into the upsampling layer of the second module for upsampling, obtaining the second image to be connected. The second image to be connected is then input into the first connection layer, concatenated with the first image to be processed, and then processed by three convolutional layers to output the final image. The first image to be upsampled, the second image to be upsampled, and the image to be output can all be used as target connected images. Hand-raising detection is performed on the target connected images to obtain the output result, which is the output label corresponding to the target connected image. The output format can be (hand-raising confidence, position of the upper left corner of the hand-raising recognition box, length and width of the hand-raising recognition box). The technical solution of this embodiment, by using a three-segment structure for image processing in the second module, not only ensures the model's detection accuracy but also improves data processing efficiency.

[0069] S230. Based on the target hand-raising detection model, process the image to be identified to obtain the recognition result of the image to be identified.

[0070] The technical solution of this embodiment obtains multiple training samples and then trains a target hand-raising detection model based on the multiple training samples. This enables the target hand-raising detection model to recognize the image to be recognized, thereby improving recognition efficiency and recognition effect.

[0071] Example 3

[0072] Figure 5 This is a schematic diagram of the structure of a video processing device according to Embodiment 3 of the present invention. Figure 5 As shown, the device includes: an extraction module 310, an image recognition module 320, a speech recognition module 330, and a video segment determination module 340.

[0073] The extraction module 310 is used to acquire the teaching video to be extracted and extract multiple frames of images to be recognized and voice data from the teaching video; the image recognition module 320 is used to determine the hand-raising time corresponding to the image to be recognized if the recognition result of the multiple frames of images to be recognized is a preset action; the voice recognition module 330 is used to perform recognition processing on the voice data to obtain the voice segment of the target object corresponding to the preset role and determine the starting voice time of the voice segment; the video segment determination module 340 is used to determine the hand-raising answer video segment from the teaching video based on the hand-raising time and the starting voice time.

[0074] The technical solution of this embodiment acquires a teaching video to be extracted and extracts multiple frames of images to be recognized and audio data from the teaching video. For multiple frames of images to be recognized, if the recognition result of the image to be recognized is a preset action, the hand-raising moment corresponding to the image to be recognized is determined. The audio data is processed to obtain the audio segment of the target object corresponding to the preset role, and the starting audio moment of the audio segment is determined. Based on the hand-raising moment and the starting audio moment, the hand-raising and answering video segment is determined from the teaching video. The teaching quality of the teaching video is evaluated based on the hand-raising and answering video segment. This solves the problem of evaluation errors caused by image detection in the prior art. It realizes that by performing action recognition on multiple frames of images to be recognized in the teaching video, and performing audio recognition on each role in the teaching video, the hand-raising action in the image and the answer of the target object in the audio are combined to comprehensively determine the hand-raising and answering video segment where the hand-raising action occurs and there is an answer. Then, the teaching video is evaluated based on the hand-raising and answering video segment, making the evaluation of the teaching process more effective and objective, which is conducive to improving the quality of education and teaching.

[0075] Optionally, based on the above-mentioned device, the speech recognition module 330 includes a speech total duration determination unit, a target object determination unit, and a speech segment determination unit.

[0076] The total voice duration determination unit is used to separate the voice data based on the role separation algorithm and determine the total voice duration corresponding to different objects;

[0077] The target object determination unit is used to determine, based on the total duration of each of the aforementioned voice recordings, a target object corresponding to the preset role from among the different objects; wherein, the preset role includes the person raising their hand to answer;

[0078] A speech segment determination unit is used to determine the speech segments of the target object.

[0079] Optionally, based on the above-mentioned device, the video segment determination module 340 is specifically used to determine a hand-raising response video segment based on the hand-raising time and the end voice time of the voice segment to which the start voice time belongs if a matching start voice time exists within a preset time after the hand-raising time, so as to evaluate the teaching quality of the teaching video based on the hand-raising response video segment.

[0080] Optionally, based on the above-described apparatus, the apparatus may further include a training sample acquisition module and a model training module.

[0081] The training sample acquisition module is used to acquire multiple training samples, which include an image to be trained and the annotation results corresponding to the image to be trained.

[0082] The model training module is used to train a target hand-raising detection model based on the multiple training samples, and to process the image to be identified based on the target hand-raising detection model to obtain the recognition result of the image to be identified.

[0083] Based on the above-mentioned device, optionally, the model training module includes a stitched image determination unit, an output label determination unit, a loss value determination unit, a parameter correction unit, and a model training unit.

[0084] The image stitching determination unit is used to input the images to be trained from the plurality of training samples into the hand-raising detection model to be trained, and to process each group of images to be trained based on the data augmentation model in the hand-raising detection model to be trained, so as to obtain a stitched image corresponding to each group of images to be trained; the group of images to be trained includes a preset number of images to be trained; the number of annotation results in the stitched image corresponds to the preset number.

[0085] The output label determination unit is used to process each of the stitched images based on the backbone network in the hand-raising detection model to obtain output labels.

[0086] The loss value determination unit is used to determine the loss value based on the annotation results in the stitched image and the output label;

[0087] A parameter correction unit is used to correct the model parameters in the hand-raising detection model to be trained based on the loss value;

[0088] The model training unit is used to converge the loss function of the hand-raising detection model to be trained as the training objective, so as to obtain the target hand-raising detection model.

[0089] Based on the above-mentioned device, optionally, the backbone network includes a first module and a second module; the first module consists of multiple convolutional layers, multiple C3 modules and an SPPF module;

[0090] The first module is used to process the stitched image based on the first module in the backbone network to obtain a first image to be processed based on the output of the second C3 module, a second image to be processed based on the output of the third C3 module, and a third image to be processed based on the output of the SPPF module.

[0091] The second module is used to process the first image to be processed, the second image to be processed, and the first image to be processed based on the second module to obtain the target connection image and the output label of the target connection image.

[0092] Based on the above-mentioned device, optionally, the second module includes a first connection layer, a second connection layer, a third convolutional layer, and an upsampling layer;

[0093] The third convolutional layer is used to process the third image to be processed based on the third convolutional layer in the second module to obtain the first image to be upsampled;

[0094] The upsampling layer is used to process the first image to be upsampled based on the upsampling layer in the second module to obtain the first image to be connected.

[0095] The second connection layer is used to process the second image to be processed and the first image to be connected based on the second connection layer in the second module to obtain the second image to be upsampled;

[0096] The upsampling layer is used to process the second image to be upsampled based on the upsampling layer in the second module to obtain the second image to be connected.

[0097] The first connection layer is used to process the first image to be processed and the second image to be connected based on the first connection layer in the second module to obtain the image to be output;

[0098] The device further includes an output module, which is used to determine a target connection image based on the first image to be upsampled, the second image to be upsampled, and the image to be output, and to determine the output label of the target connection image.

[0099] The video processing apparatus provided in the embodiments of the present invention can execute the video processing method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the method.

[0100] Example 4

[0101] Figure 6 This is a schematic diagram of the structure of an electronic device implementing the video processing method of an embodiment of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0102] like Figure 6 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0103] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0104] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as video processing methods.

[0105] In some embodiments, the video processing method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or mounted on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the video processing method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the video processing method by any other suitable means (e.g., by means of firmware).

[0106] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0107] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0108] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0109] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0110] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0111] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0112] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0113] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A video processing method, characterized in that, include: The teaching video to be extracted is obtained, and multiple frames of images and audio data to be recognized are extracted from the teaching video. For the multiple frames of images to be identified, if the identification result of the image to be identified is a preset action, then the hand-raising time corresponding to the image to be identified is determined; The voice data is processed to obtain a voice segment of the target object corresponding to a preset role, and the start time of the voice segment is determined. Based on the hand-raising moment and the start of the voice, a hand-raising and answering video segment is determined from the teaching video to evaluate the teaching quality of the teaching video based on the hand-raising and answering video segment; Multiple training samples are obtained, including the image to be trained and the corresponding annotation results of the image to be trained; A target hand-raising detection model is trained based on the multiple training samples, and the image to be identified is processed based on the target hand-raising detection model to obtain the recognition result of the image to be identified; The target hand-raising detection model trained based on the multiple training samples includes: The images to be trained from the multiple training samples are input into the hand-raising detection model to be trained. Based on the data augmentation model in the hand-raising detection model to be trained, each group of images to be trained is processed to obtain a stitched image corresponding to each group of images to be trained. The group of images to be trained includes a preset number of images to be trained. The number of annotation results in the stitched image corresponds to the preset number. The data augmentation model randomly selects a preset number of images to be trained as a group of images to be trained using a mosaic data augmentation method. For each of the stitched images, the stitched images are processed based on the backbone network in the hand-raising detection model to be trained, and output labels are obtained; The loss value is determined based on the annotation results in the stitched image and the output label; The model parameters in the hand-raising detection model to be trained are corrected based on the loss value; The convergence of the loss function of the hand-raising detection model to be trained is taken as the training objective, and the target hand-raising detection model is obtained.

2. The method according to claim 1, characterized in that, The process of recognizing and processing the voice data to obtain a voice segment of the target object corresponding to a preset role includes: The voice data is separated based on a role separation algorithm to determine the total voice duration for different objects. Based on the total duration of each of the aforementioned voice recordings, a target object corresponding to the preset role is determined from the different objects; wherein, the preset role includes the person raising their hand to answer; Identify the speech segment of the target object.

3. The method according to claim 1, characterized in that, The step of determining the video segment for raising hands and answering from the teaching video based on the moment the hand is raised and the moment the initial voice is spoken includes: If a matching start voice moment exists within a preset time after the hand-raising moment, then a hand-raising response video segment is determined based on the hand-raising moment and the end voice moment of the voice segment to which the start voice moment belongs.

4. The method according to claim 1, characterized in that, The backbone network includes a first module and a second module; the first module consists of multiple convolutional layers, multiple C3 modules, and an SPPF module; the processing of the stitched image based on the backbone network in the hand-raising detection model to obtain output labels includes: The stitched image is processed based on the first module in the backbone network to obtain a first image to be processed based on the output of the second C3 module, a second image to be processed based on the output of the third C3 module, and a third image to be processed based on the output of the SPPF module. The second module processes the first image to be processed, the second image to be processed, and the first image to be processed to obtain the target connected image and the output label of the target connected image.

5. The method according to claim 4, characterized in that, The second module includes a first connection layer, a second connection layer, a third convolutional layer, and an upsampling layer; the step of processing the first image to be processed, the second image to be processed, and the first image to be processed based on the second module to obtain a target connected image and the output label of the target connected image includes: The third image to be processed is processed based on the third convolutional layer in the second module to obtain the first image to be upsampled. The first image to be upsampled is then processed based on the upsampling layer in the second module to obtain the first image to be connected. The second image to be processed and the first image to be connected are processed based on the second connection layer in the second module to obtain the second image to be upsampled. The second image to be upsampled is then processed based on the upsampling layer in the second module to obtain the second image to be connected. The first image to be processed and the second image to be connected are processed based on the first connection layer in the second module to obtain the image to be output; Based on the first image to be upsampled, the second image to be upsampled, and the image to be output, a target connected image is determined, and the output label of the target connected image is determined.

6. A video processing apparatus, characterized in that, include: The extraction module is used to acquire the teaching video to be extracted, and to extract multiple frames of images to be recognized and audio data from the teaching video; The image recognition module is used to determine the hand-raising moment corresponding to the image to be recognized if the recognition result of the image to be recognized is a preset action for the multiple frames of images to be recognized. The speech recognition module is used to recognize and process the speech data to obtain a speech segment of the target object corresponding to a preset role, and to determine the start time of the speech segment. A video segment determination module is used to determine a video segment of raising hands and answering from the teaching video based on the time of raising hands and the time of starting voice, so as to evaluate the teaching quality of the teaching video based on the video segment of raising hands and answering. The device also includes a training sample acquisition module and a model training module; The training sample acquisition module is used to acquire multiple training samples, which include the image to be trained and the annotation results corresponding to the image to be trained. The model training module is used to train a target hand-raising detection model based on the multiple training samples, and to process the image to be identified based on the target hand-raising detection model to obtain the recognition result of the image to be identified. The model training module includes a stitched image determination unit, an output label determination unit, a loss value determination unit, a parameter correction unit, and a model training unit; The stitched image determination unit is used to input the images to be trained from the multiple training samples into the hand-raising detection model to be trained, and to process each group of images to be trained based on the data augmentation model in the hand-raising detection model to be trained, so as to obtain a stitched image corresponding to each group of images to be trained. The group of images to be trained includes a preset number of images to be trained. The number of annotation results in the stitched image corresponds to the preset number; wherein, the data augmentation model randomly selects a preset number of training images as a training image group using a mosaic data augmentation method; The output label determination unit is used to process each stitched image based on the backbone network in the hand-raising detection model to obtain an output label. The loss value determination unit is used to determine the loss value based on the annotation results in the stitched image and the output label; The parameter correction unit is used to correct the model parameters in the hand-raising detection model to be trained based on the loss value; The model training unit is used to converge the loss function of the hand-raising detection model to be trained as the training objective, so as to obtain the target hand-raising detection model.

7. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the video processing method according to any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the video processing method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Video cutting method and device, computer equipment and storage medium

    CN109743624A

  • Classroom type intelligent evaluation system based on teaching video

    CN115170064A