Methods, devices, equipment and storage media for recognizing singing videos

By using audio and video feature detection and lip movement recognition methods, video clips of real people singing are selected, which solves the problem of low recognition accuracy in existing technologies and improves resource utilization and server efficiency.

CN113762056BActive Publication Date: 2025-10-28TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110541736.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-05-18
Publication Date
2025-10-28
Estimated Expiration
2041-05-18

AI Technical Summary

Technical Problem

In existing technologies, the accuracy of video content recognition is low, the resource utilization rate in information flow content services is low, and it is difficult to effectively identify and manage videos of real people singing.

Method used

By combining audio and video feature detection, including audio detection and face detection, video segments that meet the criteria are selected, and then lip movement recognition is performed to determine that the video segments are singing videos.

Benefits of technology

It improves the accuracy of recognizing singing videos, reduces the amount of computation, and enhances the resource utilization and server operating efficiency of information flow content services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113762056B_ABST
    Figure CN113762056B_ABST
Patent Text Reader

Abstract

This application discloses a method, apparatus, device, and storage medium for recognizing singing videos, belonging to the field of computer technology. The method includes: acquiring a video segment corresponding to video data; performing audio-visual feature detection processing on the video segment to obtain audio detection results and face detection results; if the audio detection results and face detection results meet a first condition, performing lip movement recognition processing on the video segment to obtain lip movement recognition results corresponding to the video segment; if the lip movement recognition results meet a second condition, determining the video segment as a singing video segment. In the technical solution provided by this application, audio features and face features of the video segment are obtained through audio-visual feature detection. If the two features meet the first condition, lip movements are then recognized; if the recognition result meets the second condition, it is determined to be a singing video segment. This effectively improves the accuracy of singing video recognition while reducing computational load.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a method, apparatus, device, and storage medium for recognizing singing videos. Background Technology

[0002] With the rapid development of live streaming platforms and video sharing platforms, many users broadcast their live singing performances on live streaming platforms or upload pre-recorded videos of their live singing performances to video sharing platforms for sharing. These high-quality pre-recorded videos and live singing clips are suitable for recommending to other users.

[0003] Live streaming and video sharing platforms typically inspect the videos on their platforms and use the inspection results to manage the videos. For example, the aforementioned platform uses audio data from the video to determine if it contains music, thereby identifying the video type.

[0004] Among the solutions provided by related technologies, the accuracy of video content recognition is low, and the resource utilization rate in information flow content services is low. Summary of the Invention

[0005] This application provides a method, apparatus, device, and storage medium for recognizing singing videos, which can improve the recognition accuracy of live singing videos, improve resource utilization in information flow content services, and improve server operating efficiency.

[0006] According to one aspect of the embodiments of this application, a singing video recognition method is provided, the method comprising:

[0007] Retrieve the video segment corresponding to the video data;

[0008] The video segment is subjected to audio and video feature detection processing to obtain the audio detection result and face detection result of the video segment;

[0009] If the audio detection result and the face detection result meet the first condition, lip movement recognition processing is performed on the video segment to obtain the lip movement recognition result corresponding to the video segment;

[0010] If the lip movement recognition result meets the second condition, the video segment is determined to be a singing video segment.

[0011] According to one aspect of the embodiments of this application, a singing video recognition device is provided, the device comprising:

[0012] The video clip acquisition module is used to acquire video clips corresponding to video data;

[0013] The audio and video feature detection module is used to perform audio and video feature detection processing on the video segment to obtain the audio detection result and face detection result of the video segment;

[0014] The lip movement recognition module is used to perform lip movement recognition processing on the video segment when the audio detection result and the face detection result meet the first condition, so as to obtain the lip movement recognition result corresponding to the video segment;

[0015] The video type determination module is used to determine that the video segment is a singing video segment if the lip movement recognition result meets the second condition.

[0016] According to one aspect of the embodiments of this application, a computer device is provided, the computer device including a processor and a memory, the memory storing at least one instruction, at least one program, code set or instruction set, the at least one instruction, the at least one program, the code set or instruction set being loaded and executed by the processor to implement the above-described singing video recognition method.

[0017] According to one aspect of the embodiments of this application, a computer-readable storage medium is provided, wherein the storage medium stores at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the above-described singing video recognition method.

[0018] According to one aspect of the embodiments of this application, a computer program product or computer program is provided, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the aforementioned singing video recognition method.

[0019] The technical solution provided in this application can bring the following beneficial effects:

[0020] By combining multimodal information from the video, the video is segmented into segments, and audio and video features are detected on the video segments to obtain audio detection results and face detection results for the video segments. Then, based on these two aspects of feature information, a preliminary judgment is made on the video segments. Lip action recognition is only performed if the first condition is met. If the lip action recognition result meets the second condition, the singing video segment can be finally identified. This effectively improves the recognition accuracy of singing videos while reducing the amount of computation. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 This is a schematic diagram of an application runtime environment provided in one embodiment of this application;

[0023] Figures 2 to 5 These are flowcharts of the singing video recognition methods provided in various embodiments of this application;

[0024] Figure 6 An exemplary schematic diagram of a process for recognizing singing videos is shown;

[0025] Figure 7 This is a block diagram of a singing video recognition device provided in one embodiment of this application;

[0026] Figure 8 This is a structural block diagram of a computer device provided in one embodiment of this application. Detailed Implementation

[0027] The solutions provided in this application involve artificial intelligence technology and cloud technology, which will be briefly described below to facilitate understanding by those skilled in the art.

[0028] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.

[0029] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0030] Key technologies in speech technology include Automatic Speech Recognition (ASR), Text-to-Speech (TTS), and voiceprint recognition. Enabling computers to hear, see, speak, and feel is the future direction of human-computer interaction, with speech emerging as one of the most promising methods.

[0031] The goal of automatic speech recognition technology is to convert the lexical content of human speech into computer-readable input, such as keystrokes, binary codes, or character sequences. This differs from speaker recognition and speaker verification, which attempt to identify or verify the speaker uttering the speech rather than the lexical content it contains.

[0032] Machine learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and instructional learning.

[0033] Cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or local area network to achieve data computing, storage, processing, and sharing.

[0034] Cloud technology is a collective term for network technologies, information technologies, integration technologies, management platform technologies, and application technologies applied to the cloud computing business model. It can form resource pools, providing flexible and convenient on-demand access. Cloud computing technology will become a crucial support. Backend services of technical network systems require substantial computing and storage resources, such as video websites, image websites, and many portal websites. With the rapid development and application of the internet industry, every item may have its own identification mark in the future, requiring transmission to backend systems for logical processing. Data at different levels will be processed separately, and various industry data will all require robust system support, which can only be achieved through cloud computing.

[0035] Cloud computing is a computing model that distributes computing tasks across a large pool of computers, enabling various application systems to access computing power, storage space, and information services as needed. The network providing these resources is called the "cloud." From the user's perspective, resources in the "cloud" appear infinitely scalable, readily available, on-demand, and expandable, with payment based on usage.

[0036] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0037] Please refer to Figure 1 This diagram illustrates an application runtime environment provided in one embodiment of this application. The application runtime environment may include: terminal 10 and server 20.

[0038] Terminal 10 can be an electronic device such as a mobile phone, tablet computer, game console, e-book reader, multimedia playback device, wearable device, or PC (Personal Computer). Application clients can be installed on terminal 10.

[0039] In this embodiment, the application described above can be any application capable of providing information flow content services. Typically, the application is a content sharing application. Of course, in addition to content sharing applications, other types of applications can also provide information flow content services. For example, video live streaming applications, video sharing applications, content interaction applications, news applications, social applications, interactive entertainment applications, browser applications, shopping applications, virtual reality (VR) applications, augmented reality (AR) applications, etc., are not limited in this embodiment. In addition, the content created and uploaded by users will be different for different applications, and the corresponding functions will also be different. These can be pre-configured according to actual needs, and are not limited in this embodiment. Optionally, the terminal 10 runs a client of the above-mentioned application.

[0040] Server 20 provides background services for clients of applications in terminal 10. For example, server 20 can be a background server for the aforementioned applications. Server 20 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. Terminals can be smartphones, tablets, laptops, desktop computers, smart speakers, smartwatches, etc., but are not limited to these. Optionally, server 20 can simultaneously provide background services for applications in multiple terminals 10.

[0041] Optionally, terminal 10 and server 20 can communicate with each other via network 30. Terminal 10 and server 20 can be directly or indirectly connected via wired or wireless communication, which is not limited herein.

[0042] Please refer to Figure 2 The diagram illustrates a flowchart of a singing video recognition method according to an embodiment of this application. This method can be applied to a computer device, which refers to an electronic device capable of data calculation and processing. The method may include the following steps (210-240).

[0043] Step 210: Obtain the video segment corresponding to the video data.

[0044] Optionally, the video data includes recorded video data and live video data.

[0045] When the video data is pre-recorded video data, the pre-recorded video is divided into at least one video segment. The pre-recorded video data can be a single, complete video file, which can be a user-created video file uploaded to a content sharing platform. This application does not limit the video content or the video upload format. Optionally, the pre-recorded video can be divided into multiple video segments of a fixed duration.

[0046] When the video data is live video data, the system receives the live video data stream; it then performs segmentation processing on the live video data stream to obtain video clips. The aforementioned live video data can be video data uploaded by users to a content sharing platform in real time. Optionally, a fixed-duration video data segment can be extracted from the live video data stream to obtain the video clip.

[0047] Optionally, the aforementioned video data belongs to user-generated content (UGC), which typically refers to users displaying or providing their original content to other users through internet platforms.

[0048] Step 220: Perform audio and video feature detection processing on the video clip to obtain the audio detection results and face detection results of the video clip.

[0049] The aforementioned audio and video feature detection processing refers to the processing of detecting the audio and image features of the aforementioned video clips.

[0050] The audio detection results described above are used to characterize the audio features of the aforementioned video segment. These results can be the audio type of the audio content corresponding to the detected video segment. For example, the audio content of the video segment could be singing audio, speech audio, etc.

[0051] The aforementioned face detection results are used to characterize the image features of the aforementioned video clip. Optionally, the aforementioned face detection results include face images and non-face images. The aforementioned face images may be images containing faces in the image frames of the video clip, or images containing faces that meet preset face filtering conditions. The aforementioned face filtering conditions may be that the face area is greater than an area threshold, or that the face position is within a preset image area; this application embodiment does not limit the specific criteria. The aforementioned non-face images may be images where no faces are present in the image frames of the video clip, or images where no faces meet the preset face filtering conditions are present. The aforementioned face detection results are used to characterize whether the image frames in the video clip are face images.

[0052] In an exemplary embodiment, such as Figure 3 As shown, it illustrates a flowchart of a singing video recognition method provided in an embodiment of this application. Step 220 above further includes the following sub-steps (221-222).

[0053] Step 221: Perform audio event detection processing on the audio data of the video segment to obtain the audio detection result of the video segment.

[0054] Optionally, corresponding audio data can be extracted from the video clip, and then audio event detection processing can be performed on the extracted audio data to obtain the audio detection results of the video clip. The aforementioned audio event detection processing refers to the process of detecting the similarity between audio content and preset audio events. Through audio event detection processing, the probability of the audio content being each audio event can be determined. These audio events can be pre-defined entities representing various audio content; for example, audio events include singing, speech, noise, music, silence, etc. The audio type that best matches the aforementioned audio data can be determined using these probabilities.

[0055] Step 222: Perform face detection processing on the image data of the video clip to obtain the face detection results of the video clip.

[0056] Optionally, corresponding image data can be extracted from the video clip, and then face detection processing can be performed on the image data to determine whether the image content represented by the above image data is a face image containing a face.

[0057] Step 230: If the audio detection result and the face detection result meet the first condition, perform lip action recognition processing on the video segment to obtain the lip action recognition result corresponding to the video segment.

[0058] The first condition mentioned above is used to filter video segments that contain human faces and include singing but no audio. Optionally, the first condition may be that the audio detection result indicates the audio type of the video segment is singing but no audio, and the face detection result indicates that the video segment contains human faces.

[0059] Because the audio type of the singing video is "singing without voice", some irrelevant videos can be excluded by this feature, such as videos with voice calls in the background music, and non-singing videos.

[0060] Since the singing video requires real people to sing, the presence of faces in the video clip can be determined by face detection results, which can exclude other irrelevant videos, such as videos without human faces or images.

[0061] By restricting the audio and face detection results by the first condition mentioned above, a preliminary judgment can be made on the video clips, eliminating a large number of irrelevant non-singing videos. Only when the audio and face detection results meet the first condition can the above-mentioned lip movement recognition processing be performed, which can effectively reduce the amount of computation and alleviate the operating pressure on computer equipment.

[0062] The lip movement recognition processing mentioned above refers to the image processing method for recognizing facial mouth movements. Since real people need to make sounds through their mouths when singing, the singer will open and close their mouths. The opening and closing of the mouth will cause changes in the position of the lips. The lip movement recognition processing mentioned above can determine the lip movement recognition results, so as to exclude videos with static mouths or videos where the mouth movements are not speaking or singing, such as videos of eating or facial expressions.

[0063] Optionally, the above lip movement recognition result can be the recognized lip movement state, such as the lip movement state being in a speaking or singing state, an eating state, or an expression display state.

[0064] Step 240: If the lip movement recognition result meets the second condition, the video segment is determined to be a singing video segment.

[0065] The second condition mentioned above is used to filter out video clips that include lip movements that match the lip movement characteristics of a real singer. The first condition has already identified video clips containing a face and singing but without audio. By ensuring that the lip movement recognition results meet the second condition, videos with background music but whose lip movements do not match the singing / speaking state can be excluded, thus identifying the accurate videos of real singers.

[0066] Determining whether audio in a video is speech or singing solely based on the audio itself is often insufficient to distinguish genuine live performances. This results in poor accuracy and recall rates for performance video recognition strategies in information feed content services, hindering practical application. By combining multimodal video information with appropriate audio event detection, face detection, and lip movement recognition methods, it is possible to accurately identify live performance video clips in recorded or live streams.

[0067] In an exemplary embodiment, such as Figure 3 As shown, the above method also includes the following steps (250-260).

[0068] Step 250: Segment the corresponding singing video segments from the video data to generate a singing video.

[0069] Optionally, the singing video segments in the recorded video can be spliced ​​together to generate the singing video in the recorded video.

[0070] Optionally, the singing video clips in the live video can be spliced ​​together to generate the singing video in the live video.

[0071] In one possible implementation, the singing video recognition method provided in this application can automatically identify segments of real people singing in recorded videos and provide corresponding time point markers, enabling video editing systems to quickly and automatically edit exciting segments in videos based on the aforementioned time point markers, thereby improving video editing efficiency and reducing computer equipment running time.

[0072] Step 260: Push the performance video in the information flow service.

[0073] In one possible real-time approach, the singing video recognition method provided in this application can automatically identify segments of real people singing in a video, thereby discovering high-quality recorded or live video clips and quickly pushing them to users for viewing. In an exemplary embodiment, such as Figure 3 As shown, the above method also includes the following step (270).

[0074] Step 270: Determine the live video stream corresponding to the performance video segment and push the live video stream in the information flow service.

[0075] In one possible real-time approach, the singing video recognition method provided in this application can identify in real time whether there is a real person singing in a live video. Based on this, the recommendation system of various video live streaming platforms can recommend live streamers who are performing well to users in real time, thereby improving the user's viewing time.

[0076] In an exemplary embodiment, such as Figure 3 As shown, the above method also includes the following steps (280-290).

[0077] Step 280: Determine the user account corresponding to the singing video clip.

[0078] Step 290: Mark the type of user account.

[0079] Singing live streams are a major component of video live streaming platforms. In one possible implementation, the above method can also be used to automatically identify singing streamers. Within the live stream account classification systems of various platforms, singing videos can be identified first, and then the user account that posted the singing video can be determined as a singing streamer, improving user management efficiency and reducing the operational burden on computer equipment.

[0080] In an exemplary embodiment, the method further includes adding singing video tags to the singing video clips. The singing video recognition method provided in this application can automatically identify whether a short video is a live singing video, and then add corresponding singing video tags to the singing video to complete the video classification task in various video sharing platforms and improve resource management efficiency. Optionally, videos can also be recommended based on these tags in information flow content services.

[0081] As can be seen from the above exemplary embodiments, the singing video recognition method provided in this application can be applied to systems such as singing anchor recognition and classification, real-time recommendation of highlights from singing live broadcasts, quick editing of highlights from recorded videos, and classification and recommendation of short videos in video live streaming platforms.

[0082] In summary, the technical solution provided in this application combines video multimodal information to segment the video into segments and perform audio and video feature detection on the video segments to obtain audio detection results and face detection results of the video segments. Then, based on these two aspects of feature information, a preliminary judgment is made on the video segments. Lip action recognition is only performed if the first condition is met. If the lip action recognition result meets the second condition, the singing video segment can be finally identified. This effectively improves the recognition accuracy of singing videos while reducing the amount of computation.

[0083] Furthermore, the technical solutions provided in this application can also be applied to systems such as live streaming recommendation, live streamer classification, video editing systems, video classification, and video recommendation. The solutions are highly portable and flexible.

[0084] Please refer to Figure 4 The diagram illustrates a flowchart of a singing video recognition method provided in another embodiment of this application. This method can be applied to... Figure 1 The method may include the following steps (401-415) in the application runtime environment shown.

[0085] Step 401: If the video data is pre-recorded video data, divide the pre-recorded video into at least one video segment.

[0086] Step 402: If the video data is live video data, receive the live video data stream.

[0087] Step 403: The live video data stream is segmented to obtain video clips.

[0088] Optionally, the method of obtaining the above video segments is not limited to segmenting according to a fixed duration. It can also be achieved by segmenting the video into different state stages through audio event detection or scene recognition methods.

[0089] Step 404: Extract audio data from the video clip.

[0090] Step 405: Perform frame segmentation on the audio data to obtain the audio frames corresponding to each time period.

[0091] Optionally, the audio data is segmented into frames according to a preset duration to obtain audio frames corresponding to each time period. Optionally, the duration of each time period is a preset duration.

[0092] Step 406: Perform audio event detection processing on the audio frame to obtain the score of the audio frame on each audio event.

[0093] The aforementioned audio events correspond to different types of audio content. Optionally, audio events include speech events and singing events; in other application scenarios, other audio events may also be included, which is not limited in this embodiment. The aforementioned speech event indicates that speech information exists in the audio content, and the aforementioned singing event indicates that singing information exists in the audio content.

[0094] The scoring includes a speech score and a singing score. The speech score represents the probability that the audio content of an audio frame belongs to a speech event, and the singing score represents the probability that the audio content of an audio frame belongs to a singing event. Optionally, the scoring may also include a music score, a silence score, and a noise score. The music score represents the probability that the audio content of an audio frame belongs to a music playback event, the silence score represents the probability that the audio content of an audio frame belongs to a silence event, and the noise score represents the probability that the audio content of an audio frame belongs to a noise event.

[0095] In one possible implementation, the audio event detection method includes: converting audio frames into two-dimensional Mel spectrograms, extracting corresponding image feature vectors using a VGGish (Visual Geometry Group) deep neural network, and then training a classification neural network model to classify audio events, obtaining a score for each audio frame on each audio event. Optionally, the audio event detection method is not limited to the above approach and can also be implemented using other types of audio features through an end-to-end deep learning model, such as using pretrained audio neural networks (PANNs) for audio pattern recognition.

[0096] Step 407: If the speech score is less than the speech threshold and the singing score is greater than the singing threshold, determine the audio detection result of the audio frame as having singing but no speech.

[0097] The audio detection results of the audio frame represent the audio detection results of the video segment.

[0098] The audio category of an audio frame can be determined based on the score, and this audio category is used as the audio detection result for the audio frame. Optionally, if the speech score is greater than a speech threshold, the audio detection result for the audio frame is determined to have a speech category. Optionally, if the speech score is less than a speech threshold but the singing score is greater than a singing threshold, the audio detection result for the audio frame is determined to have singing but no speech category. Optionally, if the speech score is less than a speech threshold and the singing score is less than or equal to a singing threshold, the audio detection result for the audio frame is determined to have neither singing nor a speech category.

[0099] Step 408: Obtain the original sequence of image frames from the video clip.

[0100] Step 409: Perform first image frame extraction processing on the original image frame sequence to obtain the first image frame sampling sequence.

[0101] The first image frame extraction process is used to extract a first fixed number of image frames within each time period. For example, one frame per second is extracted from the video to obtain the first image frame sampling sequence. The first fixed number can be a preset number, such as 1. The time periods can be at least one time period divided by a fixed duration, or they can be time periods determined according to the video content.

[0102] Step 410: Perform face detection processing on the first image frame in the first image frame sampling sequence to obtain the face parameters of the first image frame.

[0103] The first image frame is an image frame extracted from the original sequence of image frames.

[0104] A face parameter matrix is ​​obtained through face detection methods. This matrix includes face parameters, primarily the bounding box position (center point coordinates, length, width), facial landmark positions (coordinates of both eyes, nose, and corners of the mouth), and a face score. The face score is used to measure face quality, representing the probability that the image content within the bounding box is a face.

[0105] Optionally, the face detection described above can be implemented using an SSH (Single Stage Headless) face detection model. This application does not limit the face detection method, and adjustments and replacements can be made to the method based on specific application scenarios. The SSH face detection model described above is a trained neural network model used for face detection.

[0106] Step 411: If the face parameters of the first image frame meet the threshold condition, determine that the face detection result of the first image frame is a face image.

[0107] The aforementioned face parameters meeting the threshold conditions include a face score greater than a face score threshold, a face length greater than a length threshold, and a face width greater than a width threshold. Optionally, the face parameters meeting the threshold conditions also include the face rectangle being located within a preset range of the image, wherein the face within the preset range of the image is the primary face. The aforementioned face score threshold is a threshold used to evaluate face scores, the aforementioned length threshold is a threshold used to evaluate face length, and the aforementioned width threshold is a threshold used to evaluate face width.

[0108] The face detection results of the first image frame described above represent the face detection results of the video segment.

[0109] Step 412: If the audio detection result of the audio frame is singing but no speech category, and the face detection result of the first image frame belonging to the same time period as the audio frame is a face image, then the time period corresponding to the audio frame and the first image frame is determined to be the valid time period.

[0110] The aforementioned effective time period indicates that the video contains a human face and singing without audio, which is a time period that meets the characteristics of a singing video.

[0111] Step 413: Calculate the duration of the effective time periods of the video segments to obtain the total effective duration.

[0112] The total effective duration is obtained by summing the durations corresponding to the effective time periods.

[0113] Step 414: If the ratio of the total effective duration to the total duration of the video segment is greater than the ratio threshold, perform lip action recognition processing on the video segment to obtain the lip action recognition result corresponding to the video segment.

[0114] Based on the audio categories and main facial information mentioned above, the length of audio segments containing main faces and featuring singing without spoken words is counted. If this length accounts for a proportion greater than a threshold of the total video segment length, it is preliminarily identified as a singing segment. If it is preliminarily identified as a singing segment, the next steps continue; otherwise, the process ends. The aforementioned threshold is used to assess the ratio between the total effective duration and the total duration of the video segment.

[0115] In one possible implementation, a filtering mechanism for determining the duration of audio is added. If the duration of audio exceeds a threshold or the proportion of audio duration to the total duration exceeds a threshold, the video segment can be directly determined to be a non-singing video.

[0116] Extract lip images from image frames in a video clip to obtain a lip image sequence.

[0117] Based on the extracted lip image sequence, a deep learning model is used to determine whether the lip movements of the main character in the video are in a speaking or singing state or some other state.

[0118] Step 415: If the lip movement recognition result meets the second condition, the video segment is determined to be a singing video segment.

[0119] For video clips initially identified as singing, if the lip movement recognition result indicates speaking and singing, the video clip is ultimately determined to be a performance clip (i.e., a live singing clip).

[0120] Optionally, by changing the judgment criteria, the above method can also identify real-person voice call videos. For example, if the audio detection result of the audio frame is a speech category, and the face detection result of the first image frame belonging to the same time period as the audio frame is a face image, then the portion of the video clip belonging to the aforementioned time period is determined to be a real-person voice call video.

[0121] In an exemplary embodiment, such as Figure 5 As shown, Figure 5A flowchart of a singing video recognition method provided in an embodiment of this application is shown. Step 414 above includes the following sub-steps (4141-4145).

[0122] Step 4141: If the ratio of the total effective duration to the total duration of the video segment is greater than the ratio threshold, perform second image frame extraction processing on the original image frame sequence to obtain the second image frame sampling sequence.

[0123] The second image frame extraction process is used to extract a second fixed number of image frames within each time period. The second fixed number can be the same as or different from the first fixed number. Optionally, the second fixed number is greater than the first fixed number. For example, if the first image frame extraction process extracts one frame per second from the video to obtain a first image frame sampling sequence, then the second image frame extraction process extracts multiple frames per second from the video to obtain a second image frame sampling sequence. The specific values ​​of the first and second fixed numbers are not limited in this embodiment.

[0124] Optionally, if the first fixed quantity and the second fixed quantity are the same, steps 4141 and 4142 can be omitted, and step 4143 can be executed directly using the face parameters in step 410.

[0125] Step 4142: Perform face detection processing on the second image frame in the second image frame sampling sequence to obtain the face parameters of the second image frame.

[0126] Facial parameters include the coordinates of the corners of the mouth.

[0127] Step 4143: Determine the lip region image in the second image frame based on the corner of the mouth coordinate parameters.

[0128] Using the midpoint of the coordinates of the two corners of the mouth as the center and the maximum difference between the coordinates of the two corners of the mouth as the side length, the lip image of the main character in the second image frame is extracted, and each lip image is adjusted to a uniform pixel size, such as a 64*64 pixel image.

[0129] If a second image frame does not contain a main face, it can be directly replaced with a solid color image, or the second image frame containing a main face with the smallest time difference can be found by looking forward and backward in chronological order. The lip image is then extracted based on the parameters of the corresponding main face and used as the lip image corresponding to that second image frame.

[0130] Step 4144: Generate a lip image sequence based on the lip region image in the second image frame.

[0131] Optionally, the lip region images in the second image frame are arranged in chronological order to generate a lip image sequence.

[0132] Step 4145: Perform lip motion recognition processing on each lip region image in the lip image sequence to obtain the lip motion state.

[0133] Lip movement states characterize the results of lip action recognition, including speaking and singing states.

[0134] Based on the extracted lip image sequence, a deep learning model is used to determine whether the main character in the video is speaking or singing (speaking or singing), or in other states (expression, eating, drinking, etc.).

[0135] Optionally, the lip action recognition process described above can stitch together the sequence of equal-length lip images into a large lip state map. For example, a sequence of 25 64*64 lip images spanning 5 seconds can be stitched together into a large 320*320 lip state map to construct training and testing datasets for the aforementioned speaking and singing states (speaking, singing) and other states (facial expressions, eating, drinking, etc.). Then, a grouped convolutional residual network model (SE-ResNeXt) can be trained for image classification.

[0136] This application does not limit the method of lip movement recognition processing, and other similar video action recognition methods can also be used. For example, recognition methods based on 3D convolutional networks or recognition methods based on two-stream convolutional networks.

[0137] In an exemplary embodiment, such as Figure 5 As shown, step 415 above is replaced by step 4151 below.

[0138] Step 4151: When the lip movement is in a speaking or singing state, determine that the video clip is a singing video clip.

[0139] In one example, such as Figure 6As shown, this example illustrates a schematic diagram of a process for identifying singing videos. A long target video (recorded or live) is segmented into video segments of a certain duration. Audio is extracted from the video segments to separate audio files. Audio event detection is used to obtain a score for audio event categories per second, and the audio category per second is determined based on the score. In addition to audio, the video segments are sampled one frame per second to obtain a first image frame sampling sequence. Face detection is performed on the first image frame in the first image frame sampling sequence to obtain a face parameter matrix per second, thereby determining the main face parameters. If the audio category is singing without speech and a main face is present, a preliminary judgment is made to proceed to the next steps: multiple frames per second are sampled from the video segment to obtain a second image frame sampling sequence. Face detection is performed on the second image frame in the second image frame sampling sequence to obtain a face parameter matrix per second, thereby determining the main face parameters. Based on the main face parameters, lip images are determined, and a lip image sequence is generated. The lip image sequence is then used for speaking / singing state recognition. If the above speaking / singing state recognition can ultimately determine that the video segment is a singing video segment, then...

[0140] In summary, the technical solution provided in this application, by combining video multimodal information and employing methods such as audio event detection, face recognition, and lip movement recognition, can accurately identify segments of real people singing in videos. This allows for the discovery of high-quality recorded or live video segments, which can then be quickly pushed to users for viewing. This solution directly merges indistinguishable speech and singing in the video into a single category for recognition. Then, by adding voice category recognition to the audio, it filters out situations where there is background singing but the person is actually speaking. Each individual method used in this solution has a relatively high recognition accuracy and recall rate, and their combined use also achieves good recognition accuracy and recall rates for videos of real people singing.

[0141] The following are embodiments of the apparatus of this application, which can be used to execute embodiments of the method of this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the method of this application.

[0142] Please refer to Figure 7 This diagram illustrates a block diagram of a singing video recognition device according to an embodiment of this application. The device has the function of implementing the above-described singing video recognition method; this function can be implemented in hardware or by hardware executing corresponding software. The device can be a computer device or can be installed within a computer device. The device 700 may include: a video segment acquisition module 710, an audio / video feature detection module 720, a lip movement recognition module 730, and a video type determination module 740.

[0143] The video segment acquisition module 710 is used to acquire video segments corresponding to video data.

[0144] The audio and video feature detection module 720 is used to perform audio and video feature detection processing on the video segment to obtain the audio detection result and face detection result of the video segment.

[0145] The lip movement recognition module 730 is used to perform lip movement recognition processing on the video segment when the audio detection result and the face detection result meet the first condition, so as to obtain the lip movement recognition result corresponding to the video segment.

[0146] The video type determination module 740 is used to determine that the video segment is a singing video segment when the lip movement recognition result meets the second condition.

[0147] In an exemplary embodiment, the audio and video feature detection module 720 includes an audio detection unit and a face detection unit.

[0148] An audio detection unit is used to perform audio event detection processing on the audio data of the video segment to obtain the audio detection result of the video segment.

[0149] The face detection unit is used to perform face detection processing on the image data of the video segment to obtain the face detection result of the video segment.

[0150] In an exemplary embodiment, the audio detection unit is configured to:

[0151] Extract the audio data from the video segment;

[0152] The audio data is segmented into frames to obtain audio frames corresponding to each time period, and the duration of each time period is the preset duration.

[0153] The audio frame is subjected to audio event detection processing to obtain the score of the audio frame on each audio event. Each audio event includes a speech event and a singing event. The score includes a speech score and a singing score. The speech score represents the probability that the audio content of the audio frame belongs to the speech event, and the singing score represents the probability that the audio content of the audio frame belongs to the singing event.

[0154] If the speech score is less than the speech threshold and the singing score is greater than the singing threshold, the audio detection result of the audio frame is determined to be a category of singing without speech. The audio detection result of the audio frame represents the audio detection result of the video segment.

[0155] In an exemplary embodiment, the face detection unit is used for:

[0156] Obtain the original sequence of image frames of the video segment;

[0157] The original sequence of image frames is subjected to a first image frame extraction process to obtain a first image frame sampling sequence. The first image frame extraction process is used to extract a first fixed number of image frames in each time period.

[0158] Face detection processing is performed on the first image frame in the first image frame sampling sequence to obtain the face parameters of the first image frame. The first image frame is an image frame extracted from the original image frame sequence.

[0159] If the face parameters of the first image frame meet the threshold condition, the face detection result of the first image frame is determined to be a face image, and the face detection result of the first image frame represents the face detection result of the video segment.

[0160] In an exemplary embodiment, the lip movement recognition module 730 includes: an effective time period determination unit, an effective duration statistics unit, and a lip movement recognition unit.

[0161] The effective time period determination unit is used to determine the time period corresponding to the audio frame and the first image frame as an effective time period when the audio detection result of the audio frame is the category of singing without speech, and the face detection result of the first image frame belonging to the same time period as the audio frame is the face image.

[0162] The effective duration statistics unit is used to count the duration corresponding to the effective time period of the video segment and obtain the total effective duration.

[0163] The lip movement recognition unit is used to perform lip movement recognition processing on the video segment when the ratio of the total effective duration to the total duration of the video segment is greater than a ratio threshold, so as to obtain the lip movement recognition result corresponding to the video segment.

[0164] In an exemplary embodiment, the lip movement recognition unit is configured to:

[0165] If the ratio of the total effective duration to the total duration of the video segment is greater than the ratio threshold, the original sequence of image frames is subjected to a second image frame extraction process to obtain a second image frame sampling sequence. The second image frame extraction process is used to extract a second fixed number of image frames in each time period.

[0166] Face detection processing is performed on the second image frame in the second image frame sampling sequence to obtain the face parameters of the second image frame, including the corner of the mouth coordinate parameters;

[0167] Based on the corner of the mouth coordinate parameters, the lip region image in the second image frame is determined;

[0168] A sequence of lip images is generated based on the lip region image in the second image frame;

[0169] The lip movement recognition process is performed on each lip region image in the lip image sequence to obtain the lip movement state. The lip movement state represents the lip movement recognition result and includes the speaking and singing state.

[0170] Accordingly, the video type determination module 740 is used for:

[0171] If the lip movement is in a speaking or singing state, then the video clip is determined to be a singing video clip.

[0172] In an exemplary embodiment, the device further includes a singing video push module.

[0173] The performance video push module is used to splice together the performance video segments corresponding to the video data to generate a performance video; and to push the performance video in the information flow service.

[0174] In an exemplary embodiment, the video push module is further configured to determine the live video stream corresponding to the singing video segment and push the live video stream in the information flow service.

[0175] In an exemplary embodiment, the device further includes a video account tagging module.

[0176] The video account tagging module is used to determine the user account corresponding to the singing video clip and to tag the type of the user account.

[0177] In an exemplary embodiment, the video data includes recorded video data and live video data, and the video segment acquisition module 710 is used to:

[0178] If the video data is pre-recorded video data, the pre-recorded video is divided into at least one video segment;

[0179] If the video data is live video data, the live video data stream is received; the live video data stream is then processed to obtain the video segment.

[0180] In summary, the technical solution provided in this application combines video multimodal information to segment the video into segments and perform audio and video feature detection on the video segments to obtain audio detection results and face detection results of the video segments. Then, based on these two aspects of feature information, a preliminary judgment is made on the video segments. Lip action recognition is only performed if the first condition is met. If the lip action recognition result meets the second condition, the singing video segment can be finally identified. This effectively improves the recognition accuracy of singing videos while reducing the amount of computation.

[0181] It should be noted that the apparatus provided in the above embodiments is only illustrated by the division of the above functional modules when implementing its functions. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0182] Please refer to Figure 8 This diagram illustrates a structural block diagram of a computer device according to an embodiment of this application. The computer device is used to perform the aforementioned singing video recognition method. Specifically:

[0183] Computer device 800 includes a central processing unit (CPU) 801, a system memory 804 including random access memory (RAM) 802 and read-only memory (ROM) 803, and a system bus 805 connecting the system memory 804 and the CPU 801. Computer device 800 also includes a basic input / output system (I / O system) 806 that facilitates information transfer between various devices within the computer, and a mass storage device 807 for storing the operating system 813, application programs 814, and other program modules 812.

[0184] The basic input / output system 806 includes a display 808 for displaying information and an input device 809 for user input, such as a mouse or keyboard. Both the display 808 and the input device 809 are connected to the central processing unit 801 via an input / output controller 810 connected to the system bus 805. The basic input / output system 806 may also include the input / output controller 810 for receiving and processing input from multiple other devices such as a keyboard, mouse, or electronic stylus. Similarly, the input / output controller 810 also provides output to a display screen, printer, or other types of output devices.

[0185] Mass storage device 807 is connected to central processing unit 801 via a mass storage controller (not shown) connected to system bus 805. Mass storage device 807 and its associated computer-readable media provide non-volatile storage for computer device 800. That is, mass storage device 807 may include computer-readable media (not shown) such as hard disk or CD-ROM (CompactDisc Read-Only Memory) drive.

[0186] Without loss of generality, computer-readable media can include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented using any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media include RAM, ROM, EPROM (Erasable Programmable Read Only Memory), EEPROM (Electrically Erasable Programmable Read Only Memory), flash memory or other solid-state storage technologies, CD-ROM, DVD (Digital Video Disc) or other optical storage, magnetic tape cassettes, magnetic tape, disk storage, or other magnetic storage devices. Of course, those skilled in the art will recognize that computer storage media are not limited to the above-mentioned types. The system memory 804 and mass storage device 807 described above can be collectively referred to as memory.

[0187] According to various embodiments of this application, the computer device 800 can also be connected to a remote computer on a network, such as the Internet, for operation. That is, the computer device 800 can be connected to a network 812 via a network interface unit 811 connected to the system bus 805, or the network interface unit 811 can be used to connect to other types of networks or remote computer systems (not shown).

[0188] The memory also includes a computer program stored in the memory and configured to be executed by one or more processors to implement the above-described singing video recognition method.

[0189] In an exemplary embodiment, a computer-readable storage medium is also provided, the storage medium storing at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, the at least one program, the code set, or the instruction set, when executed by a processor, implements the above-described singing video recognition method.

[0190] Optionally, the computer-readable storage medium may include: ROM (Read Only Memory), RAM (Random Access Memory), SSD (Solid State Drives), or optical disc, etc. The random access memory may include ReRAM (Resistance Random Access Memory) and DRAM (Dynamic Random Access Memory).

[0191] In an exemplary embodiment, a computer program product or computer program is also provided, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the aforementioned singing video recognition method.

[0192] It should be understood that "multiple" as used herein refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. Furthermore, the step numbers described herein are merely illustrative of one possible execution order. In some other embodiments, the steps may not be executed in numerical order, such as two steps with different numbers being executed simultaneously, or two steps with different numbers being executed in the reverse order of the illustration. This application does not limit this.

[0193] The above description is merely an exemplary embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A method for recognizing singing videos, characterized in that, The method includes: Retrieve the video segment corresponding to the video data; The video segment is subjected to audio and video feature detection processing to obtain the audio detection result and face detection result of the video segment; If the audio detection result and the face detection result meet the first condition, lip movement recognition processing is performed on the video segment to obtain the lip movement recognition result corresponding to the video segment; the first condition is used to filter video segments that contain faces and have singing but no speech. If the lip movement recognition result meets the second condition, the video segment is determined to be a singing video segment; the second condition is used to filter out video segments that include lip movements that match the lip movement characteristics of a real person singing.

2. The method according to claim 1, characterized in that, The step of performing audio and video feature detection processing on the video segment to obtain the audio detection result and face detection result of the video segment includes: Audio event detection processing is performed on the audio data of the video segment to obtain the audio detection result of the video segment; The image data of the video segment is processed for face detection to obtain the face detection result of the video segment.

3. The method according to claim 2, characterized in that, The step of performing audio event detection processing on the audio data of the video segment to obtain the audio detection result of the video segment includes: Extract the audio data from the video segment; The audio data is processed by frame segmentation to obtain audio frames corresponding to each time period, and the duration of each time period is a preset duration. The audio frame is subjected to audio event detection processing to obtain the score of the audio frame on each audio event. Each audio event includes a speech event and a singing event. The score includes a speech score and a singing score. The speech score represents the probability that the audio content of the audio frame belongs to the speech event, and the singing score represents the probability that the audio content of the audio frame belongs to the singing event. If the speech score is less than the speech threshold and the singing score is greater than the singing threshold, the audio detection result of the audio frame is determined to be a category of singing without speech. The audio detection result of the audio frame represents the audio detection result of the video segment.

4. The method according to claim 3, characterized in that, The step of performing face detection processing on the image data of the video segment to obtain the face detection result of the video segment includes: Obtain the original sequence of image frames of the video segment; The original sequence of image frames is subjected to a first image frame extraction process to obtain a first image frame sampling sequence. The first image frame extraction process is used to extract a first fixed number of image frames in each time period. Face detection processing is performed on the first image frame in the first image frame sampling sequence to obtain the face parameters of the first image frame. The first image frame is an image frame extracted from the original image frame sequence. If the face parameters of the first image frame meet the threshold condition, the face detection result of the first image frame is determined to be a face image, and the face detection result of the first image frame represents the face detection result of the video segment.

5. The method according to claim 4, characterized in that, When the audio detection result and the face detection result meet the first condition, lip movement recognition processing is performed on the video segment to obtain the lip movement recognition result corresponding to the video segment, including: If the audio detection result of the audio frame is "singing but no speech", and the face detection result of the first image frame belonging to the same time period as the audio frame is the face image, then the time period corresponding to the audio frame and the first image frame is determined to be a valid time period. The total effective duration is obtained by calculating the duration of the effective time periods of the video segments. If the ratio of the total effective duration to the total duration of the video segment is greater than a ratio threshold, lip movement recognition processing is performed on the video segment to obtain the lip movement recognition result corresponding to the video segment.

6. The method according to claim 5, characterized in that, When the ratio of the total effective duration to the total duration of the video segment is greater than a ratio threshold, lip movement recognition processing is performed on the video segment to obtain the lip movement recognition result corresponding to the video segment, including: If the ratio of the total effective duration to the total duration of the video segment is greater than the ratio threshold, the original sequence of image frames is subjected to a second image frame extraction process to obtain a second image frame sampling sequence. The second image frame extraction process is used to extract a second fixed number of image frames in each time period. Face detection processing is performed on the second image frame in the second image frame sampling sequence to obtain the face parameters of the second image frame, including the corner of the mouth coordinate parameters; Based on the corner of the mouth coordinate parameters, determine the lip region image in the second image frame; A sequence of lip images is generated based on the lip region image in the second image frame; The lip movement recognition process is performed on each lip region image in the lip image sequence to obtain the lip movement state. The lip movement state represents the lip movement recognition result and includes the speaking and singing state. Accordingly, determining the video segment as a singing video segment when the lip movement recognition result meets the second condition includes: If the lip movement is in a speaking or singing state, then the video clip is determined to be a singing video clip.

7. The method according to any one of claims 1 to 3, characterized in that, The method further includes: The video data corresponding to the singing video segments are spliced ​​together to generate a singing video; the singing video is then pushed in the information flow service.

8. The method according to any one of claims 1 to 3, characterized in that, The method further includes: The live video stream corresponding to the performance video segment is determined, and the live video stream is pushed in the information flow service.

9. The method according to claim 1, characterized in that, The video data includes recorded video data and live video data, and the step of obtaining the video segment corresponding to the video data includes: If the video data is pre-recorded video data, the pre-recorded video is divided into at least one video segment; If the video data is live video data, the live video data stream is received; the live video data stream is processed to obtain the video segment.

10. A singing video recognition device, characterized in that, The device includes: The video clip acquisition module is used to acquire video clips corresponding to video data; The audio and video feature detection module is used to perform audio and video feature detection processing on the video segment to obtain the audio detection result and face detection result of the video segment; The lip movement recognition module is used to perform lip movement recognition processing on the video segment when the audio detection result and the face detection result meet the first condition, so as to obtain the lip movement recognition result corresponding to the video segment; the first condition is used to filter video segments that contain faces and have singing but no speech. The video type determination module is used to determine that the video segment is a singing video segment if the lip movement recognition result meets the second condition; the second condition is used to filter out video segments that include lip movements that match the lip movement characteristics of a real person singing.

11. The apparatus according to claim 10, characterized in that, The audio and video feature detection module includes: an audio detection unit and a face detection unit; The audio detection unit is used to perform audio event detection processing on the audio data of the video segment to obtain the audio detection result of the video segment; The face detection unit is used to perform face detection processing on the image data of the video segment to obtain the face detection result of the video segment.

12. The apparatus according to claim 11, characterized in that, The audio detection unit is used for: Extract the audio data from the video segment; The audio data is processed by frame segmentation to obtain audio frames corresponding to each time period, and the duration of each time period is a preset duration. The audio frame is subjected to audio event detection processing to obtain the score of the audio frame on each audio event. Each audio event includes a speech event and a singing event. The score includes a speech score and a singing score. The speech score represents the probability that the audio content of the audio frame belongs to the speech event, and the singing score represents the probability that the audio content of the audio frame belongs to the singing event. If the speech score is less than the speech threshold and the singing score is greater than the singing threshold, the audio detection result of the audio frame is determined to be a category of singing without speech. The audio detection result of the audio frame represents the audio detection result of the video segment.

13. The apparatus according to claim 12, characterized in that, The face detection unit is used for: Obtain the original sequence of image frames of the video segment; The original sequence of image frames is subjected to a first image frame extraction process to obtain a first image frame sampling sequence. The first image frame extraction process is used to extract a first fixed number of image frames in each time period. Face detection processing is performed on the first image frame in the first image frame sampling sequence to obtain the face parameters of the first image frame. The first image frame is an image frame extracted from the original image frame sequence. If the face parameters of the first image frame meet the threshold condition, the face detection result of the first image frame is determined to be a face image, and the face detection result of the first image frame represents the face detection result of the video segment.

14. The apparatus according to claim 13, characterized in that, The lip movement recognition module includes: an effective time period determination unit, an effective duration statistics unit, and a lip movement recognition unit; The effective time period determination unit is used to determine the time period corresponding to the audio frame and the first image frame as an effective time period when the audio detection result of the audio frame is the category of singing without speech, and the face detection result of the first image frame belonging to the same time period as the audio frame is the face image. The effective duration statistics unit is used to count the duration corresponding to the effective time period of the video segment and obtain the total effective duration; The lip movement recognition unit is used to perform lip movement recognition processing on the video segment when the ratio of the total effective duration to the total duration of the video segment is greater than a ratio threshold, so as to obtain the lip movement recognition result corresponding to the video segment.

15. The apparatus according to claim 14, characterized in that, The lip movement recognition unit is used for: If the ratio of the total effective duration to the total duration of the video segment is greater than the ratio threshold, the original sequence of image frames is subjected to a second image frame extraction process to obtain a second image frame sampling sequence. The second image frame extraction process is used to extract a second fixed number of image frames in each time period. Face detection processing is performed on the second image frame in the second image frame sampling sequence to obtain the face parameters of the second image frame, including the corner of the mouth coordinate parameters; Based on the corner of the mouth coordinate parameters, determine the lip region image in the second image frame; A sequence of lip images is generated based on the lip region image in the second image frame; The lip movement recognition process is performed on each lip region image in the lip image sequence to obtain the lip movement state. The lip movement state represents the lip movement recognition result and includes the speaking and singing state. Accordingly, the video type determination module is used for: If the lip movement is in a speaking or singing state, then the video clip is determined to be a singing video clip.

16. The apparatus according to any one of claims 10-12, characterized in that, The device also includes a singing video push module; The singing video push module is used to splice together the singing video segments corresponding to the video data to generate a singing video; and push the singing video in the information flow service.

17. The apparatus according to any one of claims 10-12, characterized in that, The video push module is also used to determine the live video stream corresponding to the performance video segment and push the live video stream in the information flow service.

18. The apparatus according to claim 10, characterized in that, The video data includes recorded video data and live video data. The video segment acquisition module is used for: If the video data is pre-recorded video data, the pre-recorded video is divided into at least one video segment; If the video data is live video data, receive the live video data stream; The live video data stream is segmented to obtain the video segment.

19. A computer device, characterized in that, The computer device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set, or an instruction set, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the singing video recognition method as described in any one of claims 1-9.

20. A computer-readable storage medium, characterized in that, The storage medium stores at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, the at least one program, the code set, or instruction set is loaded and executed by a processor to implement the singing video recognition method as described in any one of claims 1-9.

21. A computer program product, characterized in that, The computer program product includes computer instructions stored in a computer-readable storage medium; a processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the singing video recognition method as described in any one of claims 1-9.

Citation Information

Patent Citations

  • Singing recognition method and device, equipment and storage medium

    CN114022950A