Video Processing Method, Apparatus, Electronic Device, and Computer-Readable Storage Medium
By dividing the video with audio information and performing facial recognition and speech status analysis, the problem of human-relying speech object annotation in the video is solved, and efficient and accurate speech object recognition is achieved.
Patent Information
- Application Number
- CN202110438512.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-04-22
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2041-04-22
AI Technical Summary
In the prior art, the annotation of the speaking object in the video requires a lot of manual intervention, resulting in waste of resources and inefficiency.
The video is segmented through audio information, image frames are extracted for facial recognition and speech state analysis, and the speaking object is determined in combination with multimodal analysis methods.
It improves the recognition efficiency and accuracy of speaking objects in the video, reduces manual intervention, simplifies the difficulty of object recognition, and improves the accuracy of dialogue object recognition.
Smart Images

Figure CN113762036B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of video processing, and particularly to a video processing method, apparatus, electronic device, and computer-readable storage medium. Background Art
[0002] With the development of the video industry, in some application scenarios, it is necessary to label the speaking object in the video to be recognized in order to better meet the user requirements.
[0003] For example, in a speech contest video, it may be necessary to label the speaking object in each time period. For another example, in a debate contest, it may be necessary to label the speaking object in each time period. Still for another example, in an interview program, it may be necessary to label the speaking object at each time point.
[0004] If the above-mentioned speaking object labeling work is done manually, it will waste a lot of human resources.
[0005] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present disclosure. Summary of the Invention
[0006] The purpose of the present disclosure is to provide a video processing method, apparatus, and electronic device, which can automatically and accurately determine the speaking object in the video to be recognized, improving the recognition efficiency and accuracy.
[0007] Other features and advantages of the present disclosure will become apparent through the following detailed description, or be learned in part through the practice of the present disclosure.
[0008] An embodiment of the present disclosure provides a video processing method, including: using the audio information of the video to be recognized to segment the video to be recognized, obtaining at least one video segment of the video to be recognized, where each video segment includes at least one speaking object; extracting at least two frames of images from each video segment; performing face recognition processing on each frame of image in each video segment to obtain at least one face image to be recognized from each frame of image in each video segment; performing speaking state recognition on each face image to be recognized in each frame of image in each video segment to determine the speaking state recognition result of each face image to be recognized; and determining the speaking object in each video segment according to the speaking state recognition result of the at least one face image to be recognized in the at least two frames of images in each video segment.
[0009] An embodiment of the present disclosure provides a video processing apparatus, including: a video segmentation module, a frame extraction module, a face recognition module, a speaking state recognition module, and a speaking object determination module.
[0010] Among them, the video segmentation module is used to segment the video to be recognized by using the audio information of the video to be recognized, so as to obtain at least one video segment of the video to be recognized, and each video segment includes at least one speaking object. The frame extraction module is used to extract at least two frames of images from each video segment. The face recognition module is used to perform face recognition processing on each frame of image in each video segment, so as to obtain at least one face image to be recognized from each frame of image in each video segment. The speaking state recognition module is used to perform speaking state recognition on each face image to be recognized in each frame of image in each video segment respectively, so as to determine the speaking state recognition result of each face image to be recognized. The speaking object determination module is used to determine the speaking object in each video segment according to the speaking state recognition results of the at least one face image to be recognized in the at least two frames of images in each video segment.
[0011] In some embodiments, at least one video segment of the video to be recognized includes a target video segment. Among them, the speaking object determination module includes: an object recognition processing sub-module, a clustering sub-module, a target face image to be recognized clustering determination sub-module, and a speaking object determination sub-module.
[0012] Among them, the object recognition processing sub-module is used to perform object recognition processing on each face image to be recognized in each frame of image in the target video segment, and determine the object category of each face image to be recognized. The clustering sub-module is used to cluster the at least one face image to be recognized in the at least two frames of images in the target video segment according to the object category, so as to obtain the face image to be recognized clustering in the target video segment under each object category. The target face image to be recognized clustering determination sub-module is used to determine, according to the face image to be recognized clustering in the target video segment under each object category and the speaking state recognition results of the at least one face image to be recognized in the at least two frames of images in the target video segment, the face image to be recognized clustering with the largest number of face images to be recognized whose speaking state recognition results are the speaking state as the target face image to be recognized clustering. The speaking object determination sub-module is used to determine that the object category corresponding to the target face image to be recognized clustering is the speaking object in the target video segment.
[0013] In some embodiments, at least one video segment of the video to be recognized includes a target video segment, at least two frames of images of the target video segment include target frame images, and at least one face image to be recognized of the target frame images includes a target face image to be recognized. Among them, the speaking state recognition module may include: a feature extraction sub-module, a feature comparison sub-module, and a speaking state label determination sub-module.
[0014] Among them, the feature extraction sub-module is used to extract features from the target face image to be recognized through a speaker recognition model, so as to obtain a target face feature vector of the target face image to be recognized. The speaker recognition model is trained using face image samples and the speech state labels of the face image samples. The feature comparison sub-module is used to compare the target face feature vector to be recognized with the recognized face feature vectors of the face image samples to determine the similarity between the target face feature vector to be recognized and each recognized face feature vector. The speech state label determination sub-module is used to determine the speech state label of the target face image according to the similarity between the target face feature vector to be recognized and each recognized face feature vector.
[0015] In some embodiments, the speech state label determination sub-module includes: a speech state level determination unit, a similarity distribution determination unit, and a speech state label determination unit.
[0016] Among them, the speech state level determination unit is used to determine the speech state level of each face image sample. The similarity distribution determination unit is used to determine the similarity distribution between the target face feature vector to be recognized and each recognized face feature vector according to the speech state level of each face image sample and the similarity between the target face feature vector to be recognized and each recognized face feature vector. The speech state label determination unit is used to determine the speech state label of the target face image according to the similarity distribution between the target face feature vector to be recognized and each recognized face feature vector.
[0017] In some embodiments, the speech state label determination unit includes: a unimodal distribution determination sub-unit, a distribution peak determination sub-unit, a speech state determination sub-unit, and a non-speech state determination sub-unit.
[0018] Among them, the unimodal distribution determination sub-unit is used to determine that the similarity distribution between the target face feature vector to be recognized and each recognized face feature vector is a unimodal distribution. The distribution peak determination sub-unit is used to obtain the distribution peak of the similarity distribution between the target face feature vector to be recognized and each recognized face feature vector. The speech state determination sub-unit is used to determine that the speech state label of the target face image corresponding to the target face feature vector is a speech state if the speech state level of the face image sample corresponding to the distribution peak position is greater than a first threshold. The non-speech state determination sub-unit is used to determine that the speech state label of the target face image corresponding to the target face feature vector is a non-speech state if the speech state level of the face image sample corresponding to the distribution peak position is lower than the first threshold.
[0019] In some embodiments, the speech state recognition module further includes: an initial training image acquisition sub-module, a speech object recognition intermediate model training sub-module, a training image to be labeled acquisition sub-module, a training vector to be labeled acquisition sub-module, a similarity determination sub-module, a speech state label determination sub-module, and a transfer training sub-module.
[0020] Among them, the initial training image acquisition sub-module is used to acquire the initial training image and the speech state label of the initial training image. The speech object recognition intermediate model training sub-module is used to train a target neural network according to the initial training image and the speech state label of the initial training image to obtain a speech object recognition intermediate model. The training image to be labeled acquisition sub-module is used to acquire the training image to be labeled. The training vector to be labeled acquisition sub-module is used to extract features from the training image to be labeled through the speech object recognition intermediate model to obtain a training vector to be labeled. The similarity determination sub-module is used to perform feature comparison between the training vector to be labeled and the recognized facial feature vectors of the facial image samples to determine the similarity between the training vector to be labeled and each recognized facial feature vector. The speech state label determination sub-module is used to determine the speech state label of the training image to be labeled according to the similarity between the training vector to be labeled and each recognized facial feature vector. The transfer training sub-module is used to perform transfer training on the speech object recognition intermediate model according to the training image to be labeled and the speech state label of the training image to be labeled to obtain the speech object recognition model.
[0021] In some embodiments, the speech state label determination sub-module may include: a speech state level determination unit for facial image samples, a similarity distribution determination unit for recognized facial feature vectors, and a speech state label determination unit for the training image to be labeled.
[0022] Among them, the speech state level determination unit for facial image samples determines the speech state level of each facial image sample. The similarity distribution determination unit for recognized facial feature vectors is used to determine the similarity distribution between the training vector to be labeled and each recognized facial feature vector according to the speech state level of each facial image sample and the similarity between the training vector to be labeled and each recognized facial feature vector. The speech state label determination unit for the training image to be labeled is used to determine the speech state label of the training image to be labeled according to the similarity distribution between the training vector to be labeled and each recognized facial feature vector.
[0023] In some embodiments, the speech state label determination sub-module may include: an initial confidence determination unit, a target confidence determination unit, and a speech state label determination unit for the training image to be labeled.
[0024] Among them, the initial confidence determination unit is used to determine the initial confidence of the to-be-annotated training vector relative to the recognized facial feature vectors according to the speaking state labels of the facial image samples corresponding to the recognized facial feature vectors and the similarity between each recognized facial feature vector and the to-be-annotated training vector. The target confidence determination unit is used to determine the target confidence of the to-be-annotated training vector relative to the recognized facial feature vectors according to the initial confidence and the speaking state labels of the recognized facial feature vectors. The speaking state label determination unit of the to-be-annotated training image is used to determine the speaking state label of the to-be-annotated training image according to the target confidence.
[0025] In some embodiments, the speaking state label determination unit of the to-be-annotated training image may include: a first threshold judgment sub-unit, a second threshold judgment sub-unit, and a third threshold judgment sub-unit.
[0026] Among them, the first threshold judgment sub-unit is used to determine that the speaking state label of the to-be-annotated training object corresponding to the to-be-annotated training vector is the speaking state if the target confidence is greater than the second threshold. The second threshold judgment sub-unit is used to determine that the speaking state label of the to-be-annotated training object corresponding to the to-be-annotated training vector is the non-speaking state if the target confidence is less than the third threshold. The third threshold judgment sub-unit is used to determine that the speaking state of the to-be-annotated training object corresponding to the to-be-annotated training vector is unknown if the target confidence is greater than or equal to the third threshold and less than or equal to the second threshold, so as to label the to-be-annotated training image with an unknown speaking state through manual recognition.
[0027] An embodiment of the present disclosure provides an electronic device, which includes: one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the video processing method described in any one of the above.
[0028] An embodiment of the present disclosure provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the video processing method described in any one of the above.
[0029] An embodiment of the present disclosure provides a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device implements the above video processing method.
[0030] The video processing method provided by the embodiments of the present disclosure, on the one hand, combines the multi-modal analysis method to use audio information to segment the video to be recognized into video segments with less picture information, and then performs speaker recognition on each video segment, which simplifies the difficulty of speaker recognition and improves the recognition accuracy of the dialogue objects in the video to be recognized; on the other hand, the present embodiment performs speaker state recognition processing on each face image to be recognized in each video segment of the video to be recognized, and then determines the speaker in the video segment based on at least one face image to be recognized in at least two frames of images in the video segment. This method not only improves the accuracy of recognizing the speaker state of the video segment by static classification processing of each face image to be recognized, thereby improving the accuracy of recognizing the speaker in the video to be recognized, but also determines the speaker in the video segment by processing multiple image information such as at least one face image to be recognized in at least two frames of images, improving the accuracy of determining the speaker in the video segment, and thus improving the accuracy of recognizing the speaker in the video to be recognized.
[0031] It should be understood that the above general description and the following detailed description are only exemplary and do not limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] The drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure. Obviously, the drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0033] Figure 1 The figure shows a schematic diagram of an exemplary system architecture to which the video processing method or video processing apparatus of the embodiments of the present disclosure is applied.
[0034] Figure 2 The figure shows a schematic diagram of the structure of an electronic device suitable for implementing the embodiments of the present disclosure.
[0035] Figure 3 is a flowchart of a video processing method shown according to an exemplary embodiment.
[0036] Figure 4 is a flowchart of a method for classifying audio information by voice recognition technology shown according to an exemplary embodiment.
[0037] Figure 5 is a flowchart of a method for determining a speaker state recognition result for a target face image to be recognized shown according to an exemplary embodiment.
[0038] Figure 6It is a flowchart of a method for determining a speech state label of a target face image to be recognized according to similarity, shown in accordance with an exemplary embodiment.
[0039] Figure 7 It is a flowchart of a method for determining a speaking object in a target video segment, shown in accordance with an exemplary embodiment.
[0040] Figure 8 It is a training method of a speaking object recognition model, shown in accordance with an exemplary embodiment.
[0041] Figure 9 It is a schematic diagram of a face image, shown in accordance with an exemplary embodiment.
[0042] Figure 10 It is according to Figure 8 The shown embodiment gives a physical block diagram corresponding to a training method of a speaking object recognition model.
[0043] Figure 11 It is a flowchart of a method for determining a speech state label of a training image to be labeled according to the similarity between the training vector to be labeled and each recognized face feature vector, shown in accordance with an exemplary embodiment.
[0044] Figure 12 It is a structural diagram of a video processing method, shown in accordance with an exemplary embodiment.
[0045] Figure 13 It is a flowchart of a method for performing speaking object recognition on a video segment, shown in accordance with an exemplary embodiment.
[0046] Figure 14 It is a block diagram of a video processing apparatus, shown in accordance with an exemplary embodiment. Detailed implementation manners
[0047] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the concept of the example embodiments to those skilled in the art. Identical reference numerals in the figures denote identical or similar parts, and thus their repetitive description will be omitted.
[0048] The features, structures, or characteristics described in this disclosure may be combined in one or more embodiments in any suitable manner. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of this disclosure. However, those skilled in the art will realize that one or more of the specific details may be omitted when implementing the technical solutions of this disclosure, or other methods, components, devices, steps, etc. may be adopted. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of this disclosure.
[0049] The accompanying drawings are only schematic illustrations of this disclosure, and the same reference numerals in the drawings denote the same or similar parts, thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0050] The flowcharts shown in the accompanying drawings are only exemplary illustrations, and do not necessarily include all the contents and steps, nor do they necessarily need to be executed in the described order. For example, some steps may be decomposed, while some steps may be combined or partially combined, so the actual execution order may change according to the actual situation.
[0051] In this specification, the terms "a", "an", "the", "said", and "at least one" are used to indicate the existence of one or more elements / components / etc.; the terms "comprising", "including", and "having" are used to mean an open inclusion and refer to the existence of additional elements / components / etc. in addition to the listed elements / components / etc.; the terms "first", "second", "third", etc. are only used as labels and are not a limitation on the quantity of their objects.
[0052] The exemplary embodiments of this disclosure will be described in detail below with reference to the accompanying drawings.
[0053] Figure 1 A schematic diagram of an exemplary system architecture to which the video processing method or video processing device according to the embodiments of this disclosure can be applied is shown.
[0054] As Figure 1 shown, the system architecture 100 may include terminal devices 101, 102, or 103, a network 104, and a server 105. The network 104 is used to provide a medium for a communication link between the terminal devices 101, 102, or 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0055] Users can interact with the server 105 through the network 104 using the terminal devices 101, 102, or 103 to receive or send messages, etc. For example, users can send the video to be recognized or video segments segmented from the video to be recognized to the server 105 through the terminal devices 101, 102, or 103. Among them, the terminal devices 101, 102, or 103 can be various electronic devices with a display screen and supporting web browsing, including but not limited to smartphones, tablets, laptop computers, desktop computers, wearable devices, virtual reality devices, smart homes, and so on.
[0056] The server 105 can be a server that provides various services, such as a back-end management server that supports the devices operated by users using the terminal devices 101, 102, or 103. The back-end management server can analyze and process data such as received requests, and feedback the processing results to the terminal devices.
[0057] The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The present disclosure does not limit this.
[0058] The server 105 can, for example, segment the video to be recognized using the audio information of the video to be recognized to obtain at least one video segment of the video to be recognized, and each video segment includes at least one speaking object; the server 105 can, for example, extract at least two frames of images from each video segment; the server 105 can, for example, perform face recognition processing on each frame of image in each video segment to obtain at least one face image to be recognized from each frame of image in each video segment; the server 105 can, for example, perform speaking state recognition on each face image to be recognized in each frame of image in each video segment to determine the speaking state recognition result of each face image to be recognized; the server 105 can, for example, determine the speaking object in each video segment according to the speaking state recognition results of the at least one face image to be recognized in the at least two frames of images in each video segment.
[0059] After the server 105 determines the speaking object in each video segment, it will send the speaking object recognition result of each video segment to the terminal device for display.
[0060] It should be understood, Figure 1The numbers of the terminal devices, networks, and servers in it are merely illustrative. The server 105 can be a physical server or composed of multiple servers. According to actual needs, there can be any number of terminal devices, networks, and servers.
[0061] Figure 2 FIG. shows a schematic structural diagram of an electronic device suitable for use as a terminal device or a server for implementing the embodiments of the present disclosure. It should be noted that, Figure 2 The illustrated electronic device 200 is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.
[0062] As Figure 2 shown, the electronic device 200 includes a central processing unit (CPU) 201, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 202 or the program loaded from the storage section 208 into the random access memory (RAM) 203. In the RAM 203, various programs and data required for the operation of the electronic device 200 are also stored. The CPU 201, ROM 202, and RAM 203 are connected to each other via a bus 204. The input / output (I / O) interface 205 is also connected to the bus 204.
[0063] The following components are connected to the I / O interface 205: an input section 206 including a keyboard, a mouse, etc.; an output section 207 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 208 including a hard disk, etc.; and a communication section 209 including a network interface card such as a LAN card, a modem, etc. The communication section 209 performs communication processing via a network such as the Internet. A drive 210 is also connected to the I / O interface 205 as needed. A removable medium 211, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 210 as needed so that a computer program read from it can be installed into the storage section 208 as needed.
[0064] Specifically, according to the embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments of the present disclosure include a computer program product, which includes a computer program carried on a computer-readable storage medium, and the computer program contains program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network through the communication section 209 and / or installed from the removable medium 211. When the computer program is executed by the central processing unit (CPU) 201, the above functions defined in the system of the present application are executed.
[0065] It should be noted that the computer-readable storage medium shown in this disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In this application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable storage medium other than the computer-readable storage medium, and this computer-readable storage medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable storage medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0066] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0067] The modules and / or sub-modules and / or units and / or sub-units involved in the embodiments of the present application can be implemented in software or in hardware. The described modules and / or sub-modules and / or units and / or sub-units can also be provided in a processor. For example, it can be described as: a processor includes a sending unit, an obtaining unit, a determining unit, and a first processing unit. Among them, the names of these modules and / or sub-modules and / or units and / or sub-units do not, in some cases, constitute a limitation on the modules and / or sub-modules and / or units and / or sub-units themselves.
[0068] As another aspect, the present application also provides a computer-readable storage medium, which can be included in the device described in the above embodiments; or it can exist alone without being assembled into the device. The above computer-readable storage medium carries one or more programs. When the above one or more programs are executed by the device, the functions that the device can implement include: using the audio information of the video to be recognized to segment the video to be recognized, obtaining at least one video segment of the video to be recognized, and each video segment includes at least one speaking object; extracting at least two frames of images from each video segment; performing face recognition processing on each frame of image in each video segment to obtain at least one face image to be recognized from each frame of image in each video segment; performing speaking state recognition on each face image to be recognized in each frame of image in each video segment to determine the speaking state recognition result of each face image to be recognized; and determining the speaking object in each video segment according to the speaking state recognition results of at least one face image to be recognized in at least two frames of images in each video segment.
[0069] The embodiments of the present disclosure provide a method for accurately and efficiently determining a speaking object in a video to be recognized.
[0070] Figure 3 is a flowchart of a video processing method shown according to an exemplary embodiment. The method provided by the embodiments of the present disclosure can be executed by any electronic device with computing and processing capabilities. For example, the method can be executed by the server or the terminal device in the above Figure 1 embodiments, or can be jointly executed by the server and the terminal device. In the following embodiments, the server is taken as an example of the execution entity for illustration, but the present disclosure is not limited thereto.
[0071] Referring to Figure 3 , the video processing method provided by the embodiments of the present disclosure may include the following steps.
[0072] Step S31, using the audio information of the video to be recognized to segment the video to be recognized, obtaining at least one video segment of the video to be recognized, and each video segment includes at least one speaking object.
[0073] The above-mentioned speaking object can be a person object in the real world or a fictional anime object in the anime world (such as a cat, a dog, or other fictional speakable objects), and the present disclosure does not limit this.
[0074] The video to be recognized can refer to a video including a speaking object. The video to be recognized can include one speaking object or multiple speaking objects; within the same time period, there can be one speaking object in the video to be recognized (such as a single-person speech video), or multiple speaking objects can exist simultaneously (such as a shopping mall surveillance video). The video to be recognized can include one speaking topic or different speaking topics, etc., and the present disclosure does not limit this.
[0075] In order to simplify the complexity of the speaking relationships of the speaking objects in the video to be recognized and reduce the difficulty of object recognition, in this embodiment, audio information is used to segment the video to be recognized to obtain video segments with relatively simple speaking relationships (such as video segments including only one speaking object, or video segments discussing only one topic, or video segments interviewing only one person, etc.).
[0076] For example, the audio information of a multi-person speech video can be used to segment the multi-person speech video to obtain single-person speech video segments; for example, the audio information of an interview video can be used to segment the interview video to obtain single-person speaking segments; for another example, the audio information in a conference video can be used to segment the conference video to obtain conference segments for a single topic; it can be understood that those skilled in the art can set the video segmentation results according to the actual needs of the user, and the present disclosure does not limit this.
[0077] In some embodiments, using the audio information of the video to be recognized to segment the video to be recognized may include:
[0078] Using the timbre in the audio information to segment the video to be recognized into multiple video segments, each video segment including a speaking object with the same speaking timbre; or using the speech content in the audio information to segment the video to be recognized into multiple video segments, and the speaking objects in each video segment all speak around the same topic; or using speech recognition technology to classify the audio information to determine different audio segments corresponding to different speaking objects, and then using the different audio segments to perform segmentation processing on the video to be recognized so that each video segment includes one speaking object. The present disclosure does not limit the method of segmenting the video to be recognized through audio information, and those skilled in the art can set different segmentation methods according to their actual needs.
[0079] In some embodiments, automatic speech recognition technology, such as ASR (Automatic Speech Recognition), can be used to segment the audio of the video to be recognized, and then the video to be recognized is segmented according to the segmented audio to obtain at least one video segment.
[0080] Among them, the video segmentation process can be completed by the steps as Figure 4 shown: Step S41, receive the audio information of the video to be recognized; Step S42, perform feature extraction processing on the audio information through the speech recognition technology ASR to obtain speech features; Step S43, perform decoding processing on the speech features through the speech recognition technology ASR to implement the segmentation processing of the audio information; Step S44, then segment the video to be recognized according to the segmentation result of the audio information to obtain at least one video segment, so as to segment the video to be recognized according to the segmented audio information.
[0081] Step S32, extract at least two frames of images from each video segment.
[0082] In some embodiments, frame extraction processing can be performed on each video segment to obtain at least two frames of images from each video segment. For example, frame extraction processing can be performed on each video segment at a certain time interval or at an interval of a certain number of frames to obtain at least two frames of images (for example, 1 frame can be extracted every 20 frames), or frame extraction processing can be randomly performed on each video segment to obtain at least two frames of images. The present disclosure does not limit this.
[0083] Among them, frame extraction is an image processing method of extracting several frames of images from a video at a certain frame interval or a certain period of time.
[0084] Step S33, perform face recognition processing on each frame of image in each video segment to obtain at least one face image to be recognized from each frame of image in each video segment.
[0085] Face recognition processing is to cut out a face image including face information from the frame image to be recognized.
[0086] In some embodiments, at least one face image to be recognized can be obtained from each frame of image through pixel-level image detection technology (traditional non-machine learning image processing method); face recognition processing can also be performed on each frame of image through a trained face recognition neural network to obtain at least one face image to be recognized from each frame of image. The present disclosure does not limit the face recognition processing method.
[0087] Among them, at least one face image to be recognized may be included in one frame of image, and the face image of one object to be recognized may be included in one face image to be recognized.
[0088] Step S34: For each face image to be recognized in each frame of image in each video segment, perform a speaking state recognition respectively to determine the speaking state recognition result of each face image to be recognized.
[0089] In some embodiments, the speaking state of the face image to be recognized can be determined by an image detection method at the pixel level (traditional non-machine learning image processing method), or the speaking state of each face image to be recognized can be determined by a trained speaking state recognition neural network.
[0090] Step S35: According to the speaking state recognition results of at least one face image to be recognized in at least two frames of images in each video segment, determine the speaking object in each video segment.
[0091] In the above speaking state recognition results, the speaking state or non-speaking state of the face image to be recognized can be shown through a speaking state label.
[0092] In some embodiments, the category of the speaking object in each face image to be recognized can be determined, then the face images to be recognized are classified according to the object category included in the face image, then the number of face images to be recognized in the speaking state under each object category is counted, and finally the speaking object in each video segment is determined according to the statistical result.
[0093] The above object category can refer to different human object categories (the different human object categories can include, for example, Person A, Person B, Person C), or can refer to different anime object categories (the different anime object categories can include, for example, Anime D, Anime Character E), etc. The present disclosure does not limit this.
[0094] For example, if it is determined that there may be two speaking objects in a certain video segment, the first object corresponding to the object category with the largest number of face images to be recognized in the speaking state in this video segment can be used as the first speaking object of this video segment; the second object corresponding to the object category with the second largest number of face images to be recognized in the speaking state is used as the second speaking object of this video segment. Suppose that in the video segment, the number of face images to be recognized in the speaking state of the first object S is the largest, and the number of face images to be recognized in the speaking state of the second object P is the second largest, then the first object S is the first speaking object in this video segment, and the second object P is the second speaking object in this video segment.
[0095] The present disclosure does not limit the method for determining the speaking object in each video segment according to the speaking state recognition result.
[0096] The technical solution provided in this embodiment analyzes the speaking object in the video through the method of temporal filtering. That is, for a video to be recognized, single images are obtained through frame extraction, and then static image classification is performed. Then, from the classification results in a video segment, the object category with the most occurrences of the face is selected, and the object corresponding to this object category is determined as the speaking object, achieving the effect of recognizing the speaking object. By using the above method, not only the efficiency of recognizing the speaking object is improved, but also the accuracy of recognizing the speaking object is improved. Even for a relatively small-sized face, a relatively accurate recognition result can be obtained through the above method.
[0097] Among them, temporal filtering is to filter out some redundant frame images in time sequence to obtain the key frame images in each video segment.
[0098] The technical solution provided in this embodiment, on the one hand, accurately recognizes the speaking state of each face image to be recognized through a single static image; on the other hand, through temporal filtering analysis, the speaking object is accurately determined for the video segment including multiple face images to be recognized. In short, the technical solution provided in this embodiment improves the efficiency of recognizing the speaking object in the video to be recognized while also improving the accuracy of recognizing the speaking object.
[0099] The technical solution provided in this embodiment, on the one hand, combines the multi-modal analysis method to use audio information to divide the video to be recognized into video segments with less picture information, and then recognizes the speaking object for each video segment, simplifying the difficulty of recognizing the speaking object and improving the accuracy of recognizing the conversation object in the video to be recognized; on the other hand, this embodiment performs speaking state recognition processing on each face image to be recognized in each video segment of the video to be recognized, and then determines the speaking object in the video to be recognized based on at least one face image to be recognized in at least two frames of images in the video segment. This method not only improves the accuracy of recognizing the speaking state of the video segment through the static classification processing of each face image to be recognized, thereby improving the accuracy of recognizing the speaking object in the video to be recognized, but also improves the accuracy of determining the speaking object in the video segment through the determination processing of the speaking object for the video segment using multiple image information such as at least one face image to be recognized in at least two frames of images, thereby improving the accuracy of recognizing the speaking object in the video to be recognized.
[0100] There are many video shares of excellent students in internal and external live broadcasts. In addition to single-person speech videos, there are also a large number of videos in the form of interviews and face-to-face exchanges. In these videos, the multi-person conversation scenario is mainly involved.
[0101] For these multi-person scenarios, identifying the speakers therein plays an important role in applications such as video understanding and video segmentation. We hope to effectively identify the speakers through multi-modal analysis of the data in the video and apply the results to other scenarios. For example, the speaker segments in the video can be extracted and associated with the corresponding person tags. When searching for relevant classmates, the sharing segments of the classmate in the video can be automatically matched, facilitating other classmates to obtain information about the people they are interested in. On the other hand, the video can be effectively segmented through the technical solution provided in this embodiment, saving the time for watching the video. In some audio and video conferences, corresponding functions such as person focusing can also be achieved through speaker recognition technology.
[0102] In some embodiments, at least one video segment of the video to be recognized may include a target video segment. At least two frames of images in the target video segment include target frame images, and at least one face image to be recognized in the target frame images includes a target face image to be recognized.
[0103] Figure 5 It is a flowchart of a method for determining a speech state recognition result for a target face image to be recognized shown according to an exemplary embodiment.
[0104] In some embodiments, a machine learning method in the field of artificial intelligence technology can be used to determine a speech state recognition result for the target face image to be recognized.
[0105] Among them, artificial intelligence technology is a comprehensive discipline, involving a wide range of fields, including both hardware-level technologies and software-level technologies. Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0106] Machine learning is a multi-field interdisciplinary subject, involving multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.
[0107] In some embodiments, a target neural network can be trained with a face image sample and a speech state label of the face image sample to obtain a speaker recognition model.
[0108] Among them, the target neural network can be any neural network capable of facial recognition, such as a convolutional neural network, a recurrent neural network, or a temporal neural network, and the present disclosure does not limit this.
[0109] Among them, the facial image sample can be an image including the facial information of the training object, and the speech state label of the facial image sample can be used to indicate whether the training object in the facial image is speaking. If the training object in the facial image sample is speaking, the speech state label of the facial image sample indicates that the facial image sample is in the speaking state; if the training object in the facial image sample is not speaking, the speech state label of the facial image sample indicates that the facial image sample is in the non-speaking state.
[0110] Among them, the number of the facial image samples can be one or multiple, and the present disclosure does not limit this.
[0111] In some embodiments, the speech object recognition model can be used to classify the target facial feature vector to be recognized, and directly determine the speech state label of the target facial image to be recognized.
[0112] Of course, the method shown in steps S51 - S53 can also be adopted to determine the speech state label of the target facial image to be recognized.
[0113] In step S51, the speech object recognition model is used to extract features from the target facial image to be recognized, so as to obtain the target facial feature vector of the target facial image to be recognized. The speech object recognition model is trained using the facial image sample and the speech state label of the facial image sample.
[0114] In some implementations, the speech object recognition model can be used to extract features from the facial image sample to obtain the recognized facial feature vector; and the speech object recognition model is used to extract features from the target facial image to be recognized to obtain the target facial feature vector to be recognized.
[0115] In step S52, the target facial feature vector to be recognized is compared with the recognized facial feature vectors of the facial image samples to determine the similarity between the target facial feature vector to be recognized and each recognized facial feature vector.
[0116] In some embodiments, the target distance between the target facial feature vector to be recognized and each recognized facial image feature vector can be determined as the similarity between the target facial feature vector to be recognized and each recognized facial image feature vector; or the similarity between the target facial feature vector to be recognized and each recognized facial image feature vector can be directly calculated through the cosine algorithm, and the present disclosure does not limit this.
[0117] The above target distance may refer to Euclidean distance, absolute value distance, Chebyshev distance, Mahalanobis distance, Canberra distance, etc., and the present disclosure does not limit this.
[0118] In step S53, according to the similarity between the target face feature vector to be recognized and each recognized face feature vector, determine the speaking state label of the target face image to be recognized.
[0119] In some embodiments, each recognized face feature vector can be classified according to the speaking state label, and the similarity between the target face feature vector to be recognized and each recognized face feature vector in the speaking state, as well as the similarity between the target face feature vector to be recognized and each recognized face feature vector in the non-speaking state, are counted, and the speaking state label corresponding to the target face feature vector is determined according to the above two similarities.
[0120] For example, if the average value (or median, or maximum value, or minimum value, etc.) of the similarities between the target face feature vector to be recognized and each recognized face feature vector in the speaking state is greater than a certain threshold, it is determined that the target face image to be recognized is in the speaking state.
[0121] For example, if the average value (or median, or maximum value, or minimum value, etc.) of the similarities between the target face feature vector to be recognized and each recognized face feature vector in the non-speaking state is greater than a certain threshold, it is determined that the target face image to be recognized is in the non-speaking state.
[0122] It can be understood that this embodiment only takes the target face image to be recognized as an example to explain how to determine the speaking state label for the face image to be recognized, and those skilled in the art can determine the speaking state label for each face image to be recognized according to the technical solution provided in this embodiment.
[0123] The technical solution provided in this embodiment determines the similarity between the target face feature vector to be recognized and multiple recognized face feature vectors, and then determines the speaking state label of the target face image to be recognized according to multiple face image samples, improving the accuracy of speaking state recognition.
[0124] Figure 6 It is a flowchart of a method for determining the speaking state label of a target face image to be recognized according to similarity shown in an exemplary embodiment.
[0125] Refer to Figure 6 , the method for determining the speaking state label of the above target face image to be recognized may include the following process.
[0126] In step S61, determine the speaking state level of each face image sample.
[0127] In some embodiments, the speech state level of the facial image sample can be determined according to the recognition difficulty of the speech state of the facial image sample.
[0128] In some implementations, if a facial image sample is relatively easy to be recognized as a speech state, the speech state level of the facial image sample may be relatively high; if a facial image sample is relatively easy to be recognized as a non-speech state, the speech state level of the facial image sample may be relatively low; if a facial image sample is neither easy to be judged as a speech state nor easy to be judged as a non-speech state, the speech state level of the facial image sample may be in the middle.
[0129] For example, if the target object in the facial image sample wears a headset and has an open mouth, it is very likely that the target object is speaking, and the speech state level of the facial image sample can be relatively high, such as 10; if the target object in the facial image sample is obviously in a sleeping state, it is impossible for the target object in the facial image sample to be speaking, then the speech state level of the facial image sample can be relatively low, such as 0; if the target object in the facial image sample has an open mouth as if speaking (but not so certain), the speech state level of the facial image sample may be between 0 and 10, and those skilled in the art can set the speech state level for each image sample according to actual needs.
[0130] It can be understood that if a facial image sample is relatively easy to be recognized as a speech state, the speech state level of the facial image sample can also be set to be relatively low; if a facial image sample is relatively easy to be recognized as a non-speech state, the speech state level of the facial image sample can also be set to be relatively high; if a facial image sample is neither easy to be judged as a speech state nor easy to be judged as a non-speech state, the speech state level of the facial image sample can also be set to be in the middle, and the present disclosure does not limit this.
[0131] This embodiment is only explained by taking the above method for determining the speech state level as an example, but the present disclosure is not limited thereto.
[0132] In step S62, according to the speech state levels of the respective facial image samples, the similarity between the target facial feature vector to be recognized and each of the recognized facial feature vectors, the similarity distribution between the target facial feature vector to be recognized and each of the recognized facial feature vectors is determined.
[0133] In some implementations, the speech state levels of the respective recognized facial feature vectors can be used as one variable, and the similarity between the target facial feature vector to be recognized and each of the recognized facial feature vectors can be used as another variable, so as to determine the similarity distribution between the target facial feature vector to be recognized and each of the recognized facial feature vectors.
[0134] For example, the similarity distribution between the target face feature vector to be recognized and each recognized face feature vector can be determined with the speaking state level of each recognized face feature vector as the abscissa and the similarity between the target face feature vector to be recognized and each recognized face feature vector as the ordinate.
[0135] In step S63, according to the similarity distribution between the target face feature vector to be recognized and each recognized face feature vector, the speaking state label of the target face image to be recognized is determined.
[0136] In some implementations, if the speaking state level of the face image sample that is easily recognized as the speaking state level is higher than the speaking state level of the face image sample that is easily recognized as the non-speaking state level, the speaking state label of the target face image to be recognized can be determined according to the following method:
[0137] Determine that the similarity distribution between the target face feature vector to be recognized and each recognized face feature vector is a unimodal distribution (or approaches a unimodal distribution); obtain the distribution peak of the similarity distribution between the target face feature vector to be recognized and each recognized face feature vector; if the speaking state level of the face image sample corresponding to the distribution peak position is greater than the first threshold, determine that the speaking state label of the target face image corresponding to the target face feature vector to be recognized is the speaking state; if the speaking state level of the face image sample corresponding to the distribution peak position is lower than the first threshold, determine that the speaking state label of the target face image corresponding to the target face feature vector to be recognized is the non-speaking state.
[0138] In some embodiments, if the speaking state level of the face image sample that is easily recognized as the speaking state level is lower than the speaking state level of the face image sample that is easily recognized as the non-speaking state level, the speaking state label of the target face image to be recognized can be determined according to the following method:
[0139] Determine that the similarity distribution between the target face feature vector to be recognized and each recognized face feature vector is a unimodal distribution (or a regional unimodal distribution); obtain the distribution peak of the similarity distribution between the target face feature vector to be recognized and each recognized face feature vector; if the speaking state level of the face image sample corresponding to the distribution peak position is less than the fourth threshold, determine that the speaking state label of the target face image corresponding to the target face feature vector to be recognized is the speaking state; if the speaking state level of the face image sample corresponding to the distribution peak position is higher than the fifth threshold, determine that the speaking state label of the target face image corresponding to the target face feature vector to be recognized is the non-speaking state.
[0140] The technical solution provided in this embodiment performs a hierarchical processing on the speaking states of each facial image sample, then determines the similarity between the target facial feature vector to be recognized and each recognized facial feature vector, then determines the similarity distribution between the target facial feature vector to be recognized and each recognized facial feature vector, and finally determines the speaking state label of the target facial image to be recognized according to this similarity distribution. The technical solution provided in this embodiment determines the speaking state label of the target facial feature vector from the overall distribution effect through the similarity distribution of the target facial feature vector to be recognized relative to each recognized facial feature vector, and provides the accuracy of the speaking state label recognition.
[0141] In some embodiments, at least one video segment of the video to be recognized may include a target video segment.
[0142] Figure 7 It is a flowchart of a method for determining a speaking object in a target video segment shown according to an exemplary embodiment.
[0143] Reference Figure 7 , the above-mentioned speaking object determination method may include the following steps.
[0144] In step S71, object recognition processing is performed on each facial image to be recognized in each frame of the target video segment to determine the object category of each facial image to be recognized.
[0145] The object category of the above-mentioned facial image to be recognized may refer to the object to be recognized to which the facial information in the facial image to be recognized belongs. For example, if the facial information in the facial image to be recognized belongs to object 1 to be recognized, the object category of the facial image to be recognized is object 1 to be recognized; if the facial information in the facial image to be recognized belongs to object 2, the object category of the facial image to be recognized is object 2.
[0146] It can be understood that different facial images to be recognized may correspond to the same object category, that is, the same object category may appear in different facial images to be recognized.
[0147] In some embodiments, object recognition processing may be performed on each facial image to be recognized in each frame of the target video segment through pixel-level image detection technology (image processing technology that is not machine learning) to determine the object category of each facial image to be recognized; object recognition processing may also be performed on each facial image to be recognized in each frame of the target video segment through a trained neural network model (image processing technology for classification through machine learning) to determine the object category of each facial image to be recognized, and the present disclosure does not limit this.
[0148] In step S72, at least one face image to be recognized in at least two frames of images in the target video segment is clustered according to object categories, so as to obtain the face image clusters to be recognized in the target video segment under each object category.
[0149] In some embodiments, all face images to be recognized in the target video segment can be clustered according to object categories, so as to obtain the face image clusters to be recognized in the target video segment under each object category.
[0150] Among them, the images in each face image cluster to be recognized belong to the same object category.
[0151] In step S73, according to the face image clusters to be recognized in the target video segment under each object category and the speech state recognition results of at least one face image to be recognized in at least two frames of images in the target video segment, the face image cluster with the largest number of face images to be recognized whose speech state recognition result is the speaking state is determined as the target face image cluster to be recognized.
[0152] In some embodiments, the number of face images to be recognized in the speaking state and the number of face images to be recognized in the non-speaking state in each face image cluster to be recognized can be counted, and then the face image cluster with the largest number of face images to be recognized in the speaking state is determined as the target face image cluster to be recognized.
[0153] In step S74, the object category corresponding to the target face image cluster to be recognized is determined as the speaking object in the target video segment.
[0154] The technical solution provided in this embodiment clusters at least one face image to be recognized in the target video segment by object categories, and then determines the speaking object in the target video segment according to the clustering result, and determines the speaking object in the target video to be recognized from the overall distribution effect of the speech states of all face images to be recognized in the target video segment. The technical solution provided in this embodiment determines the speaking object in the target video segment through clustering and statistical principles, and improves the accuracy of determining the speaking object.
[0155] Figure 8 It is a training method of a speaking object recognition model shown according to an exemplary embodiment.
[0156] In addition, in the process of obtaining the training images of the speaking object recognition model, the following problems need to be noted: 1) How to accurately define the label set for speaking object recognition; 2) How to obtain a certain scale of labeled data.
[0157] When it comes to the problem of defining the label set, the technicians in this field first think of defining it based on the mouth shape, converting the problem of identifying the speaking object into the problem of "open mouth recognition". However, during the labeling process, it can be found that there is a certain difference between whether a person is speaking and whether he is opening his mouth. Figure 9 As shown, although the objects in the facial images corresponding to 900A, 900B, and 900C do not open their mouths, they are wearing headphones and have rich expressions, so they are very likely talking; on the contrary, although the objects in the facial images corresponding to 900C, 900D, and 900E have opened their mouths, it can be judged from the figures that they may just be smiling, not talking.
[0158] Therefore, when selecting training images, this embodiment not only retains the information of the mouth or facial feature points, but also retains the complete information of the entire face, thereby better retaining effective information.
[0159] Therefore, the training images provided in this embodiment are facial images.
[0160] In step S81, an initial training image and a speech state label of the initial training image are obtained.
[0161] In some embodiments, the speaking state label of the initial training image may be obtained by manual annotation.
[0162] For example, some frame images in videos including speaking objects may be manually annotated to obtain initial training images, such as annotating frame images in business videos in certain business scenarios (such as interview scenarios and speeches).
[0163] It should be noted that if the speaking status of frame images in a video including a speaking subject is manually labeled, the video needs to be frame-extracted before labeling, and then the frame images obtained after the frame-extraction process are labeled with the speaking status to generate the initial training image.
[0164] The initial training images obtained by the above-mentioned “frame extraction-labeling” method can be used to perform redundancy processing on similar frame images in the video to ensure the diversity of the initial training images finally obtained.
[0165] In addition, during the labeling process, we need to pay attention to one issue: remove visually inseparable data. That is, if it is impossible to tell whether a person is speaking from a single static image, then even if we judge from the video that the person in the image is indeed speaking in the current frame, such data should not be marked as a positive sample, because if it is fed to the machine for training, the machine will also find it difficult to distinguish.
[0166] In the classification problem, the visual separability of data is particularly important. Therefore, removing this type of data from the labeled data is beneficial to the solution and optimization of subsequent problems.
[0167] Figure 10 is based on Figure 8 The following is an entity block diagram corresponding to the training method of a speaker recognition model according to the illustrated embodiment.
[0168] Combined with Figure 10 , Figure 8 The training method of the illustrated speaker recognition model may include step S82-step S87.
[0169] In step S82, a target neural network is trained according to the initial training image and the speech state label of the initial training image to obtain an intermediate speaker recognition model.
[0170] In some embodiments, the target neural network can be any neural network model that can perform image recognition. For example, it can be a MobileNet-v1 (a neural network for face key point detection with deep cascading) classification model, or it can be a MobileNet-v2 (another neural network for face key point detection with deep cascading).
[0171] Among them, MobileNet-v2 adds a shortcut on the basis of MobileNet-v1, adopts the method of first expanding the channels and then compressing to avoid the problem of too few extractable features. In addition, at the end of the network, instead of using ReLU (Rectified Linear Unit), it uses Linear (a linear transformation function) to prevent ReLU from further destroying features in the channel compression stage.
[0172] It can be understood that the number of manually labeled initial training images 1001 is insufficient and it is difficult to reach the training scale.
[0173] Therefore, in response to the above problem 2) of how to obtain a certain scale of labeled data, this embodiment proposes the following method to achieve data supplementation for the initial training image 1001.
[0174] In step S83, the training images to be labeled are obtained.
[0175] The training images to be labeled 1003 are some facial images with unknown speech states.
[0176] To ensure the robustness of the model, the facial data of the same object should preferably appear only once in the training images to be labeled. However, it is difficult to meet this requirement in the data scale of the business scenario.
[0177] Therefore, different facial images can be crawled through the network, and the facial images corresponding to different facial information can be screened out as the training images to be labeled by using facial detection algorithms and deduplication algorithms.
[0178] In step S84, the intermediate model for speaker recognition is used to extract features from the training images to be labeled, so as to obtain the training vectors to be labeled.
[0179] In step S85, the training vectors to be labeled are compared with the recognized facial feature vectors of the facial image samples to determine the similarity between the training vectors to be labeled and each recognized facial feature vector.
[0180] In some embodiments, the intermediate model 1002 for speaker recognition can also be used to extract feature processing from the facial image samples to obtain the recognized facial feature vectors.
[0181] In some embodiments, the target distance between the training vectors to be labeled and each recognized facial image feature vector can be determined as the similarity between the training vectors to be labeled and each recognized facial image feature vector; alternatively, the similarity between the training vectors to be labeled and each recognized facial image feature vector can be directly calculated by the cosine algorithm, and the present disclosure does not limit this. The above target distance can refer to Euclidean distance, absolute value distance, Chebyshev distance, Mahalanobis distance, Canberra distance, etc., and the present disclosure does not limit this.
[0182] In step S86, according to the similarity between the training vectors to be labeled and each recognized facial feature vector, the speech state label of the training images to be labeled is determined.
[0183] In some embodiments, the recognized facial feature vectors can be classified according to the speech state label, and the similarity between the training vector to be labeled and each recognized facial feature vector in the speech state, as well as the similarity between the training vector to be labeled and each recognized facial feature vector in the non-speech state, can be counted, and the speech state label 1004 corresponding to the training vector to be labeled can be determined according to the above two similarities.
[0184] For example, if the average value (or median, or maximum value, or minimum value, etc.) of the similarities between the training vector to be labeled and each recognized facial feature vector in the speech state is greater than a certain threshold, it is determined that the training image to be labeled is in the speech state.
[0185] For example, if the average value (or median, or maximum value, or minimum value, etc.) of the similarities between the training vector to be labeled and each recognized facial feature vector in the non-speech state is greater than a certain threshold, it is determined that the training image to be labeled is in the non-speech state.
[0186] In some other embodiments, the speaking state label of the training image to be labeled can also be determined by the following method: determining the speaking state level of each facial image sample; determining the similarity distribution of the training vector to be labeled and each recognized facial feature vector according to the speaking state level of each facial image sample and the similarity between the training vector to be labeled and each recognized facial feature vector; and determining the speaking state label of the training image to be labeled according to the similarity distribution of the training vector to be labeled and each recognized facial feature vector.
[0187] Among them, the speaking state label of the training image to be labeled can be determined by the following method according to the similarity distribution of the training vector to be labeled and each recognized facial feature vector: determining that the similarity distribution of the training vector to be labeled and each recognized facial feature vector is a unimodal distribution; obtaining the distribution peak of the similarity distribution of the training vector to be labeled and each recognized facial feature vector; if the speaking state level of the facial image sample corresponding to the distribution peak position is greater than a certain set threshold, determining that the speaking state label of the training image to be labeled corresponding to the training vector to be labeled is the speaking state; if the speaking state level of the facial image sample corresponding to the distribution peak position is lower than another set threshold, determining that the speaking state label of the training image to be labeled corresponding to the training vector to be labeled is the non-speaking state.
[0188] In step S87, transfer training is performed on the intermediate model for speaker recognition according to the training image to be labeled and the speaking state label of the training image to be labeled, so as to obtain the speaker recognition model 1005.
[0189] The technical solution provided in this embodiment, on the one hand, trains the target neural network with the initial training images manually labeled, ensuring the accuracy of the training results; on the other hand, transfer training is performed on the intermediate model for speaker recognition by using the labels of the training images to be labeled classified and recognized by the intermediate model for speaker recognition and the training images to be labeled, further improving the intermediate model for speaker recognition, so as to obtain a target speaker recognition model with a relatively high accuracy of speaking state recognition. Among them, the intermediate model for speaker recognition is used to determine the speaking state label of the training image to be labeled, which can not only improve the accuracy of determining the speaking state label but also improve the speed of determining the speaking state label.
[0190] In some other embodiments, the speaking state label of the training image to be labeled can also be determined by the following method.
[0191] Figure 11 FIG. is a flowchart of a method for determining the speaking state label of a training image to be labeled according to the similarity between the training vector to be labeled and each recognized facial feature vector shown in an exemplary embodiment.
[0192] In step S111, based on the speech state labels of the facial image samples corresponding to the recognized facial feature vectors and the similarities between the recognized facial feature vectors and the training vector to be labeled, determine the initial confidence level of the training vector to be labeled relative to the recognized facial feature vectors.
[0193] The initial confidence level is a parameter that can be used to measure the overall correlation between the training vector to be labeled and all recognized facial feature vectors identified as being in a speech state.
[0194] In some embodiments, the target distance between the training vector to be labeled and each recognized facial feature vector can be determined as the similarity s between the training vector to be labeled and each recognized facial feature vector i ; alternatively, the similarity s between the training vector to be labeled and each recognized facial image feature vector can be directly calculated through the cosine algorithm i , and the present disclosure does not limit this. Here, i is an integer greater than or equal to 1 and less than or equal to the number of recognized facial feature vectors.
[0195] For example, it can be determined by S i = cos<f si , f b > the similarity between the training vector f to be labeled b and each recognized facial feature vector f si .
[0196] In some embodiments, the initial confidence level of the training vector to be labeled relative to the recognized facial feature vectors can be determined according to formula (1).
[0197]
[0198] Among them, ψ i represents the label value corresponding to the i-th recognized facial feature vector. If the label corresponding to the i-th recognized facial feature vector is the speech state, then this ψ i is 1. If the label corresponding to the i-th recognized facial feature vector is the non-speech state, then this ψ i is 0.
[0199] In step S112, based on the initial confidence level and the speech state labels of the recognized facial feature vectors, determine the target confidence level of the training vector to be labeled relative to the recognized facial feature vectors.
[0200] The target confidence level is a parameter that can be used to measure the credibility of the training vector to be labeled as being in a speech state.
[0201] In some embodiments, the target confidence level of the training vector to be labeled relative to the recognized facial feature vectors can be determined according to formula (2).
[0202]
[0203] Among them, ψ i represents the label value corresponding to the i-th recognized facial feature vector. If the label corresponding to the i-th recognized facial feature vector is the speaking state, then this ψ i is 1. If the label corresponding to the i-th recognized facial feature vector is the non-speaking state, then this ψ i is 0; s i represents the similarity between the training vector to be labeled and each recognized facial image feature vector, and i is an integer greater than or equal to 1 and less than or equal to the number n of recognized facial feature vectors.
[0204] In step S113, determine the speaking state label of the training image to be labeled according to the target confidence level.
[0205] In some embodiments, if the target confidence level is greater than the second threshold, determine that the speaking state label of the training object corresponding to the training vector to be labeled is the speaking state; if the target confidence level is less than the third threshold, determine that the speaking state label of the training object corresponding to the training vector to be labeled is the non-speaking state; if the target confidence level is greater than or equal to the third threshold and less than or equal to the second threshold, determine that the speaking state of the training object corresponding to the training vector to be labeled is unknown, so as to label the training image to be labeled with an unknown speaking state through manual recognition.
[0206] The technical solution improved in this embodiment determines the speaking state label of the training image to be labeled according to the target confidence level, improving the accuracy and speed of determining the speaking state label.
[0207] Figure 12 is a structural diagram of a video processing method shown according to an exemplary embodiment. Refer to Figure 12 , the above video processing method may include the following steps.
[0208] Step S121, perform video segmentation on the video 1201 to be recognized to obtain at least one video clip clipi 1202, 1203, or 1204, etc., where i is an integer greater than or equal to 1 and less than or equal to N, and N is the number of video clips; Step S122 (passport step S123, step S124), perform speaker recognition on each video clip to determine the speakers 1205, 1206, or 1207 in each video clip.
[0209] Among them, the above video clip may include the target video clip. Among them, step S124, performing speaker recognition on the target video clip may include the following steps:
[0210] Perform frame extraction on the target video segment to obtain at least two target frame images; perform face detection processing on each of the at least two target frame images to detect at least one face image to be recognized from the at least two target frame images; perform object recognition processing on the at least one face image to be recognized to determine the object category of each face image to be recognized; perform speech state recognition on the at least one face image to be recognized in the target video to obtain a speech state recognition result to determine the speech state of the at least one face image to be recognized; cluster the at least one face image to be recognized in at least two frames of the target video segment according to the object category to obtain the clustering of the face images to be recognized in the target video segment under each object category; according to the clustering of the face images to be recognized in the target video segment under each object category and the speech state recognition result of the at least one face image to be recognized in the at least two target frame images of the target video segment, determine the clustering of the face images to be recognized with the largest number of face images to be recognized whose speech state recognition result is the speech state as the target clustering of face images to be recognized; determine the object category corresponding to the target clustering of face images to be recognized as the speaking object in the target video segment.
[0211] Figure 13 It is a flowchart of a method for recognizing a speaking object in a video segment shown according to an exemplary embodiment.
[0212] Wherein, each video segment may include a target video segment, and performing speaking object recognition on the target video segment respectively may include the following steps:
[0213] Step S131, obtain the target video segment. Step S132, perform frame extraction on the target video to obtain at least two target frame images. Step S133, perform face detection processing on each of the at least two target frame images to detect at least one face image to be recognized from the at least two target frame images. In step S134, perform object recognition processing on the at least one face image to be recognized to determine the object category of each face image to be recognized. In step S135, calculate the confidence of each face image to be recognized. In step S136, determine the speech state label of the face image to be recognized according to the confidence of the face image to be recognized. In step S137, perform temporal analysis on each face image to be recognized in the target video segment to determine the speaking object of the target video segment.
[0214] Among them, calculating the confidence level for each face image to be recognized includes: extracting features from at least one face image to be recognized in the target video to obtain at least one face feature vector to be recognized; determining the initial confidence level of at least one face feature vector to be recognized relative to the recognized face feature vectors according to the speaking state labels of the face image samples corresponding to the recognized face feature vectors and the similarity between the recognized face feature vectors and at least one face feature vector to be recognized; determining the target confidence level of at least one face feature vector to be recognized relative to the recognized face feature vectors according to the initial confidence level and the speaking state labels of the recognized face feature vectors; and determining the speaking state label of at least one face image to be recognized according to the target confidence level.
[0215] Among them, performing temporal analysis on each face image to be recognized in the target video segment to determine the speaking object of the target video segment includes: clustering at least one face image to be recognized in at least two target frame images in the target video segment according to the object category to obtain the clustering of face images to be recognized in the target video segment under each object category; determining, according to the clustering of face images to be recognized in the target video segment under each object category and the speaking state recognition results of at least one face image to be recognized in at least two frames of images in the target video segment, the clustering of face images to be recognized with the largest number of face images to be recognized whose speaking state recognition results are the speaking state as the target clustering of face images to be recognized; and determining the object category corresponding to the target clustering of face images to be recognized as the speaking object in the target video segment.
[0216] Figure 14 is a block diagram of a video processing apparatus shown according to an exemplary embodiment. Referring to Figure 14 , the video processing apparatus 1400 provided in the embodiments of the present disclosure may include: a video segmentation module 1401, a frame extraction module 1402, a face recognition module 1403, a speaking state recognition module 1404, and a speaking object determination module 1405.
[0217] Among them, the video segmentation module 1401 can be used to segment the video to be recognized by using the audio information of the video to be recognized, so as to obtain at least one video segment of the video to be recognized, and each video segment includes at least one speaking object. The frame extraction module 1402 can be used to extract at least two frames of images from each video segment. The face recognition module 1403 can be used to perform face recognition processing on each frame of image in each video segment, so as to obtain at least one face image to be recognized from each frame of image in each video segment. The speaking state recognition module 1404 can be used to perform speaking state recognition on each face image to be recognized in each frame of image in each video segment respectively, so as to determine the speaking state recognition result of each face image to be recognized. The speaking object determination module 1405 can be used to determine the speaking object in each video segment according to the speaking state recognition results of at least one face image to be recognized in at least two frames of images in each video segment.
[0218] In some embodiments, at least one video segment of the video to be recognized includes a target video segment. Among them, the speaking object determination module 1405 can include: an object recognition processing sub-module, a clustering sub-module, a target face image to be recognized clustering determination sub-module, and a speaking object determination sub-module.
[0219] Among them, the object recognition processing sub-module can be used to perform object recognition processing on each face image to be recognized in each frame of image in the target video segment, and determine the object category of each face image to be recognized. The clustering sub-module can be used to cluster at least one face image to be recognized in at least two frames of images in the target video segment according to the object category, so as to obtain the clustering of face images to be recognized in the target video segment under each object category. The target face image to be recognized clustering determination sub-module can be used to determine, according to the clustering of face images to be recognized in the target video segment under each object category and the speaking state recognition results of at least one face image to be recognized in at least two frames of images in the target video segment, the clustering of face images to be recognized with the largest number of face images to be recognized whose speaking state recognition results are the speaking state as the target clustering of face images to be recognized. The speaking object determination sub-module can be used to determine the object category corresponding to the target clustering of face images to be recognized as the speaking object in the target video segment.
[0220] In some embodiments, at least one video segment of the video to be recognized includes a target video segment, at least two frames of images of the target video segment include target frame images, and at least one face image to be recognized of the target frame images includes a target face image to be recognized. Among them, the speaking state recognition module 1404 can include: a feature extraction sub-module, a feature comparison sub-module, and a speaking state label determination sub-module.
[0221] Among them, the feature extraction sub-module can be used to extract features from the target facial image to be recognized through a speaker recognition model, so as to obtain the target facial feature vector of the target facial image to be recognized. The speaker recognition model is trained using facial image samples and the speech state labels of the facial image samples. The feature comparison sub-module can be used to compare the target facial feature vector to be recognized with the recognized facial feature vectors of the facial image samples to determine the similarity between the target facial feature vector to be recognized and each recognized facial feature vector. The speech state label determination sub-module can be used to determine the speech state label of the target facial image to be recognized according to the similarity between the target facial feature vector to be recognized and each recognized facial feature vector.
[0222] In some embodiments, the speech state label determination sub-module may include: a speech state level determination unit, a similarity distribution determination unit, and a speech state label determination unit.
[0223] Among them, the speech state level determination unit can be used to determine the speech state level of each facial image sample. The similarity distribution determination unit can be used to determine the similarity distribution of the target facial feature vector to be recognized and each recognized facial feature vector according to the speech state level of each facial image sample and the similarity between the target facial feature vector to be recognized and each recognized facial feature vector. The speech state label determination unit can be used to determine the speech state label of the target facial image to be recognized according to the similarity distribution of the target facial feature vector to be recognized and each recognized facial feature vector.
[0224] In some embodiments, the speech state label determination unit includes: a unimodal distribution determination sub-unit, a distribution peak determination sub-unit, a speech state determination sub-unit, and a non-speech state determination sub-unit.
[0225] Among them, the unimodal distribution determination sub-unit can be used to determine that the similarity distribution of the target facial feature vector to be recognized and each recognized facial feature vector is a unimodal distribution. The distribution peak determination sub-unit can be used to obtain the distribution peak of the similarity distribution of the target facial feature vector to be recognized and each recognized facial feature vector. The speech state determination sub-unit can be used to determine that the speech state label of the target facial image corresponding to the target facial feature vector is the speech state if the speech state level of the facial image sample corresponding to the distribution peak position is greater than the first threshold. The non-speech state determination sub-unit can be used to determine that the speech state label of the target facial image corresponding to the target facial feature vector is the non-speech state if the speech state level of the facial image sample corresponding to the distribution peak position is lower than the first threshold.
[0226] In some embodiments, the speech state recognition module 1404 may further include: an initial training image acquisition sub-module, a speech object recognition intermediate model training sub-module, a to-be-annotated training image acquisition sub-module, a to-be-annotated training vector acquisition sub-module, a similarity determination sub-module, a speech state label determination sub-module, and a transfer training sub-module.
[0227] Among them, the initial training image acquisition sub-module may be configured to acquire initial training images and speech state labels of the initial training images. The speech object recognition intermediate model training sub-module may be configured to train a target neural network according to the initial training images and speech state labels of the initial training images to obtain a speech object recognition intermediate model. The to-be-annotated training image acquisition sub-module may be configured to acquire to-be-annotated training images. The to-be-annotated training vector acquisition sub-module may be configured to perform feature extraction on the to-be-annotated training images through the speech object recognition intermediate model to obtain to-be-annotated training vectors. The similarity determination sub-module may be configured to perform feature comparison between the to-be-annotated training vectors and the recognized facial feature vectors of the facial image samples to determine the similarity between the to-be-annotated training vectors and each recognized facial feature vector. The speech state label determination sub-module may be configured to determine the speech state label of the to-be-annotated training image according to the similarity between the to-be-annotated training vector and each recognized facial feature vector. The transfer training sub-module may be configured to perform transfer training on the speech object recognition intermediate model according to the to-be-annotated training images and speech state labels of the to-be-annotated training images to obtain a speech object recognition model.
[0228] In some embodiments, the speech state label determination sub-module may include: a speech state level determination unit for facial image samples, a similarity distribution determination unit for recognized facial feature vectors, and a speech state label determination unit for to-be-annotated training images.
[0229] Among them, the speech state level determination unit for facial image samples determines the speech state levels of each facial image sample. The similarity distribution determination unit for recognized facial feature vectors may be configured to determine the similarity distribution between the to-be-annotated training vector and each recognized facial feature vector according to the speech state levels of each facial image sample and the similarity between the to-be-annotated training vector and each recognized facial feature vector. The speech state label determination unit for to-be-annotated training images may be configured to determine the speech state label of the to-be-annotated training image according to the similarity distribution between the to-be-annotated training vector and each recognized facial feature vector.
[0230] In some embodiments, the speech state label determination sub-module may include: an initial confidence determination unit, a target confidence determination unit, and a speech state label determination unit for to-be-annotated training images.
[0231] Among them, the initial confidence determination unit can be used to determine the initial confidence of the training vector to be labeled relative to the recognized facial feature vectors according to the speech state labels of the facial image samples corresponding to the recognized facial feature vectors and the similarities between the recognized facial feature vectors and the training vector to be labeled. The target confidence determination unit can be used to determine the target confidence of the training vector to be labeled relative to the recognized facial feature vectors according to the initial confidence and the speech state labels of the recognized facial feature vectors. The speech state label determination unit of the training image to be labeled can be used to determine the speech state label of the training image to be labeled according to the target confidence.
[0232] In some embodiments, the speech state label determination unit of the training image to be labeled may include: a first threshold judgment subunit, a second threshold judgment subunit, and a third threshold judgment subunit.
[0233] Among them, the first threshold judgment subunit can be used to determine that the speech state label of the training object to be labeled corresponding to the training vector to be labeled is the speech state if the target confidence is greater than the second threshold. The second threshold judgment subunit can be used to determine that the speech state label of the training object to be labeled corresponding to the training vector to be labeled is the non-speech state if the target confidence is less than the third threshold. The third threshold judgment subunit can be used to determine that the speech state of the training object to be labeled corresponding to the training vector to be labeled is unknown if the target confidence is greater than or equal to the third threshold and less than or equal to the second threshold, so as to label the training image to be labeled with an unknown speech state through manual recognition.
[0234] Since the functions of device 1400 have been described in detail in their corresponding method embodiments, the present disclosure will not be elaborated herein.
[0235] Through the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solution of the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.), including several instructions for causing a computing device (which can be a personal computer, a server, a mobile terminal, or a smart device, etc.) to execute the method according to the embodiments of the present disclosure, such as Figure 5 one or more of the steps shown.
[0236] In addition, the above drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present disclosure, rather than for limiting purposes. It is easy to understand that the processes shown in the above drawings do not indicate or limit the time sequence of these processes. Additionally, it is also easy to understand that these processes can be executed synchronously or asynchronously in, for example, multiple modules.
[0237] Other embodiments of the present disclosure will be readily apparent to those skilled in the art in view of the specification and practice of the disclosure herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field of the present disclosure that are not claimed in the present disclosure. The specification and examples are to be considered exemplary only, and the true scope and spirit of the present disclosure are pointed out by the claims.
[0238] It should be understood that the present disclosure is not limited to the detailed structures, drawing methods, or implementation methods shown herein. On the contrary, the present disclosure is intended to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims.
Claims
1. A video processing method, characterized in that, Including: Segmenting the video to be recognized using the audio information of the video to be recognized to obtain at least one video segment of the video to be recognized, including: segmenting the video to be recognized into multiple video segments using the timbre of the audio information, each video segment including a speaking object with the same speaking timbre; or segmenting the video to be recognized into multiple video segments using the speech content in the audio information, and the speaking objects in each video segment speak around the same theme; or classifying the audio information using speech recognition technology to determine different audio segments corresponding to different speaking objects, and segmenting the video to be recognized using different audio segments so that each video segment includes one speaking object; Extracting at least two frames of images from each video segment; Performing face recognition processing on each frame of image in each video segment to obtain at least one face image to be recognized from each frame of image in each video segment; Performing speaking state recognition on each face image to be recognized in each frame of image in each video segment to determine the speaking state recognition result of each face image to be recognized; Determining the speaking object in each video segment according to the speaking state recognition result of the at least one face image to be recognized in the at least two frames of images in each video segment.
2. The method according to claim 1, wherein At least one video segment of the video to be recognized includes a target video segment; wherein, determining the speaking object in each video segment according to the speaking state recognition result of the at least one face image to be recognized in the at least two frames of images in each video segment includes: Performing object recognition processing on each face image to be recognized in each frame of image in the target video segment to determine the object category of each face image to be recognized; Clustering the at least one face image to be recognized in the at least two frames of images in the target video segment according to the object category to obtain the clustering of face images to be recognized in the target video segment under each object category; Determining the clustering of face images to be recognized with the largest number of face images to be recognized whose speaking state recognition result is the speaking state as the target clustering of face images to be recognized according to the clustering of face images to be recognized in the target video segment under each object category and the speaking state recognition result of the at least one face image to be recognized in the at least two frames of images in the target video segment; Determining the object category corresponding to the target clustering of face images to be recognized as the speaking object in the target video segment.
3. The method according to claim 1, characterized in that, At least one video segment of the video to be recognized includes a target video segment, at least two frames of images of the target video segment include target frame images, and at least one face image to be recognized of the target frame image includes a target face image to be recognized; wherein, performing speaking state recognition on each face image to be recognized in each frame of image in each video segment to determine the speaking state recognition result of each face image to be recognized includes: Extract features from the target face image to be recognized through a speaker recognition model to obtain a target face feature vector of the target face image to be recognized, where the speaker recognition model is trained using face image samples and the speech state labels of the face image samples; Compare the target face feature vector to be recognized with the recognized face feature vectors of the face image samples to determine the similarity between the target face feature vector to be recognized and each recognized face feature vector; Determine the speech state label of the target face image to be recognized according to the similarity between the target face feature vector to be recognized and each recognized face feature vector.
4. The method according to claim 3, characterized in that, Determine the speech state label of the target face image to be recognized according to the similarity between the target face feature vector to be recognized and each recognized face feature vector, including: Determine the speech state level of each face image sample; Determine the similarity distribution between the target face feature vector to be recognized and each recognized face feature vector according to the speech state level of each face image sample and the similarity between the target face feature vector to be recognized and each recognized face feature vector; Determine the speech state label of the target face image to be recognized according to the similarity distribution between the target face feature vector to be recognized and each recognized face feature vector.
5. The method according to claim 4, characterized in that, Determine the speech state label of the target face image to be recognized according to the similarity distribution between the target face feature vector to be recognized and each recognized face feature vector, including: Determine that the similarity distribution between the target face feature vector to be recognized and each recognized face feature vector is a unimodal distribution; Obtain the distribution peak of the similarity distribution between the target face feature vector to be recognized and each recognized face feature vector; If the speech state level of the face image sample corresponding to the distribution peak position is greater than the first threshold, determine that the speech state label of the target face image corresponding to the target face feature vector to be recognized is the speech state; If the speech state level of the face image sample corresponding to the distribution peak position is lower than the first threshold, determine that the speech state label of the target face image corresponding to the target face feature vector to be recognized is the non-speech state.
6. The method according to claim 3, characterized in that, Before extracting features from the target face image to be recognized through a speaker recognition model, it includes: Obtain an initial training image and the speech state label of the initial training image; Train a target neural network according to the initial training image and the speech state label of the initial training image to obtain an intermediate speaker recognition model; Obtain a training image to be labeled; Extract features from the training image to be labeled through the intermediate speaker recognition model to obtain a training vector to be labeled; Compare the training vector to be labeled with the recognized face feature vectors of the face image samples to determine the similarity between the training vector to be labeled and each recognized face feature vector; Determine the speech state label of the training image to be labeled according to the similarity between the training vector to be labeled and each recognized face feature vector. Transfer training is performed on the intermediate model for speaker recognition according to the training image to be labeled and the speech state label of the training image to be labeled, so as to obtain the speaker recognition model.
7. The method according to claim 6, wherein Determining the speech state label of the training image to be labeled according to the similarity between the training vector to be labeled and each recognized facial feature vector includes: Determining the speech state level of each facial image sample; Determining the similarity distribution between the training vector to be labeled and each recognized facial feature vector according to the speech state level of each facial image sample and the similarity between the training vector to be labeled and each recognized facial feature vector; Determining the speech state label of the training image to be labeled according to the similarity distribution between the training vector to be labeled and each recognized facial feature vector.
8. The method according to claim 6, wherein Determining the speech state label of the training image to be labeled according to the similarity between the training vector to be labeled and each recognized facial feature vector includes: Determining the initial confidence of the training vector to be labeled relative to the recognized facial feature vector according to the speech state label of the facial image sample corresponding to each recognized facial feature vector and the similarity between each recognized facial feature vector and the training vector to be labeled; Determining the target confidence of the training vector to be labeled relative to the recognized facial feature vector according to the initial confidence and the speech state label of each recognized facial feature vector; Determining the speech state label of the training image to be labeled according to the target confidence.
9. The method according to claim 8, characterized in that, Determining the speech state label of the training image to be labeled according to the target confidence includes: If the target confidence is greater than a second threshold, determining that the speech state label of the training object corresponding to the training vector to be labeled is the speech state; If the target confidence is less than a third threshold, determining that the speech state label of the training object corresponding to the training vector to be labeled is the non-speech state; If the target confidence is greater than or equal to the third threshold and less than or equal to the second threshold, determining that the speech state of the training object corresponding to the training vector to be labeled is unknown, so as to label the training image to be labeled with an unknown speech state by means of manual recognition.
10. A video processing apparatus, characterized in that, Including: A video segmentation module, configured to segment the video to be recognized by using the audio information of the video to be recognized, so as to obtain at least one video segment of the video to be recognized, including: segmenting the video to be recognized into multiple video segments by using the timbre of the audio information, where each video segment includes a speaker with the same speaking timbre; or segmenting the video to be recognized into multiple video segments by using the speech content in the audio information, where the speakers in each video segment speak around the same theme; or classifying the audio information by using speech recognition technology to determine different audio segments corresponding to different speakers, and performing segmentation processing on the video to be recognized by using different audio segments, so that each video segment includes one speaker; A frame extraction module, configured to extract at least two frames of images from each video segment; A face recognition module, configured to perform face recognition processing on each frame image in each video segment to obtain at least one face image to be recognized in each frame image of each video segment; A speaking state recognition module, configured to perform speaking state recognition on each face image to be recognized in each frame image of each video segment to determine the speaking state recognition result of each face image to be recognized; A speaking object determination module, configured to determine the speaking object in each video segment according to the speaking state recognition results of the at least one face image to be recognized in the at least two frame images of each video segment.
11. An electronic device, characterized in that, Comprising: A memory; And A processor coupled to the memory, the processor being configured to execute the video processing method according to any one of claims 1-9 based on instructions stored in the memory.
12. A computer-readable storage medium, having a program stored thereon, which when executed by a processor implements the video processing method according to any one of claims 1-9.
13. A computer program product, which includes computer instructions, and when the computer instructions are executed, the video processing method according to any one of claims 1-9 is implemented.
Citation Information
Patent Citations
No-supervision multi-speaker identification device based on audio and video and method thereof
CN109410954A
Robot, control method thereof and computer readable storage medium
CN111241922A