Audio and video processing method and device, electronic equipment and computer readable medium

By performing face segmentation and feature extraction on surveillance videos, and combining image and text encoder models, the problem of insufficient efficiency and accuracy in surveillance video detection in existing technologies has been solved, achieving efficient user identification and audio optimization.

CN119992455BActive Publication Date: 2026-01-23HAINAN RES INST OF ZHEJIANG UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510080884.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2026-01-23
Estimated Expiration
2045-01-20

AI Technical Summary

Technical Problem

In existing technologies, manual inspection of surveillance videos is inefficient, and conventional video detection algorithms are sensitive to noise and have poor separation of complex backgrounds, resulting in poor accuracy in video image detection.

Method used

By extracting video frames from surveillance videos, performing face segmentation and text labeling, and using a pre-trained image segmentation and classification model combined with image and text encoders to extract image features, generating segmentation quality classification information, identifying target faces, optimizing audio processing, and sending the data to the monitoring and management terminal.

Benefits of technology

It improves the detection efficiency of surveillance videos and the accuracy of image segmentation, enabling comprehensive detection of the monitored area and enhancing user identification and audio optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992455B_ABST
    Figure CN119992455B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose an audio and video processing method, device, electronic equipment and computer readable medium. A specific implementation of the method comprises: performing face segmentation on the face in each monitoring video image to generate a face segmentation image; performing image preprocessing on the face segmentation image; inputting the preprocessed face segmentation image into an image feature extraction network included in an image segmentation classification model; inputting the preprocessed face segmentation image into an image encoder, and inputting text label information into a text encoder; inputting the first face image feature information and the text feature information into an output layer; performing user recognition on the generated target face segmentation image sequence; performing speech optimization processing on the audio corresponding to the monitoring video; and sending the user recognition result sequence and the optimized audio to a monitoring management terminal. The implementation improves the detection efficiency of the monitoring video, and improves the verification efficiency and accuracy of image segmentation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments disclosed herein relate to the field of computers, and more particularly to audio and video processing methods, apparatuses, electronic devices, and computer-readable media. Background Technology

[0002] Current defense platforms typically focus on protecting important facilities and areas, such as secure bases or critical infrastructure. Currently, the processing of surveillance video in important areas usually involves either real-time manual inspection or detection and identification using conventional video detection algorithms.

[0003] However, the above methods usually have the following technical problems: manual inspection of surveillance videos is inefficient, conventional video detection algorithms are sensitive to noise, have poor separation performance in complex backgrounds, and cannot handle blurred edges, resulting in poor accuracy of video image detection.

[0004] The information disclosed in this background section is only intended to enhance the understanding of the background of the inventive concept, and therefore may contain information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0005] The summary portion of this disclosure is intended to provide a brief overview of the concepts, which will be described in detail in the detailed description portion. This summary portion is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.

[0006] Some embodiments of this disclosure provide audio and video processing methods, apparatuses, electronic devices, and computer-readable media to solve one or more of the technical problems mentioned in the background section above.

[0007] In a first aspect, some embodiments of this disclosure provide an audio and video processing method, which includes: extracting video frames from a surveillance video to obtain a surveillance video image sequence; performing face segmentation on faces in each surveillance video image in the surveillance video image sequence to generate face segmentation images, obtaining a face segmentation image sequence; and adding text tags to each face segmentation image in the face segmentation image sequence; for each face segmentation image in the face segmentation image sequence, performing the following processing steps: performing image preprocessing on the face segmentation images to obtain preprocessed face segmentation images, wherein the face segmentation images correspond to text tag information; inputting the preprocessed face segmentation images into an image feature extraction network included in a pre-trained image segmentation and classification model to obtain first face image feature information, wherein the image segmentation and classification model includes image feature extraction network. The system comprises a feature extraction network, a semantic feature extraction network, and an output layer. The semantic feature extraction network includes a text encoder and an image encoder. The preprocessed face segmentation image is input into the image encoder, and the text tag information is input into the text encoder to obtain text feature information. The first face image feature information and the text feature information are input into the output layer to obtain segmentation quality classification information corresponding to the face segmentation image. In response to determining that the segmentation quality classification information meets preset segmentation conditions, the face segmentation image is identified as the target face segmentation image. User recognition is performed on the generated target face segmentation image sequence to obtain a user recognition result sequence. Speech optimization processing is performed on the audio corresponding to the surveillance video to obtain optimized audio. The user recognition result sequence and the optimized audio are sent to the associated monitoring management terminal.

[0008] Secondly, some embodiments of this disclosure provide an audio and video processing apparatus, comprising: an extraction unit configured to extract video frames from a surveillance video to obtain a surveillance video image sequence; a segmentation unit configured to perform face segmentation on faces in each surveillance video image in the surveillance video image sequence to generate face segmentation images, obtaining a face segmentation image sequence, and to perform text tagging on each face segmentation image in the face segmentation image sequence; and a segmentation detection unit configured to perform the following processing steps on each face segmentation image in the face segmentation image sequence: performing image preprocessing on the face segmentation images to obtain preprocessed face segmentation images, wherein the face segmentation images correspond to text tagging information; and inputting the preprocessed face segmentation images into an image feature extraction network included in a pre-trained image segmentation classification model to obtain first face image feature information, wherein the image segmentation classification model includes image feature extraction network. The system comprises a feature extraction network, a semantic feature extraction network, and an output layer. The semantic feature extraction network includes a text encoder and an image encoder. The preprocessed face segmentation image is input into the image encoder, and the text tagging information is input into the text encoder to obtain text feature information. The first face image feature information and the text feature information are input into the output layer to obtain segmentation quality classification information corresponding to the face segmentation image. In response to determining that the segmentation quality classification information meets preset segmentation conditions, the face segmentation image is identified as a target face segmentation image. A recognition unit is configured to perform user recognition on the generated target face segmentation image sequence to obtain a user recognition result sequence. An optimization unit is configured to perform voice optimization processing on the audio corresponding to the monitoring video to obtain optimized audio. A sending unit is configured to send the user recognition result sequence and the optimized audio to an associated monitoring management terminal.

[0009] Thirdly, some embodiments of this disclosure provide an electronic device, including: one or more processors; and a storage device having one or more programs stored thereon, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation of the first aspect above.

[0010] Fourthly, some embodiments of this disclosure provide a computer-readable medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the method described in any of the implementations of the first aspect above.

[0011] Manual inspection of surveillance videos is inefficient. Conventional video detection algorithms are sensitive to noise, have poor separation performance in complex backgrounds, and cannot handle blurred edges, resulting in poor accuracy in video image detection.

[0012] The above-described embodiments of this disclosure have the following beneficial effects: The audio and video processing methods of some embodiments of this disclosure improve the detection efficiency of surveillance videos and enhance the verification efficiency and accuracy of image segmentation. Furthermore, they can simultaneously detect audio in the acquired target area, facilitating comprehensive detection of the situation within the monitored area. First, video frames are extracted from the surveillance video to obtain a sequence of surveillance video images. This facilitates frame-by-frame analysis. Second, face segmentation is performed on the faces in each surveillance video image in the above-described surveillance video image sequence to generate face segmentation images, resulting in a sequence of face segmentation images. Text labels are then applied to each face segmentation image in the above-described face segmentation image sequence. This facilitates preliminary differentiation of each face image. Next, for each face segmentation image in the aforementioned face segmentation image sequence, the following processing steps are performed: Image preprocessing is performed on the face segmentation image to obtain a preprocessed face segmentation image, wherein the face segmentation image corresponds to text tagging information; the preprocessed face segmentation image is input into the image feature extraction network of a pre-trained image segmentation classification model to obtain first face image feature information, wherein the image segmentation classification model includes an image feature extraction network, a semantic feature extraction network, and an output layer, and the semantic feature extraction network includes a text encoder and an image encoder; the preprocessed face segmentation image is input into the image encoder, and the text tagging information is input into the text encoder to obtain text feature information; the first face image feature information and the text feature information are input into the output layer to obtain segmentation quality classification information corresponding to the face segmentation image; in response to determining that the segmentation quality classification information meets preset segmentation conditions, the face segmentation image is identified as the target face segmentation image. Thus, by combining the image dimension-related features and semantic dimension-related features of the face image, the segmentation quality classification information of the segmented image can be determined, thereby achieving the detection of the image segmentation result. Because the segmentation quality category information is determined by combining features from both image and text dimensions, extracting text features enhances the completeness of the extracted features and improves the accuracy of image segmentation detection. Then, user identification is performed on the generated target face segmentation image sequence to obtain a user identification result sequence. This allows for the identification of each high-quality face image to determine if it belongs to an unknown user. Finally, the audio corresponding to the aforementioned surveillance video undergoes speech optimization processing to obtain optimized audio; the user identification result sequence and the optimized audio are then sent to the associated monitoring management terminal. This improves the detection efficiency of surveillance video and enhances the verification efficiency and accuracy of image segmentation. Furthermore, audio from the collected target area can be detected simultaneously for comprehensive monitoring of the area. Attached Figure Description

[0013] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.

[0014] Figure 1 This is a flowchart of some embodiments of the audio and video processing method according to the present disclosure;

[0015] Figure 2 These are schematic diagrams illustrating the structure of some embodiments of the audio and video processing apparatus according to this disclosure;

[0016] Figure 3 This is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. Detailed Implementation

[0017] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0018] It should also be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. Unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.

[0019] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0020] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0021] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0022] This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.

[0023] Figure 1This is a flowchart of some embodiments of the audio and video processing method according to the present disclosure. A flowchart 100 of some embodiments of the audio and video processing method according to the present disclosure is shown. The audio and video processing method includes the following steps:

[0024] Step 101: Extract video frames from the surveillance video to obtain a sequence of surveillance video images.

[0025] In some embodiments, the entity executing the audio / video processing method (e.g., a computing device) can extract video frames from the surveillance video to obtain a sequence of surveillance video images. The surveillance video can refer to surveillance video of a target area. The target area can refer to an area requiring security monitoring. For example, the target area could be a research institute, a bank, or a restricted area. The surveillance video can refer to video covering images of users' faces within the target area.

[0026] Step 102: Perform face segmentation on the faces in each surveillance video image in the above surveillance video image sequence to generate face segmentation images, obtain a face segmentation image sequence, and add text labels to each face segmentation image in the above face segmentation image sequence.

[0027] In some embodiments, the execution entity may perform face segmentation on the faces in each monitoring video image in the monitoring video image sequence to generate face segmentation images, obtain a face segmentation image sequence, and perform text labeling on each face segmentation image in the face segmentation image sequence.

[0028] For example, firstly, surveillance video images containing complete faces can be selected from the aforementioned sequence of surveillance video images as candidate surveillance video images, resulting in a candidate surveillance video image sequence. Secondly, face region segmentation is performed on the faces in each candidate surveillance video image in the candidate sequence to generate face segmentation images, resulting in a face segmentation image sequence. Then, text labeling is applied to each face segmentation image in the aforementioned face segmentation image sequence. Here, text labeling can be simple labeling information such as the user's gender, facial features, and clothing features. It should be noted that this text labeling can be either manual or automatic.

[0029] Step 103: For each face segmentation image in the above face segmentation image sequence, perform the following processing steps:

[0030] Step 1031: Perform image preprocessing on the face segmentation image to obtain a preprocessed face segmentation image.

[0031] In some embodiments, the execution entity can perform image preprocessing on the face segmentation image to obtain a preprocessed face segmentation image. The face segmentation image includes corresponding text tagging information. For example, the face segmentation image can be resized to obtain the preprocessed face segmentation image. For instance, the size of the face segmentation image can be adjusted to a preset image size to obtain the preprocessed face segmentation image. As another example, the face segmentation image can be vectorized to obtain the preprocessed face segmentation image.

[0032] Step 1032: Input the preprocessed face segmentation image above into the image feature extraction network included in the pre-trained image segmentation and classification model to obtain the first face image feature information.

[0033] In some embodiments, the execution entity can input the preprocessed face segmentation image into the image feature extraction network of a pre-trained image segmentation and classification model to obtain first face image feature information. The image segmentation and classification model includes an image feature extraction network, a semantic feature extraction network, and an output layer. The semantic feature extraction network includes a text encoder and an image encoder. The image segmentation and classification model can be a pre-trained neural network model that takes the preprocessed face segmentation image and text labeling information as input and outputs the segmentation quality classification information corresponding to the preprocessed face segmentation image. The image feature extraction network can be used for edge feature extraction and may include a convolutional neural network or a residual network. The image feature extraction network may also include an edge detection algorithm. The edge detection algorithm may include the Sobel operator, Prewitt operator, Canny operator, and Canny edge detection. The semantic feature extraction network can be used to extract semantic features. The semantic feature extraction network may include a text encoder and an image encoder. The text encoder can be used to extract text features from the input text. The image encoder can be used to extract image features from the input image. The text encoder may include a Transformer block or BERT. The image encoder may include a VisionTransformer (ViT). The semantic feature extraction network described above can perform feature alignment on the extracted text features and image features. For example, the execution entity can input the preprocessed face segmentation image into the convolutional layers of the image feature extraction network to obtain first face image feature information. As another example, edge detection processing can be performed on the preprocessed face segmentation image to obtain first face image feature information. The first image feature information may include: corner feature information, texture feature information, and shape feature information.

[0034] Step 1033: Input the preprocessed face segmentation image into the image encoder and input the text tagging information into the text encoder to obtain text feature information.

[0035] In some embodiments, the execution entity can input the preprocessed face segmentation image into the image encoder and the text tagging information into the text encoder to obtain text feature information. For example, when the image encoder includes multi-layer ViT blocks and the text encoder includes multi-layer Transformer blocks, the execution entity can divide the preprocessed face segmentation image into a preset number of image blocks, and flatten each image block and linearly project it onto a preset dimension to obtain each image feature vector corresponding to the preprocessed face segmentation image. Then, each image feature vector can be input into the multi-layer ViT block to obtain an image representation. Next, the execution entity can perform word segmentation processing on the text tagging information to obtain a word segmentation sequence. Then, each word in the word segmentation sequence can be mapped to an embedding vector, and a learnable position embedding can be added to obtain each text feature vector corresponding to the text tagging information. Then, each text feature vector can be input into the multi-layer Transformer block to obtain a text representation. Finally, feature alignment processing can be performed on the image representation and text representation to obtain text feature information corresponding to the preprocessed face segmentation image. The image representation output by the image encoder can be a global representation obtained using the [CLS] token or average pooling. The text representation output by the text encoder can be a global text representation using the [CLS] token.

[0036] In practice, the aforementioned execution entity can input the preprocessed face segmentation image into the image encoder and the text tagging information into the text encoder through the following steps:

[0037] The first step is to input the preprocessed face segmentation image into the image encoder to obtain the second face image feature information.

[0038] The second step is to input the above text tag information into the above text encoder to obtain text tag feature information.

[0039] The third step is to perform feature alignment processing on the aforementioned second face image feature information and the aforementioned text marker feature information to obtain text feature information.

[0040] The third step mentioned above may include the following sub-steps:

[0041] The first sub-step involves normalizing the aforementioned second face image feature information to obtain normalized image feature information. For example, the aforementioned second face image feature information can be normalized using the L2 norm to obtain normalized image feature information.

[0042] The second sub-step involves normalizing the aforementioned text tag feature information to obtain normalized text tag feature information. For example, L2 norm normalization can be performed on the aforementioned text tag feature information to obtain normalized text tag feature information.

[0043] The third sub-step involves mapping the normalized image feature information and the normalized text tag feature information to the joint embedding space to obtain the mapped image feature information and the mapped text tag feature information.

[0044] The fourth sub-step involves performing the following processing steps for each pixel of the aforementioned face segmentation image:

[0045] 1. The feature vector corresponding to the above-mentioned pixel in the above-mentioned mapped image feature information is determined as the first feature vector.

[0046] 2. The similarity between the first feature vector and each feature vector in the mapped text tag feature information is determined as the feature vector similarity. For example, the similarity can be cosine similarity.

[0047] 3. The feature vectors whose similarity among the feature vectors included in the above-mentioned mapped text tag feature information meet the preset similarity conditions are determined as the target feature vectors. The preset similarity conditions can be: maximum feature vector similarity.

[0048] The fifth sub-step involves determining the obtained target feature vectors as text feature information corresponding to the above-mentioned mapped image feature information.

[0049] Step 1034: Input the first face image feature information and the text feature information into the output layer to obtain the segmentation quality classification information corresponding to the face segmentation image.

[0050] In some embodiments, the executing entity can input the first face image feature information and the text feature information into the output layer to obtain segmentation quality classification information corresponding to the face segmentation image. The activation function of the output layer can be a sigmoid function for binary classification output. The output segmentation quality classification information can indicate whether the face segmentation image is qualified or unqualified. For example, if the face in the segmented face image is incomplete, it is unqualified.

[0051] In practice, the aforementioned executing entity can input the aforementioned first facial image feature information and the aforementioned text feature information into the aforementioned output layer through the following steps:

[0052] The first step is to normalize the aforementioned first face image feature information to obtain first normalized face image feature information. For example, the aforementioned first face image feature information can be normalized using the L2 norm to obtain first normalized face image feature information.

[0053] The second step is to normalize the above text feature information to obtain normalized text feature information. For example, the above text feature information can be normalized using the L2 norm to obtain normalized text feature information.

[0054] The third step involves generating image segmentation quality confidence information based on the first normalized face image feature information and the normalized text feature information. This image segmentation quality confidence information characterizes the segmentation quality of the face segmentation image. For example, the executing entity can first generate a similarity matrix between the first normalized face image feature information and the normalized text feature information. One similarity value in the similarity matrix corresponds to one pixel, and the similarity value can be the cosine similarity between the image features and semantic features corresponding to that pixel. Then, based on the similarity matrix, a confidence value corresponding to the face segmentation image can be generated as the image segmentation quality confidence information. For example, the executing entity can input the similarity matrix into a normalized exponential function to obtain the pixel confidence value corresponding to each pixel. Then, the mean of the obtained pixel confidence values ​​can be determined as the image segmentation quality confidence information corresponding to the face segmentation image.

[0055] Fourth, in response to determining that the above image segmentation quality confidence information meets the preset confidence condition, the segmentation result representing the qualified segmentation is determined as the segmentation quality classification information. The preset confidence condition can be that the confidence level represented by the image segmentation quality confidence information is greater than or equal to the preset confidence level.

[0056] It should be noted that the sample dataset for training the image segmentation and classification model can be individual segmented images that have been pre-labeled with segmentation quality category information and text label information. The image segmentation and classification model can be trained using a batch training method.

[0057] Step 1035: In response to determining that the above segmentation quality classification information meets the preset segmentation conditions, the above face segmentation image is determined as the target face segmentation image.

[0058] In some embodiments, the execution entity may determine the face segmentation image as the target face segmentation image in response to determining that the segmentation quality classification information meets preset segmentation conditions. The preset segmentation conditions may be: the segmentation quality classification information indicates that the face segmentation image is qualified.

[0059] Step 104: Perform user recognition on the generated target face segmentation image sequence to obtain the user recognition result sequence.

[0060] In some embodiments, the execution entity can perform user identification on the generated target face segmentation image sequence to obtain a user identification result sequence. For example, the execution entity can use a pre-trained face recognition model to perform face recognition on each target face segmentation image, thereby obtaining a user identification result. The user identification result can indicate whether the user represented by the target face segmentation image is a "stranger". As another example, a similarity calculation can be performed on each target face segmentation image using a pre-established face image database to generate a user identification result. Thus, it can be determined whether the user corresponding to the target face segmentation image is a "stranger".

[0061] Step 105: Perform voice optimization processing on the audio corresponding to the above-mentioned surveillance video to obtain optimized audio.

[0062] In some embodiments, the aforementioned executing entity may perform voice optimization processing on the audio corresponding to the aforementioned surveillance video to obtain optimized audio. For example, audio noise reduction optimization processing may be performed on the audio corresponding to the aforementioned surveillance video to obtain optimized audio. The audio corresponding to the surveillance video may be audio synchronously captured during video acquisition.

[0063] In practice, the aforementioned implementing entities can perform voice optimization processing on the audio corresponding to the aforementioned surveillance video through the following steps:

[0064] The first step is to input the aforementioned audio into a pre-trained dropped-frame speech generation model to obtain dropped-frame speech data. This dropped-frame speech data includes the dropped frame location and the reconstructed dropped-frame speech. The dropped-frame speech generation model can be a pre-trained dropped-frame speech detection and generation model that takes audio as input and dropped-frame speech data as output. For example, the dropped-frame speech generation model can be a CPC (Contrastive Predictive Coding) model or an APC (Autoregressive Predictive Coding) model. CPC (Contrastive Predictive Coding): CPC is an earlier proposed pre-trained speech model. Its model structure involves segmenting the speech signal and inputting it into a CNN for feature extraction. Then, the output of the CNN layer is used as the input to a GRU layer to obtain an output with temporal information. The output with temporal information at the current time step is used to predict the CNN layer output at subsequent time steps. Contrastive loss is used during training to optimize the model and enhance prediction accuracy. APC (Autoregressive Predictive Coding) and its improved VQ-APC: APC consists of 3 layers of LSTM and is trained using contrastive loss to enhance the model's predictive ability. VQ-APC adds a VQ layer to APC, which makes the learned vector representation more advantageous through the vector quantization process, thus performing better in downstream tasks such as speech classification.

[0065] The second step is to add the reconstructed speech from the lost frames, which is included in the lost frame speech data, to the corresponding lost frame position in the audio to optimize the audio and obtain optimized audio.

[0066] The dropped-frame speech generation model can be trained through the following steps:

[0067] The first step is to obtain audio data samples. These samples include: audio data after frame loss and frame loss information. The audio data after frame loss typically contains audio data with speech loss. The frame loss information is usually used to characterize whether speech is lost in each frame of the audio data.

[0068] The second step involves inputting the audio data samples into the first encoder of the initial frame-dropping speech generation model to obtain the first audio features. The first encoder is used to extract features from the audio data samples that are related to the audio at the frame drop location. For example, the first encoder can employ a multi-layer convolutional neural network structure. As an example, the executing entity can input the audio data samples into the first encoder and perform feature extraction through multiple convolutional layers. The audio features output by the first encoder can then be used as the first audio features. Alternatively, the first encoder can also adopt the encoder structure in the generator of a frame-dropping compensation model based on a GAN (Generative Adversarial Network) framework.

[0069] The third step involves inputting the aforementioned audio data samples into the second encoder included in the initial dropped-frame speech generation model to obtain the second audio features. This second encoder employs the encoder from a pre-trained automatic speech recognition model to extract features from the audio data samples. The second encoder can be the encoder from a pre-trained automatic speech recognition (ASR) model. For example, an open-source speech recognition model (whisper-large) structure can be used. Whisper is a large ASR model that has undergone extensive training. The encoder part of Whisper uses a multi-layered transformer (a model using a self-attention mechanism) encoder stack structure to extract ASR-related features, which are then provided to the decoder to generate the recognition result. Alternatively, the second encoder can be the encoder from other large ASR models.

[0070] The fourth step involves fusing the first audio feature with the second audio feature to generate a fused audio feature. For example, the executing entity can superimpose and fuse the first audio feature with the second audio feature to obtain the fused audio feature.

[0071] The fourth step mentioned above may include the following sub-steps:

[0072] The first sub-step is to generate the query vector corresponding to the first audio feature mentioned above.

[0073] The second sub-step generates the key vector and value vector corresponding to the second audio feature mentioned above.

[0074] The third sub-step involves generating attention weights based on the query vector and the key vector described above.

[0075] The fourth sub-step involves weighting and summing the attention weights and the value vectors to obtain the attention result of the second audio feature relative to the first audio feature.

[0076] The fifth sub-step involves concatenating the attention results with the first audio feature to obtain the fused audio feature.

[0077] For example, the input to the attention fusion module is the first audio feature output from the encoder's intermediate layer and the second audio feature output from the ASR encoder. The first audio feature is transformed linearly through the query matrix to obtain the query vector Q. The second audio feature is transformed linearly to obtain the key vector K and the value vector V, respectively. The linear transformation also makes the second dimension of the query vector and the key vector the same, which facilitates the calculation of the influence of the ASR encoder on the shallow encoding of the encoder at different times by matrix multiplication of the query vector and the transpose of the key vector. Then, the multiplication result is masked and softmaxed to obtain weights representing the degree of influence. The obtained weight information is multiplied by the corresponding value vector and then summed to obtain the attention result of the ASR encoder encoding relative to the shallow encoding. This result is then concatenated with the original first audio feature along the feature dimension to obtain the final fused audio feature.

[0078] Fifth, based on the fused audio features described above, generate the predicted speech at the frame loss points of the audio data samples. The fused audio features can be input into the decoder to obtain the predicted speech at the frame loss points. The decoder may contain a deconvolution layer corresponding to the first encoder, which restores the feature representation encoded by the first encoder to the lost speech data.

[0079] The sixth step is to determine the speech loss value between the predicted speech and the corresponding sample label. For example, a preset loss function can be used to determine the speech loss value between the predicted speech and the corresponding sample label. The preset loss function can be a cross-entropy loss function or a hinge loss function.

[0080] Step 7: In response to determining that the above speech loss value is less than or equal to the preset loss value, the initial dropped frame speech generation model is determined as the trained dropped frame speech generation model.

[0081] Therefore, by fusing audio features, the extracted audio features related to the speech at the frame loss point can contain ASR coding information, thus enriching the extracted audio feature data. Using this ASR coding information to assist in generating predicted speech at the frame loss point can improve the prediction accuracy and compensation effect of the lost frame speech. This, in turn, can improve the monitoring audio and ensure the effectiveness of subsequent detection.

[0082] Step 106: Send the above user identification result sequence and the above optimized audio to the associated monitoring and management terminal.

[0083] In some embodiments, the aforementioned executing entity may send the aforementioned user identification result sequence and the aforementioned optimized audio to an associated monitoring and management terminal. The monitoring and management terminal may be a terminal that further analyzes and detects the user identification result sequence and the aforementioned optimized audio.

[0084] Further reference Figure 2 As an implementation of the methods shown in the above figures, this disclosure provides some embodiments of an audio and video processing apparatus, which are similar to... Figure 1 Corresponding to the method embodiments shown, this audio and video processing device can be specifically applied to various electronic devices.

[0085] like Figure 2 As shown, the audio / video processing apparatus 200 in some embodiments includes: an extraction unit 201, a segmentation unit 202, a segmentation detection unit 203, a recognition unit 204, an optimization unit 205, and a transmission unit 206. The extraction unit 201 is configured to extract video frames from a surveillance video to obtain a surveillance video image sequence; the segmentation unit 202 is configured to perform face segmentation on the faces in each surveillance video image in the surveillance video image sequence to generate face segmentation images, obtaining a face segmentation image sequence, and to add text tags to each face segmentation image in the face segmentation image sequence; the segmentation detection unit 203 is configured to perform the following processing steps on each face segmentation image in the face segmentation image sequence: perform image preprocessing on the face segmentation image to obtain a preprocessed face segmentation image, wherein the face segmentation image corresponds to text tag information; input the preprocessed face segmentation image into an image feature extraction network included in a pre-trained image segmentation classification model to obtain first face image feature information, wherein the image segmentation classification model includes an image feature extraction network and a semantic feature extraction network. The semantic feature extraction network, including a text encoder and an image encoder, is configured to: input the preprocessed face segmentation image into the image encoder and input the text tag information into the text encoder to obtain text feature information; input the first face image feature information and the text feature information into the output layer to obtain segmentation quality classification information corresponding to the face segmentation image; in response to determining that the segmentation quality classification information meets preset segmentation conditions, the face segmentation image is identified as the target face segmentation image; the recognition unit 204 is configured to perform user recognition on the generated target face segmentation image sequence to obtain a user recognition result sequence; the optimization unit 205 is configured to perform voice optimization processing on the audio corresponding to the monitoring video to obtain optimized audio; and the sending unit 206 is configured to send the user recognition result sequence and the optimized audio to an associated monitoring management terminal.

[0086] It is understandable that the units described in the audio / video processing device 200 are related to the reference. Figure 1 The steps in the described method correspond to each other. Therefore, the operations, features, and beneficial effects described above for the method also apply to the audio / video processing apparatus 200 and the units contained therein, and will not be repeated here.

[0087] The following is for reference. Figure 3 This document illustrates a schematic diagram of an electronic device 300 (e.g., a computing device) suitable for implementing some embodiments of the present disclosure. The electronic devices in some embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 3 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this disclosure.

[0088] like Figure 3 As shown, the electronic device 300 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. The RAM 303 also stores various programs and data required for the operation of the electronic device 300. The processing unit 301, ROM 302, and RAM 303 are interconnected via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0089] Typically, the following devices can be connected to I / O interface 305: input devices 306 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 307 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 308 including, for example, magnetic tapes, hard disks, etc.; and communication devices 309. Communication device 309 allows electronic device 300 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 3 An electronic device 300 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively. Figure 3 Each box shown can represent a device or multiple devices as needed.

[0090] In particular, according to some embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 309, or installed from a storage device 308, or installed from a ROM 302. When the computer program is executed by the processing device 301, it performs the functions defined in the methods of some embodiments of this disclosure.

[0091] It should be noted that, in some embodiments of this disclosure, the computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In some embodiments of this disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0092] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0093] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device. The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to: extract video frames from the surveillance video to obtain a sequence of surveillance video images; perform face segmentation on the faces in each surveillance video image in the sequence of surveillance video images to generate face segmentation images, obtaining a sequence of face segmentation images; and add text tags to each face segmentation image in the sequence of face segmentation images; for each face segmentation image in the sequence of face segmentation images, perform the following processing steps: perform image preprocessing on the face segmentation images to obtain preprocessed face segmentation images, wherein the face segmentation images correspond to text tagging information; input the preprocessed face segmentation images into the image feature extraction network included in a pre-trained image segmentation and classification model to obtain first face image feature information, wherein the image segmentation... The classification model includes an image feature extraction network, a semantic feature extraction network, and an output layer. The semantic feature extraction network includes a text encoder and an image encoder. The preprocessed face segmentation image is input into the image encoder, and the text tagging information is input into the text encoder to obtain text feature information. The first face image feature information and the text feature information are input into the output layer to obtain segmentation quality classification information corresponding to the face segmentation image. In response to determining that the segmentation quality classification information meets the preset segmentation conditions, the face segmentation image is identified as the target face segmentation image. User recognition is performed on the generated target face segmentation image sequence to obtain a user recognition result sequence. Speech optimization processing is performed on the audio corresponding to the surveillance video to obtain optimized audio. The user recognition result sequence and the optimized audio are sent to the associated monitoring management terminal.

[0094] Computer program code for performing operations of some embodiments of this disclosure can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0095] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0096] The units described in some embodiments of this disclosure can be implemented in software or hardware. The described units can also be housed in a processor; for example, a processor can be described as including: an extraction unit, a segmentation unit, a segmentation detection unit, an identification unit, an optimization unit, and a transmission unit. The names of these units do not necessarily limit the specific unit; for example, the transmission unit can also be described as "a unit that sends the aforementioned user identification result sequence and the aforementioned optimized audio to an associated monitoring and management terminal."

[0097] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.

[0098] The above description is merely a selection of preferred embodiments of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions disclosed in the embodiments of this disclosure.

Claims

1. An audio / video processing method, comprising: Video frames are extracted from the surveillance video to obtain a sequence of surveillance video images; Face segmentation is performed on the faces in each surveillance video image in the surveillance video image sequence to generate face segmentation images, resulting in a face segmentation image sequence. Text tags are added to each face segmentation image in the face segmentation image sequence, where the text tags are information that marks the user's gender, facial features, and clothing features. For each face segmentation image in the face segmentation image sequence, perform the following processing steps: The face segmentation image is preprocessed to obtain a preprocessed face segmentation image, wherein the face segmentation image corresponds to text tagging information; The preprocessed face segmentation image is input into the image feature extraction network of the pre-trained image segmentation and classification model to obtain the first face image feature information. The image segmentation and classification model includes an image feature extraction network, a semantic feature extraction network and an output layer. The semantic feature extraction network includes a text encoder and an image encoder. The preprocessed face segmentation image is input into the image encoder, and the text tagging information is input into the text encoder to obtain text feature information, including: The preprocessed face segmentation image is input into the image encoder to obtain the second face image feature information; The text tagging information is input into the text encoder to obtain text tagging feature information; The text feature information is obtained by performing feature alignment processing on the second face image feature information and the text marker feature information, including: normalizing the second face image feature information to obtain normalized image feature information; normalizing the text marker feature information to obtain normalized text marker feature information; mapping the normalized image feature information and the normalized text marker feature information to a joint embedding space to obtain mapped image feature information and mapped text marker feature information; for each pixel of the face segmentation image, the following processing steps are performed: determining the feature vector corresponding to the pixel in the mapped image feature information as a first feature vector; determining the similarity between the first feature vector and each feature vector in the mapped text marker feature information as a feature vector similarity; determining the feature vectors whose feature vector similarity satisfies a preset similarity condition among the feature vectors included in the mapped text marker feature information as target feature vectors; and determining each obtained target feature vector as the text feature information corresponding to the mapped image feature information. The first face image feature information and the text feature information are input into the output layer to obtain the segmentation quality classification information corresponding to the face segmentation image; In response to determining that the segmentation quality classification information meets the preset segmentation conditions, the face segmentation image is identified as the target face segmentation image; User recognition is performed on the generated target face segmentation image sequence to obtain the user recognition result sequence; The audio corresponding to the surveillance video is subjected to voice optimization processing to obtain optimized audio; The user identification result sequence and the optimized audio are sent to the associated monitoring and management terminal.

2. The method according to claim 1, wherein, The step of performing voice optimization processing on the audio corresponding to the surveillance video to obtain optimized audio includes: The audio is input into a pre-trained frame-dropped speech generation model to obtain frame-dropped speech data, wherein the frame-dropped speech data includes: frame-dropped position and frame-dropped reconstructed speech. The reconstructed speech from the lost frames, included in the lost frame speech data, is added to the corresponding lost frame position in the audio to optimize the audio and obtain optimized audio.

3. The method according to claim 2, wherein, Before inputting the audio into a pre-trained dropped-frame speech generation model to obtain dropped-frame speech data, the method further includes: Obtain audio data samples, wherein the audio data samples include: audio data after frame loss and frame loss information; The audio data sample is input into the first encoder of the initial frame-dropped speech generation model to obtain the first audio feature, wherein the first encoder is used to extract the features in the audio data sample that are related to the audio at the frame drop point; The audio data sample is input into the second encoder included in the initial frame-dropping speech generation model to obtain the second audio features, wherein the second encoder is an encoder in a pre-trained automatic speech recognition model, used to extract features from the audio data sample; The first audio feature and the second audio feature are fused to generate a fused audio feature; Based on the fused audio features, predictive speech is generated for the audio data sample at the frame loss location; Determine the speech loss value between the predicted speech and the corresponding sample label; In response to determining that the speech loss value is less than or equal to a preset loss value, the initial dropped frame speech generation model is determined as the trained dropped frame speech generation model.

4. An audio / video processing apparatus, comprising: The extraction unit is configured to extract video frames from the surveillance video to obtain a sequence of surveillance video images. The segmentation unit is configured to perform face segmentation on the faces in each surveillance video image in the surveillance video image sequence to generate face segmentation images, obtain a face segmentation image sequence, and perform text tagging on each face segmentation image in the face segmentation image sequence. The text tagging is tagging information that marks the user's gender, facial features, and clothing features. The segmentation detection unit is configured to perform the following processing steps for each face segmentation image in the face segmentation image sequence: preprocessing the face segmentation images to obtain preprocessed face segmentation images, wherein the face segmentation images correspond to text tagging information; inputting the preprocessed face segmentation images into an image feature extraction network included in a pre-trained image segmentation classification model to obtain first face image feature information, wherein the image segmentation classification model includes an image feature extraction network, a semantic feature extraction network, and an output layer, and the semantic feature extraction network includes a text encoder and an image encoder; inputting the preprocessed face segmentation images into the image encoder and inputting the text tagging information into the text encoder to obtain text feature information, including: inputting the preprocessed face segmentation images into the image encoder to obtain second face image feature information; inputting the text tagging information into the text encoder to obtain text tagging feature information; performing feature alignment processing on the second face image feature information and the text tagging feature information to obtain text feature information, including: normalizing the second face image feature information. Normalized image feature information is obtained; the text marker feature information is normalized to obtain normalized text marker feature information; the normalized image feature information and the normalized text marker feature information are mapped to a joint embedding space to obtain mapped image feature information and mapped text marker feature information; for each pixel of the face segmentation image, the following processing steps are performed: the feature vector corresponding to the pixel in the mapped image feature information is determined as a first feature vector; the similarity between the first feature vector and each feature vector in the mapped text marker feature information is determined as feature vector similarity; the feature vectors whose feature vector similarity satisfies a preset similarity condition among the feature vectors included in the mapped text marker feature information are determined as target feature vectors; each target feature vector is determined as text feature information corresponding to the mapped image feature information; the first face image feature information and the text feature information are input into the output layer to obtain segmentation quality classification information corresponding to the face segmentation image; in response to determining that the segmentation quality classification information satisfies the preset segmentation condition, the face segmentation image is determined as the target face segmentation image; The recognition unit is configured to perform user recognition on the generated target face segmentation image sequence to obtain a user recognition result sequence. The optimization unit is configured to perform voice optimization processing on the audio corresponding to the monitoring video to obtain optimized audio. The sending unit is configured to send the user identification result sequence and the optimized audio to an associated monitoring and management terminal.

5. An electronic device, comprising: One or more processors; A storage device on which one or more programs are stored; When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-3.

6. A computer-readable medium having a computer program stored thereon, wherein, When the program is executed by the processor, it implements the method as described in any one of claims 1-3.

Citation Information

Patent Citations

  • Video security monitoring and face recognition system

    CN115188057A

  • Image processing method and device, computer equipment, storage medium and program product

    CN116980584A