Audio and video processing method and device, electronic equipment and computer readable medium
By extracting video frames, facial segmentation and text marking of surveillance videos, and combining with the segmentation quality classification of image feature extraction networks, the problems of low efficiency and poor accuracy in surveillance video processing are solved, and efficient user identification and audio optimization processing are achieved.
Patent Information
- Application Number
- CN202510080884.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-20
AI Technical Summary
The prior art has low efficiency in monitoring video processing, is sensitive to noise, has poor complex background resolution, and is unable to handle fuzzy edges, resulting in poor detection accuracy.
By performing video frame extraction, face segmentation and text marking on the surveillance video, combining preprocessing and image feature extraction networks, segmentation quality classification is performed, target face segmentation images are determined, and user recognition and audio optimization are performed.
The detection efficiency of surveillance video and the verification efficiency and accuracy of image segmentation are improved, and the situation in the monitoring area is fully detected.
Smart Images

Figure CN119992455A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the computer field, and in particular to an audio and video processing method, device, electronic device, and computer-readable medium. Background Art
[0002] Defense platforms in today's society usually focus on protecting important facilities and areas, such as monitoring and protecting confidential bases or critical infrastructure. At present, the processing of surveillance videos in important areas is usually done by manually inspecting the surveillance videos in real time or by detecting and identifying the surveillance videos through conventional video detection algorithms.
[0003] However, the above method usually has the following technical problems: the efficiency of manual inspection of surveillance videos is low, conventional video detection algorithms are sensitive to noise, have poor separation effect on complex backgrounds, and cannot handle blurred edges, resulting in poor accuracy of video image detection.
[0004] The above information disclosed in this Background section is only for enhancement of understanding of the background of the inventive concept and therefore it may contain information that does not form the prior art that is already known in this country to a person of ordinary skill in the art. Summary of the invention
[0005] The content of this disclosure is used to introduce concepts in a brief form, which will be described in detail in the detailed implementation section below. The content of this disclosure is not intended to identify the key features or essential features of the technical solution claimed for protection, nor is it intended to limit the scope of the technical solution claimed for protection.
[0006] Some embodiments of the present disclosure propose audio and video processing methods, devices, electronic devices, and computer-readable media to solve one or more of the technical problems mentioned in the above background technology section.
[0007] In a first aspect, some embodiments of the present disclosure provide an audio and video processing method, the method comprising: extracting video frames from a surveillance video to obtain a surveillance video image sequence; performing face segmentation on a face in each surveillance video image in the surveillance video image sequence to generate a face segmentation image to obtain a face segmentation image sequence, and performing text labeling on each face segmentation image in the face segmentation image sequence; for each face segmentation image in the face segmentation image sequence, performing the following processing steps: performing image preprocessing on the face segmentation image to obtain a preprocessed face segmentation image, wherein the face segmentation image corresponds to text labeling information; inputting the preprocessed face segmentation image into an image feature extraction network included in a pre-trained image segmentation classification model to obtain first face image feature information, wherein the image segmentation classification model includes an image feature extraction network; The method comprises the following steps: a first face segmentation image processing method and a second face segmentation image processing method; ...
[0008] In a second aspect, some embodiments of the present disclosure provide an audio and video processing device, the device comprising: an extraction unit, configured to extract video frames from a surveillance video to obtain a surveillance video image sequence; a segmentation unit, configured to perform face segmentation on a face in each surveillance video image in the surveillance video image sequence to generate a face segmentation image to obtain a face segmentation image sequence, and to perform text marking on each face segmentation image in the face segmentation image sequence; a segmentation detection unit, configured to perform the following processing steps for each face segmentation image in the face segmentation image sequence: perform image preprocessing on the face segmentation image to obtain a preprocessed face segmentation image, wherein the face segmentation image corresponds to text marking information; input the preprocessed face segmentation image into an image feature extraction network included in a pre-trained image segmentation classification model to obtain first face image feature information, wherein the image segmentation classification model includes an image A feature extraction network, a semantic feature extraction network and an output layer, wherein the semantic feature extraction network includes a text encoder and an image encoder; the preprocessed face segmentation image is input into the image encoder, and the text tag information is input into the text encoder to obtain text feature information; the first face image feature information and the text feature information are input into the output layer to obtain segmentation quality classification information corresponding to the face segmentation image; in response to determining that the segmentation quality classification information meets a preset segmentation condition, the face segmentation image is determined as a target face segmentation image; a recognition unit is configured to perform user recognition on the generated target face segmentation image sequence to obtain a user recognition result sequence; an optimization unit is configured to perform voice optimization processing on the audio corresponding to the monitoring video to obtain an optimized audio; a sending unit is configured to send the user recognition result sequence and the optimized audio to an associated monitoring management terminal.
[0009] In a third aspect, some embodiments of the present disclosure provide an electronic device comprising: one or more processors; a storage device on which one or more programs are stored, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation manner of the above-mentioned first aspect.
[0010] In a fourth aspect, some embodiments of the present disclosure provide a computer-readable medium having a computer program stored thereon, wherein when the program is executed by a processor, the method described in any implementation manner of the above-mentioned first aspect is implemented.
[0011] The efficiency of manual inspection of surveillance videos is low. Conventional video detection algorithms are sensitive to noise, have poor separation effects on complex backgrounds, and cannot handle blurred edges, resulting in poor accuracy of video image detection. The above-mentioned various embodiments of the present disclosure have the following beneficial effects: through the audio and video processing methods of some embodiments of the present disclosure, the detection efficiency of the surveillance video is improved, and the verification efficiency and accuracy of the image segmentation are improved. In addition, the audio of the collected target area can also be detected synchronously to facilitate the all-round detection of the situation in the surveillance area. First, the surveillance video is extracted for video frames to obtain a surveillance video image sequence. This facilitates frame-by-frame analysis. Secondly, face segmentation is performed on the faces in each surveillance video image in the above-mentioned surveillance video image sequence to generate a face segmentation image, obtain a face segmentation image sequence, and text labeling is performed on each face segmentation image in the above-mentioned face segmentation image sequence. This facilitates the preliminary distinction of each face image. Next, for each face segmentation image in the above face segmentation image sequence, the following processing steps are performed: image preprocessing is performed on the face segmentation image to obtain a preprocessed face segmentation image, wherein the face segmentation image corresponds to text label information; the preprocessed face segmentation image is input into the image feature extraction network included in the pre-trained image segmentation classification model to obtain first face image feature information, wherein the image segmentation classification model includes an image feature extraction network, a semantic feature extraction network and an output layer, and the semantic feature extraction network includes a text encoder and an image encoder; the preprocessed face segmentation image is input into the image encoder, and the text label information is input into the text encoder to obtain text feature information; the first face image feature information and the text feature information are input into the output layer to obtain segmentation quality classification information corresponding to the face segmentation image; in response to determining that the segmentation quality classification information meets the preset segmentation condition, the face segmentation image is determined as the target face segmentation image. Thus, the segmentation quality classification information of the segmented image can be determined in combination with the image dimension-related features and the semantic dimension-related features of the face image to achieve the detection of the image segmentation result. Also, because the segmentation quality category information is determined by combining the features of the image dimension and the text dimension, the extraction of text features can enhance the integrity of the extracted features and improve the accuracy of image segmentation detection. Then, the generated target face segmentation image sequence is subjected to user identification to obtain a user identification result sequence. In this way, each high-quality face image can be identified to determine whether it is a strange user. Finally, the audio corresponding to the above-mentioned surveillance video is subjected to voice optimization processing to obtain optimized audio; the above-mentioned user identification result sequence and the above-mentioned optimized audio are sent to the associated monitoring management terminal. In this way, the detection efficiency of the surveillance video is improved, and the verification efficiency and accuracy of the image segmentation are improved. In addition, the audio of the collected target area can also be detected synchronously to facilitate the all-round detection of the situation in the monitoring area. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that components and elements are not necessarily drawn to scale.
[0013] Figure 1 is a flow chart of some embodiments of the audio and video processing method according to the present disclosure; Figure 2 is a schematic diagram of the structure of some embodiments of the audio and video processing device according to the present disclosure; Figure 3 It is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION
[0014] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments set forth herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.
[0015] It should also be noted that, for ease of description, only the parts related to the invention are shown in the drawings. In the absence of conflict, the embodiments and features in the embodiments of the present disclosure can be combined with each other.
[0016] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0017] It should be noted that the modifications of "one" and "plurality" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".
[0018] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0019] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.
[0020] Figure 11 is a flow chart of some embodiments of the audio and video processing method according to the present disclosure. 100 of some embodiments of the audio and video processing method according to the present disclosure is shown. The audio and video processing method comprises the following steps: Step 101: extract video frames from a surveillance video to obtain a surveillance video image sequence.
[0021] In some embodiments, the execution subject (e.g., a computing device) of the audio and video processing method may extract video frames from the surveillance video to obtain a surveillance video image sequence. The surveillance video may refer to a surveillance video of a target area. The target area may refer to an area that needs to be monitored for security. For example, the target area may be a research institute, a bank, or a restricted area. The surveillance video may refer to a video covering the face of a user in the target area.
[0022] Step 102, performing face segmentation on the face in each surveillance video image in the surveillance video image sequence to generate a face segmentation image, obtain a face segmentation image sequence, and perform text marking on each face segmentation image in the face segmentation image sequence.
[0023] In some embodiments, the above-mentioned execution entity can perform face segmentation on the face in each surveillance video image in the above-mentioned surveillance video image sequence to generate a face segmentation image, obtain a face segmentation image sequence, and perform text marking on each face segmentation image in the above-mentioned face segmentation image sequence.
[0024] For example, first, a surveillance video image containing a complete face can be selected from the above-mentioned surveillance video image sequence as an alternative surveillance video image to obtain an alternative surveillance video image sequence. Secondly, the face area of each alternative surveillance video image in the alternative surveillance video image sequence is captured and segmented to generate a face segmentation image to obtain a face segmentation image sequence. Then, each face segmentation image in the above-mentioned face segmentation image sequence is text-marked. Here, the text mark can be a simple mark of the user's gender, appearance characteristics, clothing characteristics and other marking information. It should be noted that the text mark here can be either manual marking or automatic marking.
[0025] Step 103, for each face segmentation image in the above face segmentation image sequence, perform the following processing steps: Step 1031, performing image preprocessing on the face segmentation image to obtain a preprocessed face segmentation image.
[0026] In some embodiments, the execution subject may perform image preprocessing on the face segmentation image to obtain a preprocessed face segmentation image. The face segmentation image corresponds to text tag information. For example, the face segmentation image may be resized to obtain a preprocessed face segmentation image. For example, the face segmentation image may be resized to a preset image size to obtain a preprocessed face segmentation image. For another example, the face segmentation image may be vectorized to obtain a preprocessed face segmentation image.
[0027] Step 1032: input the preprocessed face segmentation image into an image feature extraction network included in a pre-trained image segmentation classification model to obtain first face image feature information.
[0028] In some embodiments, the execution subject may input the pre-processed face segmentation image into the image feature extraction network included in the pre-trained image segmentation classification model to obtain the first face image feature information. The image segmentation classification model includes an image feature extraction network, a semantic feature extraction network and an output layer, and the semantic feature extraction network includes a text encoder and an image encoder. The image segmentation classification model may be a pre-trained neural network model that takes the pre-processed face segmentation image and text tag information as input and takes the segmentation quality classification information corresponding to the pre-processed face segmentation image as output. The image feature extraction network may be used for edge feature extraction and may include: a convolutional neural network or a residual network. The image feature extraction network may also include an edge detection algorithm. The edge detection algorithm may include: a Sobel operator, a Prewitt operator, a Canny operator, and a Canny edge detection. The semantic feature extraction network may be used to extract semantic features. The semantic feature extraction network may include a text encoder and an image encoder. The text encoder may be used to extract text features of the input text. The image encoder may be used to extract image features of the input image. The text encoder may include a Transformer block or a bert. The image encoder may include a vision transformer (ViT). The semantic feature extraction network may perform feature alignment on the extracted text features and image features. For example, the execution entity may input the preprocessed face segmentation image into the convolution layer included in the image feature extraction network to obtain first face image feature information. For another example, edge detection may be performed on the preprocessed face segmentation image to obtain first face image feature information. The first image feature information may include: corner feature information, texture feature information, and shape feature information.
[0029] Step 1033: input the preprocessed face segmentation image into the image encoder, and input the text tag information into the text encoder to obtain text feature information.
[0030] In some embodiments, the execution subject may input the preprocessed face segmentation image into the image encoder, and input the text tag information into the text encoder to obtain text feature information. For example, when the image encoder includes a multi-layer ViT block and the text encoder includes a multi-layer Transformer block, the execution subject may divide the preprocessed face segmentation image into a preset number of image blocks, and flatten and linearly project each image block to a preset dimension to obtain each image feature vector corresponding to the preprocessed face segmentation image. Then, each image feature vector may be input into a multi-layer ViT block to obtain an image representation. Secondly, the execution subject may perform word segmentation processing on the text tag information to obtain a word segmentation sequence. Next, each word segmentation in the word segmentation sequence may be mapped to an embedding vector, and a learnable position embedding may be added to obtain each text feature vector corresponding to the text tag information. Then, each text feature vector may be input into a multi-layer Transformer block to obtain a text representation. Finally, the image representation and the text representation may be feature aligned to obtain text feature information corresponding to the preprocessed face segmentation image. The image representation output by the image encoder can be a global representation obtained by using the [CLS] token or average pooling. The text representation output by the text encoder can be a global text representation using the [CLS] token.
[0031] In practice, the execution subject may input the pre-processed face segmentation image into the image encoder and input the text tag information into the text encoder through the following steps: The first step is to input the preprocessed face segmentation image into the image encoder to obtain feature information of the second face image.
[0032] The second step is to input the above text mark information into the above text encoder to obtain text mark feature information.
[0033] The third step is to perform feature alignment processing on the second face image feature information and the text mark feature information to obtain text feature information.
[0034] The third step may include the following sub-steps: The first sub-step is to perform normalization processing on the second facial image feature information to obtain normalized image feature information. For example, L2 norm normalization processing can be performed on the second facial image feature information to obtain normalized image feature information.
[0035] The second sub-step is to perform normalization processing on the above text mark feature information to obtain normalized text mark feature information. For example, the above text mark feature information can be normalized by performing L2 norm normalization processing to obtain normalized text mark feature information.
[0036] In a third sub-step, the normalized image feature information and the normalized text mark feature information are mapped into a joint embedding space to obtain mapped image feature information and mapped text mark feature information.
[0037] The fourth sub-step is to perform the following processing steps for each pixel of the above face segmentation image: 1. Determine the feature vector corresponding to the pixel in the mapped image feature information as the first feature vector.
[0038] 2. Determine the similarity between the first feature vector and each feature vector in the mapped text mark feature information as feature vector similarity. For example, the similarity may be cosine similarity.
[0039] 3. Determine the feature vector whose feature vector similarity among the feature vectors included in the above-mentioned mapping text mark feature information satisfies the preset similarity condition as the target feature vector. The above-mentioned preset similarity condition may be: the feature vector similarity is the largest.
[0040] The fifth sub-step is to determine each obtained target feature vector as text feature information corresponding to the above-mentioned mapped image feature information.
[0041] Step 1034, inputting the first facial image feature information and the text feature information into the output layer to obtain segmentation quality classification information corresponding to the facial segmentation image.
[0042] In some embodiments, the execution subject may input the first face image feature information and the text feature information into the output layer to obtain the segmentation quality classification information corresponding to the face segmentation image. The activation function of the output layer may be a Sigmoid function to perform a binary classification output. The segmentation quality classification information may indicate whether the face segmentation image is qualified or unqualified. For example, if the face in the segmented face segmentation image is incomplete, it is unqualified.
[0043] In practice, the execution subject may input the first facial image feature information and the text feature information into the output layer through the following steps: The first step is to normalize the first facial image feature information to obtain first normalized facial image feature information. For example, the first facial image feature information may be normalized using the L2 norm to obtain first normalized facial image feature information.
[0044] The second step is to perform normalization processing on the above text feature information to obtain normalized text feature information. For example, the above text feature information can be normalized by performing L2 norm processing to obtain normalized text feature information.
[0045] The third step is to generate image segmentation quality confidence information based on the first normalized face image feature information and the normalized text feature information. The image segmentation quality confidence information can be used to characterize the segmentation quality of the face segmentation image. For example, the execution subject can first generate a similarity matrix of the first normalized face image feature information and the normalized text feature information. One similarity in the similarity matrix corresponds to one pixel, and the similarity can be the cosine similarity between the image feature and the semantic feature corresponding to the pixel. Then, according to the similarity matrix, the confidence corresponding to the face segmentation image can be generated as the image segmentation quality confidence information. For example, the execution subject can input the similarity matrix into a normalized exponential function to obtain the pixel confidence corresponding to each pixel. Then, the average of the obtained pixel confidences can be determined as the image segmentation quality confidence information corresponding to the face segmentation image.
[0046] In step 4, in response to determining that the image segmentation quality confidence information satisfies a preset confidence condition, determining the qualified segmentation result as segmentation quality classification information. The preset confidence condition may be that the confidence represented by the image segmentation quality confidence information is greater than or equal to the preset confidence.
[0047] It should be noted that the sample data set for training the image segmentation classification model may be each segmented image pre-labeled with segmentation quality category information and text label information. The image segmentation classification model may be trained in a batch training manner.
[0048] Step 1035: In response to determining that the segmentation quality classification information satisfies a preset segmentation condition, the face segmentation image is determined as a target face segmentation image.
[0049] In some embodiments, the execution subject may determine the face segmentation image as the target face segmentation image in response to determining that the segmentation quality classification information satisfies a preset segmentation condition. The preset segmentation condition may be: the segmentation quality classification information indicates that the face segmentation image is qualified.
[0050] Step 104, performing user recognition on the generated target face segmentation image sequence to obtain a user recognition result sequence.
[0051] In some embodiments, the execution subject may perform user recognition on the generated target face segmentation image sequence to obtain a user recognition result sequence. For example, the execution subject may perform face recognition on each target face segmentation image through a pre-trained face identity recognition model to obtain a user recognition result. The user recognition result may indicate whether the user represented by the target face segmentation image is a "stranger". For another example, a similarity calculation may be performed on each target face segmentation image through a pre-established face image library to generate a user recognition result. Thus, it is determined whether the user corresponding to the target face segmentation image is a "stranger".
[0052] Step 105, performing voice optimization processing on the audio corresponding to the above-mentioned monitoring video to obtain optimized audio.
[0053] In some embodiments, the execution subject may perform voice optimization processing on the audio corresponding to the surveillance video to obtain optimized audio. For example, the audio corresponding to the surveillance video may be subjected to audio noise reduction optimization processing to obtain optimized audio. The audio corresponding to the surveillance video may be the audio collected synchronously when the video is collected.
[0054] In practice, the above-mentioned execution subject can perform voice optimization processing on the audio corresponding to the above-mentioned surveillance video through the following steps: In the first step, the above audio is input into a pre-trained frame loss speech generation model to obtain frame loss speech data. The above frame loss speech data includes: the position of the lost frame and the reconstructed speech of the lost frame. The frame loss speech generation model can be a pre-trained frame loss speech detection generation model that takes audio as input and takes frame loss speech data as output. For example. The frame loss speech generation model can be a CPC (Contrastive Predictive Coding) model or an APC (Autoregressive Predictive Coding) model. CPC (Contrastive Predictive Coding): CPC is an earlier proposed pre-trained speech model. Its model structure includes dividing the speech signal into segments and inputting them into CNN for feature extraction. Then, the output of the CNN layer is used as the input of the GRU layer to obtain an output with time sequence information. The output with time sequence information at the current moment is used to predict the output of the CNN layer at the subsequent moment. During training, contrast loss is used to optimize the model to enhance the accuracy of the prediction. APC (Autoregressive Predictive Coding) and its improved VQ-APC: APC consists of 3 layers of LSTM and is trained through contrast loss to enhance the model prediction ability. VQ-APC adds a VQ layer on the basis of APC, and makes the learned vector representation more advantageous through the vector quantization process, so that it performs better in downstream tasks such as speech classification.
[0055] The second step is to add the lost frame reconstructed speech included in the lost frame speech data to the corresponding lost frame position in the audio to optimize the audio to obtain optimized audio.
[0056] The frame-dropping speech generation model can be trained by the following steps: The first step is to obtain an audio data sample. The audio data sample includes: audio data after frame loss and frame loss information. The audio data after frame loss is usually audio data with voice loss. The frame loss information is usually used to characterize whether there is voice loss in each frame of the audio data.
[0057] In the second step, the audio data sample is input into the first encoder included in the initial frame loss speech generation model to obtain the first audio feature. The first encoder is used to extract features related to the audio at the frame loss in the audio data sample. For example, the first encoder can adopt a multi-layer convolutional neural network structure. As an example, the execution entity can input the audio data sample into the first encoder and perform feature extraction through multiple layers of convolutional layers. Then, the audio feature output by the first encoder can be used as the first audio feature. For another example, the first encoder can also adopt the encoder structure in the generator of the frame loss compensation model based on the GAN (Generative Adversarial Network) framework.
[0058] In the third step, the audio data sample is input into the second encoder included in the initial frame loss speech generation model to obtain the second audio feature. The second encoder uses an encoder in a pre-trained automatic speech recognition model to extract the features of the audio data sample. The second encoder can use an encoder in a pre-trained automatic speech recognition model (ASR, Automatic Speech Recognition). For example, an ASR model with an open source speech recognition model (whisper-large) structure can be used. Whisper is a large ASR model that has been trained extensively. The encoder part of whisper uses a multi-layer transformer (a model using a self-attention mechanism) encoder stacking structure to extract ASR-related features and provide these features to the decoder to generate recognition results. The second encoder here can also use encoders of other large ASR models.
[0059] In the fourth step, the first audio feature and the second audio feature are fused to generate a fused audio feature. For example, the execution entity may superimpose and fuse the first audio feature and the second audio feature to obtain a fused audio feature.
[0060] The fourth step may include the following sub-steps: The first sub-step is to generate a query vector corresponding to the first audio feature.
[0061] The second sub-step is to generate a key vector and a value vector corresponding to the second audio feature.
[0062] The third sub-step is to generate attention weights based on the query vector and the key vector.
[0063] The fourth sub-step is to perform weighted summation of the attention weight and the value vector to obtain an attention result of the second audio feature relative to the first audio feature.
[0064] In the fifth sub-step, the attention result is concatenated with the first audio feature to obtain a fused audio feature.
[0065] For example, the input of the attention fusion module is the first audio feature output by the encoder intermediate layer and the second audio feature output by the ASRencoder. The first audio feature is linearly transformed through the query matrix to obtain the query vector Q. The second audio feature is linearly transformed to obtain the key vector K and the value vector V respectively. The linear transformation also makes the second dimension of the query vector and the key vector the same, which is convenient for calculating the influence of the ASR encoder on the shallow encoding of the encoder at different times through the matrix multiplication of the query vector and the transposed key vector. After that, the multiplication result is subjected to the mask and softmax operations to obtain the weight indicating the degree of influence. The obtained weight information is multiplied by the corresponding value vector for weighted summation, that is, the attention result of the ASRencoder encoding relative to the shallow encoding is obtained. It is concatenated (i.e., concat) with the original first audio feature in the feature dimension to obtain the final fused audio feature.
[0066] In the fifth step, based on the fused audio features, the predicted speech of the audio data sample at the frame loss is generated. The fused audio features can be input into a decoder to obtain the predicted speech at the frame loss. The decoder can include a deconvolution layer corresponding to the first encoder, which restores the feature representation encoded by the first encoder to the lost speech data.
[0067] Step 6: Determine the speech loss value between the predicted speech and the corresponding sample label. For example, the speech loss value between the predicted speech and the corresponding sample label can be determined by a preset loss function. The preset loss function can be a cross entropy loss function or a hinge loss function.
[0068] In the seventh step, in response to determining that the speech loss value is less than or equal to the preset loss value, the initial frame-loss speech generation model is determined as the trained frame-loss speech generation model.
[0069] Therefore, through audio feature fusion, the extracted audio features related to the speech at the frame loss location can contain the ASR coding information, thereby enriching the extracted audio feature data. In this way, the ASR coding information is used to assist in generating the predicted speech at the frame loss location, which can improve the prediction accuracy and compensation effect of the frame loss speech. Therefore, the monitoring audio can be improved to ensure the effect of subsequent detection.
[0070] Step 106: Send the user identification result sequence and the optimized audio to the associated monitoring management terminal.
[0071] In some embodiments, the execution subject may send the user identification result sequence and the optimized audio to an associated monitoring management terminal. The monitoring management terminal may be a terminal that further analyzes and detects the user identification result sequence and the optimized audio.
[0072] Further references Figure 2 As an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of an audio and video processing device. These audio and video processing device embodiments are Figure 1 Corresponding to the method embodiments shown, the audio and video processing device can be specifically applied to various electronic devices.
[0073] like Figure 2As shown, the audio and video processing device 200 of some embodiments includes: an extraction unit 201, a segmentation unit 202, a segmentation detection unit 203, a recognition unit 204, an optimization unit 205 and a sending unit 206. The extraction unit 201 is configured to extract video frames from the surveillance video to obtain a surveillance video image sequence; the segmentation unit 202 is configured to perform face segmentation on the face in each surveillance video image in the surveillance video image sequence to generate a face segmentation image to obtain a face segmentation image sequence, and to perform text marking on each face segmentation image in the face segmentation image sequence; the segmentation detection unit 203 is configured to perform the following processing steps for each face segmentation image in the face segmentation image sequence: perform image preprocessing on the face segmentation image to obtain a preprocessed face segmentation image, wherein the face segmentation image corresponds to text marking information; input the preprocessed face segmentation image into an image feature extraction network included in a pre-trained image segmentation classification model to obtain first face image feature information, wherein the image segmentation classification model includes an image feature extraction network, a semantic feature extraction network and an output layer, wherein the semantic feature extraction network includes a text encoder and an image encoder; the preprocessed face segmentation image is input into the image encoder, and the text tag information is input into the text encoder to obtain text feature information; the first face image feature information and the text feature information are input into the output layer to obtain segmentation quality classification information corresponding to the face segmentation image; in response to determining that the segmentation quality classification information satisfies a preset segmentation condition, the face segmentation image is determined as a target face segmentation image; the recognition unit 204 is configured to perform user recognition on the generated target face segmentation image sequence to obtain a user recognition result sequence; the optimization unit 205 is configured to perform voice optimization processing on the audio corresponding to the monitoring video to obtain optimized audio; the sending unit 206 is configured to send the user recognition result sequence and the optimized audio to the associated monitoring management terminal.
[0074] It is understandable that the units described in the audio and video processing device 200 are similar to those described in the reference Figure 1 Therefore, the operations, features and beneficial effects described above for the method are also applicable to the audio and video processing device 200 and the units included therein, and will not be described in detail here.
[0075] Reference below Figure 3, which shows a schematic diagram of the structure of an electronic device 300 (e.g., a computing device) suitable for implementing some embodiments of the present disclosure. The electronic devices in some embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 3 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0076] like Figure 3 As shown, the electronic device 300 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 to a random access memory (RAM) 303. In the RAM 303, various programs and data required for the operation of the electronic device 300 are also stored. The processing device 301, the ROM 302, and the RAM 303 are connected to each other via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0077] Typically, the following devices may be connected to the I / O interface 305: an input device 306 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 308 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 309. The communication device 309 may allow the electronic device 300 to communicate with other devices wirelessly or by wire to exchange data. Figure 3 The electronic device 300 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead. Figure 3 Each block shown in the figure may represent one device, or may represent multiple devices as required.
[0078] In particular, according to some embodiments of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, some embodiments of the present disclosure include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In some such embodiments, the computer program can be downloaded and installed from the network through the communication device 309, or installed from the storage device 308, or installed from the ROM 302. When the computer program is executed by the processing device 301, the above-mentioned functions defined in the method of some embodiments of the present disclosure are executed.
[0079] It should be noted that the computer-readable medium recorded in some embodiments of the present disclosure may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, device or device. In some embodiments of the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer readable signal media may also be any computer readable medium other than computer readable storage media, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0080] In some embodiments, the client and the server may communicate using any currently known or future developed network protocol such as HTTP (HyperText Transfer Protocol), and may be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.
[0081] The above-mentioned computer-readable medium may be included in the above-mentioned electronic device; or it may exist independently without being assembled into the electronic device. The above-mentioned computer-readable medium carries one or more programs. When the above-mentioned one or more programs are executed by the electronic device, the electronic device: extracts video frames from the surveillance video to obtain a surveillance video image sequence; performs face segmentation on the face in each surveillance video image in the above-mentioned surveillance video image sequence to generate a face segmentation image to obtain a face segmentation image sequence, and performs text labeling on each face segmentation image in the above-mentioned face segmentation image sequence; for each face segmentation image in the above-mentioned face segmentation image sequence, performs the following processing steps: performs image preprocessing on the face segmentation image to obtain a preprocessed face segmentation image, wherein the above-mentioned face segmentation image corresponds to text labeling information; inputs the above-mentioned preprocessed face segmentation image into an image feature extraction network included in a pre-trained image segmentation classification model to obtain first face image feature information, wherein the above-mentioned image segmentation The classification model includes an image feature extraction network, a semantic feature extraction network and an output layer, wherein the semantic feature extraction network includes a text encoder and an image encoder; the preprocessed face segmentation image is input into the image encoder, and the text tag information is input into the text encoder to obtain text feature information; the first face image feature information and the text feature information are input into the output layer to obtain segmentation quality classification information corresponding to the face segmentation image; in response to determining that the segmentation quality classification information meets a preset segmentation condition, the face segmentation image is determined as a target face segmentation image; user recognition is performed on the generated target face segmentation image sequence to obtain a user recognition result sequence; voice optimization processing is performed on the audio corresponding to the monitoring video to obtain optimized audio; the user recognition result sequence and the optimized audio are sent to an associated monitoring management terminal.
[0082] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0083] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0084] The units described in some embodiments of the present disclosure may be implemented by software or hardware. The described units may also be set in a processor, for example, it may be described as: a processor includes: an extraction unit, a segmentation unit, a segmentation detection unit, a recognition unit, an optimization unit and a sending unit. Among them, the names of these units do not constitute a limitation on the unit itself under certain circumstances. For example, the sending unit may also be described as "a unit that sends the above-mentioned user recognition result sequence and the above-mentioned optimized audio to the associated monitoring management terminal".
[0085] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.
[0086] The above descriptions are only some preferred embodiments of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the above features are replaced with (but not limited to) technical features with similar functions disclosed in the embodiments of the present disclosure.
Claims
1. An audio and video processing method, comprising: Extracting video frames from surveillance video to obtain surveillance video image sequences; Performing face segmentation on the face in each surveillance video image in the surveillance video image sequence to generate a face segmentation image, obtaining a face segmentation image sequence, and performing text marking on each face segmentation image in the face segmentation image sequence; For each face segmentation image in the face segmentation image sequence, the following processing steps are performed: Performing image preprocessing on the face segmentation image to obtain a preprocessed face segmentation image, wherein the face segmentation image corresponds to text labeling information; Inputting the preprocessed face segmentation image into an image feature extraction network included in a pre-trained image segmentation classification model to obtain first face image feature information, wherein the image segmentation classification model includes an image feature extraction network, a semantic feature extraction network and an output layer, and the semantic feature extraction network includes a text encoder and an image encoder; Inputting the preprocessed face segmentation image into the image encoder, and inputting the text tag information into the text encoder to obtain text feature information; Inputting the first face image feature information and the text feature information into the output layer to obtain segmentation quality classification information corresponding to the face segmentation image; In response to determining that the segmentation quality classification information satisfies a preset segmentation condition, determining the face segmentation image as a target face segmentation image; Perform user recognition on the generated target face segmentation image sequence to obtain a user recognition result sequence; Performing voice optimization processing on the audio corresponding to the surveillance video to obtain optimized audio; The user identification result sequence and the optimized audio are sent to an associated monitoring management terminal.
2. The method according to claim 1, wherein: The performing voice optimization processing on the audio corresponding to the surveillance video to obtain optimized audio includes: Input the audio into a pre-trained frame-dropped speech generation model to obtain frame-dropped speech data, wherein the frame-dropped speech data includes: a frame-dropped position and a frame-dropped reconstructed speech; The lost frame reconstructed speech included in the lost frame speech data is added to the corresponding lost frame position in the audio to optimize the audio and obtain optimized audio.
3. The method according to claim 2, wherein: Before inputting the audio into a pre-trained frame-dropped speech generation model to obtain frame-dropped speech data, the method further includes: Acquire an audio data sample, wherein the audio data sample includes: audio data after frame loss and frame loss information; Inputting the audio data sample into a first encoder included in the initial frame loss speech generation model to obtain a first audio feature, wherein the first encoder is used to extract features related to the audio at the frame loss in the audio data sample; Inputting the audio data sample into a second encoder included in the initial frame-dropping speech generation model to obtain a second audio feature, wherein the second encoder uses an encoder in a pre-trained automatic speech recognition model to extract features of the audio data sample; Fusing the first audio feature with the second audio feature to generate a fused audio feature; Based on the fused audio features, generating predicted speech of the audio data sample at the frame loss position; Determining a speech loss value between the predicted speech and the corresponding sample label; In response to determining that the speech loss value is less than or equal to a preset loss value, the initial frame-loss speech generation model is determined as a trained frame-loss speech generation model.
4. The method according to claim 1, wherein: The step of inputting the preprocessed face segmentation image into the image encoder and inputting the text tag information into the text encoder to obtain text feature information includes: Inputting the preprocessed face segmentation image into the image encoder to obtain feature information of a second face image; Inputting the text mark information into the text encoder to obtain text mark feature information; Perform feature alignment processing on the second facial image feature information and the text mark feature information to obtain text feature information.
5. An audio and video processing device, comprising: An extraction unit is configured to extract video frames from the surveillance video to obtain a surveillance video image sequence; a segmentation unit configured to perform face segmentation on a face in each surveillance video image in the surveillance video image sequence to generate a face segmentation image, obtain a face segmentation image sequence, and perform text marking on each face segmentation image in the face segmentation image sequence; The segmentation detection unit is configured to perform the following processing steps for each face segmentation image in the face segmentation image sequence: perform image preprocessing on the face segmentation image to obtain a preprocessed face segmentation image, wherein the face segmentation image corresponds to text labeling information; input the preprocessed face segmentation image into an image feature extraction network included in a pre-trained image segmentation classification model to obtain first face image feature information, wherein the image segmentation classification model includes an image feature extraction network, a semantic feature extraction network and an output layer, and the semantic feature extraction network includes a text encoder and an image encoder; input the preprocessed face segmentation image into the image encoder, and input the text labeling information into the text encoder to obtain text feature information; input the first face image feature information and the text feature information into the output layer to obtain segmentation quality classification information corresponding to the face segmentation image; in response to determining that the segmentation quality classification information meets a preset segmentation condition, determine the face segmentation image as a target face segmentation image; The recognition unit is configured to perform user recognition on the generated target face segmentation image sequence to obtain a user recognition result sequence; An optimization unit is configured to perform voice optimization processing on the audio corresponding to the surveillance video to obtain optimized audio; The sending unit is configured to send the user identification result sequence and the optimized audio to an associated monitoring management terminal.
6. An electronic device comprising: one or more processors; a storage device having one or more programs stored thereon; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 4.
7. A computer readable medium having a computer program stored thereon, wherein: When the program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.
Citation Information
Patent Citations
Image processing method and device, electronic equipment and computer readable storage medium
CN113435354A
Video security monitoring and face recognition system
CN115188057A
Image processing method and device, computer equipment, storage medium and program product
CN116980584A
Face image recognition method and system based on computer vision
CN117690178A
Method and device for warning dangerous events of elderly living alone based on household Internet of Things controller
CN119152641A