Cloud video collaborative processing method and system, and storage medium
By using speech recognition and speech synthesis technologies, sensitive sentences in cloud videos can be identified and replaced, solving the problem of logical breaks in cloud videos and improving the viewing experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-09
- Publication Date
- 2026-03-24
AI Technical Summary
In the current cloud video processing, the muting of sensitive words causes logical breaks in the cloud video content, reducing the viewing experience.
Voice data is acquired through speech recognition technology, sensitive sentences are identified and replaced, and a speech synthesis model is used to generate replacement speech. The replacement speech is then applied to the video to maintain semantic coherence.
Effectively remove sensitive content, maintain the semantic and logical consistency of cloud videos, and improve the viewing experience.
Smart Images

Figure CN121483241B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of speech recognition technology, and in particular to a cloud video collaborative processing method, system, and storage medium. Background Technology
[0002] Cloudvideo refers to a video network platform service based on a cloud computing business model. On the cloud platform, all video providers, agents, planning service providers, producers, industry associations, management agencies, industry media, legal structures, etc., are centrally integrated into a resource pool. These resources can be displayed and interacted with each other, communicate on demand, and reach agreements, thereby reducing costs and improving efficiency.
[0003] In existing cloud video processing, sensitive words in the cloud video are usually muted, which leads to logical breaks in the cloud video content and reduces the cloud video viewing experience. Summary of the Invention
[0004] The purpose of this invention is to provide a cloud video collaborative processing method, system, and storage medium to solve the problem of poor cloud video viewing experience in the prior art.
[0005] This invention is implemented as follows: a cloud video collaborative processing method, the method comprising:
[0006] Obtain the cloud video to be processed uploaded by the user, and extract the audio data from the cloud video to be processed;
[0007] The audio data is classified to obtain the dialogue speech, and the dialogue speech is then subjected to speech recognition to obtain the speech recognition result;
[0008] Based on the speech recognition results, sensitive words in the dialogue are identified, and the video sentences corresponding to the sensitive words are identified as sensitive sentences.
[0009] Semantic recognition is performed on the sensitive sentences to obtain the semantics of the sensitive sentences, and a sensitive vector is constructed based on the sensitive words in the speech and the semantics of the sensitive sentences;
[0010] The replacement speech is determined based on the sensitivity vector, and the sensitive sentences are replaced in the cloud video to be processed based on the replacement speech to obtain the output cloud video.
[0011] Determining the replacement speech based on the sensitivity vector includes:
[0012] The sensitive vector is matched with the replacement database to obtain the replacement text. The replacement database stores the correspondence between different sensitive vectors and corresponding replacement texts. The replacement text is text with the same or similar semantics preset for the corresponding sensitive vector and without sensitive content.
[0013] The timbre features of the sensitive sentences are extracted, and the timbre features and the replacement text are input into a pre-trained speech synthesis model to synthesize the replacement speech.
[0014] Preferably, constructing a sensitivity vector based on the semantics of the speech-sensitive words and the sensitive sentences includes:
[0015] Obtain the sensitivity type and sensitivity score of the speech sensitive words, and construct a sensitive word vector based on the sensitivity type and sensitivity score;
[0016] Calculate the semantic similarity between the sensitive sentence semantics and the preset semantics, and determine the preset semantics corresponding to the maximum semantic similarity as the target semantics. The preset semantics is a pre-constructed semantic vector library, which contains a variety of common semantic categories and their corresponding vector representations.
[0017] The semantic vector corresponding to the target semantic is determined as the sensitive semantic vector, and the sensitive semantic vector and the sensitive word vector are combined to obtain the sensitive vector.
[0018] Preferably, the speech data is classified into audio segments to obtain dialogue speech, including:
[0019] Zero-crossing detection is performed on the speech data, and silence filtering is performed on the speech data based on the zero-crossing detection results to obtain silence-filtered speech;
[0020] The silence-filtered speech is subjected to frequency domain filtering to obtain frequency domain filtered speech, and the speech features of the frequency domain filtered speech are extracted.
[0021] The speech features are classified to obtain the dialogue speech.
[0022] Preferably, the silence-filtering speech is subjected to frequency domain filtering to obtain frequency domain filtered speech, including:
[0023] The odd-numbered speech signals in the silence-filtered speech are flipped according to the first signal flipping function to obtain the odd-numbered flipped signal.
[0024] The even-numbered speech signals in the silence-filtered speech are flipped according to the second signal flipping function to obtain even-numbered flipped signals.
[0025] The odd-numbered flip signal and the even-numbered flip signal are combined to obtain the speech reconstruction signal, and the real part of the speech reconstruction signal is processed by fast Fourier transform to obtain the real part feature signal.
[0026] The imaginary part of the reconstructed speech signal is subjected to discrete cosine transform to obtain the imaginary part feature signal, and the real part feature signal and the imaginary part feature signal are concatenated to obtain the concatenated feature signal;
[0027] The spliced feature signal is subjected to an inverse real part fast Fourier transform to obtain an inverse transform feature signal, and the inverse transform feature signal is filtered according to a preset frequency domain signal to obtain the frequency domain filtered speech.
[0028] Preferably, after replacing the sensitive statement according to the replaced speech, the method further includes:
[0029] The replacement text is segmented into words to obtain text segments. The text segments are then matched with a mouth opening and closing degree query database to obtain the mouth opening and closing degree. The mouth opening and closing degree query database stores the correspondence between different text segments and the corresponding mouth opening and closing degree.
[0030] The process involves obtaining the image of the sensitive statement and determining the speaker in the image based on the statement identifier of the sensitive statement. The statement identifier is used to mark the image position of the speaker in the cloud video to be processed.
[0031] In the statement screen, the speaker's mouth shape is modified according to the degree of mouth opening to obtain a replacement screen, and the statement screen is replaced according to the replacement screen.
[0032] Preferably, in the sentence frame, the speaker's lip shape is modified according to the degree of mouth opening to obtain a replacement frame, including:
[0033] Obtain an image of the mouth region in the statement frame, and determine a mouth-related image based on the mouth region image;
[0034] Vector transformation is performed on the image features of the mouth region image and the mouth-related image respectively to obtain image feature vector and related feature vector;
[0035] The first matrix dimension is constructed based on the high-level semantic vector, the mid-level image vector, and the low-level appearance vector in the image feature vector;
[0036] Calculate the feature association strength between the mouth region image and the mouth-related image, and construct a second matrix dimension based on the feature association strength;
[0037] The third matrix dimension is determined based on the mouth opening degree, the temporal weight of the cloud video to be processed is obtained, and the fourth matrix dimension is constructed based on the temporal weight.
[0038] A feature mapping matrix is constructed based on the first matrix dimension, the second matrix dimension, the third matrix dimension, and the fourth matrix dimension, and the feature mapping matrix is decoded to obtain the replacement features;
[0039] A replacement image is generated based on the replacement features, and in the sentence screen, the mouth region image and the mouth-related image are replaced based on the replacement image to obtain the replacement screen.
[0040] Another objective of this invention is to provide a cloud video collaborative processing system, the system comprising:
[0041] The voice extraction module is used to acquire cloud videos uploaded by users and extract voice data from the cloud videos.
[0042] The speech recognition module is used to classify the speech data to obtain dialogue speech, and to perform speech recognition on the dialogue speech to obtain speech recognition results;
[0043] The sensitivity detection module is used to determine the sensitive words in the dialogue based on the speech recognition results, and to identify the video sentences corresponding to the sensitive words as sensitive sentences.
[0044] The vector construction module is used to perform semantic recognition on the sensitive sentences, obtain the semantics of the sensitive sentences, and construct sensitive vectors based on the sensitive words in the speech and the semantics of the sensitive sentences;
[0045] The collaborative processing module is used to determine the replacement speech based on the sensitivity vector, and to replace the sensitive sentences in the cloud video to be processed based on the replacement speech to obtain the output cloud video;
[0046] The collaborative processing module is further configured to: match the sensitive vector with the replacement database to obtain replacement text, wherein the replacement database stores the correspondence between different sensitive vectors and corresponding replacement texts, and the replacement text is text with the same or similar semantics preset for the corresponding sensitive vector and without sensitive content;
[0047] The timbre features of the sensitive sentences are extracted, and the timbre features and the replacement text are input into a pre-trained speech synthesis model to synthesize the replacement speech.
[0048] Preferably, the vector construction module is further used for:
[0049] Obtain the sensitivity type and sensitivity score of the speech sensitive words, and construct a sensitive word vector based on the sensitivity type and sensitivity score;
[0050] Calculate the semantic similarity between the sensitive sentence semantics and the preset semantics, and determine the preset semantics corresponding to the maximum semantic similarity as the target semantics. The preset semantics is a pre-constructed semantic vector library, which contains a variety of common semantic categories and their corresponding vector representations.
[0051] The semantic vector corresponding to the target semantic is determined as the sensitive semantic vector, and the sensitive semantic vector and the sensitive word vector are combined to obtain the sensitive vector.
[0052] In this embodiment of the invention, by performing audio classification on the speech data, the dialogue speech in the speech data can be effectively obtained. By performing speech recognition on the dialogue speech, sensitive words and sentences in the dialogue speech can be effectively identified based on the speech recognition results. Sensitive vectors are constructed based on the semantics of the sensitive words and sentences. Based on the sensitive vectors, the speech can be effectively replaced. By replacing the sensitive sentences with the replaced speech, the sensitive content in the cloud video to be processed can be effectively deleted, while ensuring the semantic logic coherence in the cloud video to be processed, thus improving the cloud video viewing experience. Attached Figure Description
[0053] Figure 1 This is a flowchart of the cloud video collaborative processing method provided in the first embodiment of the present invention;
[0054] Figure 2 This is a schematic diagram of the cloud video collaborative processing system provided in the second embodiment of the present invention;
[0055] Figure 3 This is a schematic diagram of the structure of the terminal device provided in the third embodiment of the present invention. Detailed Implementation
[0056] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0057] To illustrate the technical solution described in this invention, specific embodiments are described below.
[0058] Example 1
[0059] Please see Figure 1 This is a flowchart of a cloud video collaborative processing method provided in the first embodiment of the present invention. This cloud video collaborative processing method can be applied to any device or system, and includes the following steps:
[0060] Step S10: Obtain the cloud video to be processed uploaded by the user, and extract the voice data from the cloud video to be processed;
[0061] Step S20: Perform audio classification on the voice data to obtain dialogue voice, and perform speech recognition on the dialogue voice to obtain speech recognition results;
[0062] In this process, the audio acoustic features of each audio segment in the speech data can be extracted, and the similarity between the audio acoustic features and preset features can be calculated to obtain the feature similarity. Audio with a feature similarity greater than the preset similarity is identified as dialogue speech. The preset features and preset similarity can be set according to the requirements. In this step, the dialogue speech is converted into text to obtain the speech recognition result.
[0063] Optionally, the speech data is subjected to audio classification to obtain dialogue speech, including:
[0064] Zero-crossing detection is performed on the speech data, and silence filtering is performed on the speech data based on the zero-crossing detection results to obtain silence-filtered speech; wherein, by performing zero-crossing detection on the speech data, the speech position of the silence segment in the speech data is obtained, and silence filtering is performed on the silence segment based on the speech position of the silence segment to obtain silence-filtered speech.
[0065] The silence-filtered speech is subjected to frequency domain filtering to obtain frequency domain filtered speech. The speech features of the frequency domain filtered speech are extracted and classified to obtain the dialogue speech. In this process, frequency domain filtering of the silence-filtered speech removes frequency domain interference to the speech signal, thereby improving the speech quality of the frequency domain filtered speech. The dialogue speech in the speech data is extracted by classifying the speech features.
[0066] Further, the silence-filtered speech is subjected to frequency domain filtering to obtain frequency domain filtered speech, including:
[0067] The odd-numbered speech signals in the silence-filtered speech are flipped according to the first signal flipping function to obtain the odd-numbered flipped signal; wherein, the first signal flipping function can be set according to requirements, and the first signal flipping function is used to flip the odd-numbered speech signals in the frequency domain.
[0068] The even-numbered speech signals in the silence filtering speech are flipped according to the second signal flipping function to obtain an even-numbered flipped signal; wherein, the second signal flipping function can be set according to requirements, and the second signal flipping function is used to flip the even-numbered speech signals in the frequency domain;
[0069] The odd-numbered flip signal and the even-numbered flip signal are combined to obtain the speech reconstruction signal. The real part of the speech reconstruction signal is then processed by Fast Fourier Transform to obtain the real part feature signal. By processing the real part of the speech reconstruction signal with Fast Fourier Transform, the real part of the speech reconstruction signal can be effectively transformed from the time domain to the frequency domain, and the time domain fluctuations that are difficult to analyze directly are transformed into intuitive frequency distribution features, which facilitates the subsequent signal splicing operation.
[0070] The imaginary part of the reconstructed speech signal is subjected to discrete cosine transform to obtain an imaginary feature signal. The real feature signal and the imaginary feature signal are then concatenated to obtain a concatenated feature signal. In this process, the imaginary part of the reconstructed speech signal is subjected to discrete cosine transform and combined with the real part Fourier transform feature to form a two-dimensional feature combination of real part energy and imaginary part phase. This supplements spectral details (such as formant phase boundaries and phase synchronization of vocal organ vibrations) and solves the classification ambiguity problem of single real feature in low signal-to-noise ratio and multi-interference scenarios.
[0071] The spliced feature signal is subjected to an inverse real part fast Fourier transform to obtain an inverse transform feature signal. This inverse transform feature signal is then filtered according to a preset frequency domain signal to obtain the frequency domain filtered speech. The spliced feature signal has already undergone frequency domain filtering. The inverse real part fast Fourier transform can convert the filtered effective frequency domain features into a clean time domain signal, preserving both the amplitude and rhythm characteristics of the dialogue speech while filtering out non-dialogue interference components, thus improving the signal-to-noise ratio. This preset frequency domain signal can be set as needed. The inverse transform feature signal is filtered using this preset frequency domain signal to remove noise other than human voice from the inverse transform feature signal. The preset frequency domain signal represents the frequency domain signal corresponding to noise other than human voice.
[0072] Step S30: Determine the speech-sensitive words in the dialogue based on the speech recognition results, and identify the video sentences corresponding to the speech-sensitive words as sensitive sentences;
[0073] In this process, the speech text in the speech recognition result is matched with preset sensitive words to obtain speech sensitive words. The preset sensitive words can be set according to needs.
[0074] Step S40: Perform semantic recognition on the sensitive sentences to obtain the semantics of the sensitive sentences, and construct a sensitive vector based on the sensitive words and the semantics of the sensitive sentences;
[0075] Among these methods, rule-based semantic matching algorithms or deep learning algorithms can be used to perform semantic recognition on sensitive sentences and obtain the semantics of sensitive sentences.
[0076] Optionally, a sensitivity vector is constructed based on the semantics of the speech-sensitive words and the sensitive sentences, including:
[0077] The sensitivity type and sensitivity score of the speech sensitive words are obtained, and a sensitive word vector is constructed based on the sensitivity type and sensitivity score. Specifically, the speech sensitive words are matched with a type lookup table and a score lookup table to obtain the sensitivity type and sensitivity score. The type lookup table stores the correspondence between different speech sensitive words and their corresponding sensitivity types, and the score lookup table stores the correspondence between different speech sensitive words and their corresponding sensitivity scores. In this step, the sensitivity type and sensitivity score are mapped according to a preset mapping relationship to obtain a type sequence and a score sequence. The type sequence and score sequence are combined to obtain a combined sequence, and the combined sequence is vectorized to obtain the sensitive word vector.
[0078] Calculate the semantic similarity between the sensitive sentence semantics and the preset semantics, and determine the preset semantics corresponding to the maximum semantic similarity as the target semantics. The preset semantics is a pre-constructed semantic vector library, which contains a variety of common semantic categories and their corresponding vector representations.
[0079] The semantic vector corresponding to the target semantic is determined as the sensitive semantic vector, and the sensitive semantic vector and the sensitive word vector are combined to obtain the sensitive vector.
[0080] Step S50: Determine the replacement speech based on the sensitivity vector, and replace the sensitive sentences in the cloud video to be processed based on the replacement speech to obtain the output cloud video;
[0081] Optionally, determining the replacement speech based on the sensitivity vector includes:
[0082] The sensitive vector is matched with the replacement database to obtain the replacement text. The replacement database stores the correspondence between different sensitive vectors and corresponding replacement texts. The replacement text is text with the same or similar semantics preset for the corresponding sensitive vector and without sensitive content.
[0083] The timbre features of the sensitive sentences are extracted, and the timbre features and the replacement text are input into a pre-trained speech synthesis model to synthesize the replacement speech.
[0084] Preferably, before performing speech synthesis on the pre-trained speech synthesis model using timbre features and replacement text input, the method further includes:
[0085] The sample timbre and sample text are input into the speech synthesis model to synthesize speech and obtain synthesized speech. The model loss is calculated based on the sample speech and synthesized speech corresponding to the sample timbre and sample text. The parameters of the speech synthesis model are updated based on the model loss until convergence, and the pre-trained speech synthesis model is obtained.
[0086] Furthermore, after replacing the sensitive statement based on the replaced speech, the process further includes:
[0087] The replacement text is segmented into words to obtain text words. The text words are then matched with the mouth opening and closing degree query database to obtain the mouth opening and closing degree. The mouth opening and closing degree query database stores the correspondence between different text words and corresponding mouth opening and closing degrees.
[0088] The process involves obtaining the image of the sensitive statement and determining the speaker in the image based on the statement identifier of the sensitive statement. The statement identifier is used to mark the image position of the speaker in the cloud video to be processed.
[0089] In the statement screen, the speaker's mouth shape is modified according to the degree of mouth opening to obtain a replacement screen, and the statement screen is replaced according to the replacement screen.
[0090] Furthermore, in the stated sentence screen, the speaker's lip shape is modified according to the degree of mouth opening to obtain a replacement screen, including:
[0091] The process involves acquiring a mouth region image from the sentence frame and determining a mouth-related image based on the mouth region image. Specifically, the process includes extracting the speaker's facial contour from the sentence frame, performing contour analysis on the facial contour, determining the speaker's mouth contour position based on the contour analysis results, identifying the image corresponding to the mouth contour position as the mouth region image, and identifying an image within a preset region of the mouth region image as the related image. This preset region can be set according to requirements.
[0092] Vector transformation is performed on the image features of the mouth region image and the mouth-related image respectively to obtain image feature vector and related feature vector;
[0093] A first matrix dimension is constructed based on the high-level semantic vector, mid-level image vector, and low-level appearance vector in the image feature vector. The first matrix dimension includes three matrix sub-blocks. The high-level semantic vector is filled into the first matrix sub-block, the mid-level image vector is filled into the second matrix sub-block, and the low-level appearance vector is filled into the third matrix sub-block to obtain the first matrix dimension. The dimension index of each vector in the first matrix dimension is set, the vector similarity between different components is calculated, and the dimension index of the component with a vector similarity greater than the similarity threshold is set as the adjacent index to ensure the binding between related components and strengthen the spatial correlation between matrix sub-blocks.
[0094] The feature association strength between the mouth region image and the mouth-related image is calculated. A second matrix dimension is constructed based on the feature association strength, and a third matrix dimension is determined based on the mouth opening degree. The image association relationship between the mouth region image and the mouth-related image is obtained, including key point association relationships, such as the correspondence between the lip valley key point and the nose tip key point. Based on the image association relationship, the associated sub-images between the mouth region image and the mouth-related image are combined to obtain a combined image.
[0095] Obtain the sub-features of the combined image in the corresponding image features to obtain the associated combined features. For the associated combined features corresponding to high-level semantic features, mid-level image features, and low-level appearance features, calculate the feature similarity between the features within the combination, and perform numerical mapping on the feature similarity to obtain the sub-association strength. The feature association strength includes semantic association set, image association set, and appearance association set. Each association set includes the sub-association strength corresponding to different combined images. In each association set, sort the sub-association strength in ascending order. Fill the semantic association set into the fourth matrix sub-block, the image association set into the fifth matrix sub-block, and the appearance association set into the sixth matrix sub-block to obtain the second matrix dimension. Obtain the matrix information corresponding to the mouth opening degree in the mouth opening degree query database to obtain the third matrix dimension.
[0096] The temporal weights of the cloud video to be processed are obtained, and a fourth matrix dimension is constructed based on the temporal weights. Specifically, the pixel differences between adjacent video frames in the cloud video to be processed are calculated, candidate frames are determined based on the pixel differences, style clustering is performed on the candidate frames, the candidate frames are filtered based on the style clustering results to obtain key frames, video frames between different key frames are set as transition frames, the pre-set weight values of key frames and transition frames are obtained to obtain temporal weights, the corresponding temporal weights are filled into the seventh matrix sub-block according to the temporal order of the key frames, and the corresponding temporal weights are filled into the eighth matrix sub-block according to the temporal order of the transition frames to obtain the fourth matrix dimension.
[0097] A feature mapping matrix is constructed based on the first matrix dimension, the second matrix dimension, the third matrix dimension, and the fourth matrix dimension, and the feature mapping matrix is decoded to obtain replacement features; wherein, according to a preset matrix mapping relationship, matrix blocks and matrix sub-blocks in the first matrix dimension, the second matrix dimension, the third matrix dimension, and the fourth matrix dimension are mapped to obtain a feature mapping matrix, and the feature mapping matrix is input into a pre-trained decoder for decoding to obtain replacement features;
[0098] A replacement image is generated based on the replacement features, and in the sentence screen, the mouth region image and the mouth-related image are replaced based on the replacement image to obtain the replacement screen.
[0099] Optionally, in this embodiment, after determining the video statement corresponding to the voice-sensitive word as a sensitive statement, the method further includes:
[0100] Extract the semantic features of sensitive statements and their context statements, and combine them with the attribute information of sensitive words to construct a sensitive vector that integrates contextual semantics;
[0101] Based on the sensitivity vector, semantically coherent replacement text is retrieved from a preset replacement library or generated through a text generation model, and the replacement text is synthesized into replacement speech.
[0102] In the cloud video to be processed, the original audio of sensitive sentences is replaced with replacement speech. Based on the duration and characteristics of the replacement speech, the video frames before and after the replacement time point are time-series adjusted and the picture is adapted to maintain audio-visual synchronization and visual coherence, thus obtaining the processed cloud video.
[0103] In this embodiment, by performing audio classification on the speech data, the dialogue speech in the speech data can be effectively obtained. By performing speech recognition on the dialogue speech, sensitive words and sentences in the dialogue speech can be effectively identified based on the speech recognition results. Sensitive vectors are constructed based on the semantics of the sensitive words and sentences. Based on the sensitive vectors, the speech can be effectively replaced. By replacing the sensitive sentences with the replaced speech, the sensitive content in the cloud video to be processed can be effectively deleted, while ensuring the semantic logic coherence in the cloud video to be processed and improving the cloud video viewing experience.
[0104] Example 2
[0105] Please see Figure 2 This is a schematic diagram of the structure of the cloud video collaborative processing system 100 provided in the second embodiment of the present invention, including:
[0106] The voice extraction module 10 is used to acquire cloud videos uploaded by users and extract voice data from the cloud videos.
[0107] The speech recognition module 11 is used to classify the speech data to obtain dialogue speech, and to perform speech recognition on the dialogue speech to obtain speech recognition results.
[0108] Optionally, the speech recognition module 11 is further configured to: perform zero-crossing detection on the speech data, and perform silence filtering on the speech data based on the zero-crossing detection result to obtain silence-filtered speech;
[0109] The silence-filtered speech is subjected to frequency domain filtering to obtain frequency domain filtered speech, and the speech features of the frequency domain filtered speech are extracted.
[0110] The speech features are classified to obtain the dialogue speech.
[0111] Furthermore, the speech recognition module 11 is also used to: perform signal flipping on the odd-numbered speech signals in the silence-filtered speech according to the first signal flipping function to obtain an odd-numbered flipped signal;
[0112] The even-numbered speech signals in the silence-filtered speech are flipped according to the second signal flipping function to obtain even-numbered flipped signals.
[0113] The odd-numbered flip signal and the even-numbered flip signal are combined to obtain the speech reconstruction signal, and the real part of the speech reconstruction signal is processed by fast Fourier transform to obtain the real part feature signal.
[0114] The imaginary part of the reconstructed speech signal is subjected to discrete cosine transform to obtain the imaginary part feature signal, and the real part feature signal and the imaginary part feature signal are concatenated to obtain the concatenated feature signal;
[0115] The spliced feature signal is subjected to an inverse real part fast Fourier transform to obtain an inverse transform feature signal, and the inverse transform feature signal is filtered according to a preset frequency domain signal to obtain the frequency domain filtered speech.
[0116] The sensitivity detection module 12 is used to determine the sensitive words in the dialogue based on the speech recognition result, and to identify the video sentences corresponding to the sensitive words as sensitive sentences.
[0117] The vector construction module 13 is used to perform semantic recognition on the sensitive sentences, obtain the semantics of the sensitive sentences, and construct a sensitive vector based on the sensitive words and the semantics of the sensitive sentences.
[0118] Optionally, the vector construction module 13 is further configured to: obtain the sensitivity type and sensitivity score of the speech sensitive words, and construct a sensitive word vector based on the sensitivity type and the sensitivity score;
[0119] Calculate the semantic similarity between the sensitive sentence semantics and the preset semantics, and determine the preset semantics corresponding to the maximum semantic similarity as the target semantics. The preset semantics is a pre-constructed semantic vector library, which contains a variety of common semantic categories and their corresponding vector representations.
[0120] The semantic vector corresponding to the target semantic is determined as the sensitive semantic vector, and the sensitive semantic vector and the sensitive word vector are combined to obtain the sensitive vector.
[0121] The collaborative processing module 14 is used to determine the replacement speech based on the sensitivity vector, and to replace the sensitive sentences in the cloud video to be processed based on the replacement speech to obtain the output cloud video.
[0122] Optionally, the collaborative processing module 14 is further configured to: match the sensitive vector with the replacement database to obtain replacement text, wherein the replacement database stores the correspondence between different sensitive vectors and corresponding replacement texts, and the replacement text is text with the same or similar semantics preset for the corresponding sensitive vector and without sensitive content;
[0123] The timbre features of the sensitive sentences are extracted, and the timbre features and the replacement text are input into a pre-trained speech synthesis model to synthesize the replacement speech.
[0124] Furthermore, the collaborative processing module 14 is also used to: segment the replacement text to obtain text segments, match the text segments with the mouth opening and closing degree query database to obtain the mouth opening and closing degree, wherein the mouth opening and closing degree query database stores the correspondence between different text segments and the corresponding mouth opening and closing degree;
[0125] The process involves obtaining the image of the sensitive statement and determining the speaker in the image based on the statement identifier of the sensitive statement. The statement identifier is used to mark the image position of the speaker in the cloud video to be processed.
[0126] In the statement screen, the speaker's mouth shape is modified according to the degree of mouth opening to obtain a replacement screen, and the statement screen is replaced according to the replacement screen.
[0127] Furthermore, the collaborative processing module 14 is also used to: acquire a mouth region image in the statement screen, and determine a mouth-related image based on the mouth region image;
[0128] Vector transformation is performed on the image features of the mouth region image and the mouth-related image respectively to obtain image feature vector and related feature vector;
[0129] The first matrix dimension is constructed based on the high-level semantic vector, the mid-level image vector, and the low-level appearance vector in the image feature vector;
[0130] Calculate the feature association strength between the mouth region image and the mouth-related image, and construct a second matrix dimension based on the feature association strength;
[0131] The third matrix dimension is determined based on the mouth opening degree, the temporal weight of the cloud video to be processed is obtained, and the fourth matrix dimension is constructed based on the temporal weight.
[0132] A feature mapping matrix is constructed based on the first matrix dimension, the second matrix dimension, the third matrix dimension, and the fourth matrix dimension, and the feature mapping matrix is decoded to obtain the replacement features;
[0133] A replacement image is generated based on the replacement features, and in the sentence screen, the mouth region image and the mouth-related image are replaced based on the replacement image to obtain the replacement screen.
[0134] In this embodiment, by performing audio classification on the speech data, the dialogue speech in the speech data can be effectively obtained. By performing speech recognition on the dialogue speech, sensitive words and sentences in the dialogue speech can be effectively identified based on the speech recognition results. Sensitive vectors are constructed based on the semantics of the sensitive words and sentences. Based on the sensitive vectors, the speech can be effectively replaced. By replacing the sensitive sentences with the replaced speech, the sensitive content in the cloud video to be processed can be effectively deleted, while ensuring the semantic logic coherence in the cloud video to be processed and improving the cloud video viewing experience.
[0135] Example 3
[0136] Figure 3 This is a structural block diagram of a terminal device 2 provided in the third embodiment of this application. For example... Figure 3 As shown, the terminal device 2 in this embodiment includes: a processor 20, a memory 21, and a computer program 22 stored in the memory 21 and executable on the processor 20, such as a program for a cloud video collaborative processing method. When the processor 20 executes the computer program 22, it implements the steps in the various embodiments of the cloud video collaborative processing methods described above.
[0137] For example, the computer program 22 may be divided into one or more modules, which are stored in the memory 21 and executed by the processor 20 to complete this application. The one or more modules may be a series of computer program instruction segments capable of performing specific functions, which describe the execution process of the computer program 22 in the terminal device 2. The terminal device may include, but is not limited to, the processor 20 and the memory 21.
[0138] The processor 20 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.
[0139] The memory 21 can be an internal storage unit of the terminal device 2, such as a hard drive or memory of the terminal device 2. The memory 21 can also be an external storage device of the terminal device 2, such as a plug-in hard drive, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card equipped on the terminal device 2. Furthermore, the memory 21 can include both internal and external storage units of the terminal device 2. The memory 21 is used to store the computer program and other programs and data required by the terminal device. The memory 21 can also be used to temporarily store data that has been output or will be output.
[0140] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0141] If an integrated module is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. This computer-readable storage medium can be non-volatile or volatile. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable storage medium can include: any entity or device capable of carrying computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the contents of a computer-readable storage medium may be appropriately added to or subtracted from the contents as required by the legislation and patent practice in a jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, a computer-readable storage medium may not include electrical carrier signals and telecommunication signals.
[0142] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A cloud video collaborative processing method, characterized in that, The method includes: Obtain the cloud video to be processed uploaded by the user, and extract the audio data from the cloud video to be processed; The audio data is classified to obtain the dialogue speech, and the dialogue speech is then subjected to speech recognition to obtain the speech recognition result; Based on the speech recognition results, sensitive words in the dialogue are identified, and the video sentences corresponding to the sensitive words are identified as sensitive sentences. Semantic recognition is performed on the sensitive sentences to obtain the semantics of the sensitive sentences, and a sensitive vector is constructed based on the sensitive words in the speech and the semantics of the sensitive sentences; The replacement speech is determined based on the sensitivity vector, and the sensitive sentences are replaced in the cloud video to be processed based on the replacement speech to obtain the output cloud video. Determining the replacement speech based on the sensitivity vector includes: The sensitive vector is matched with the replacement database to obtain the replacement text. The replacement database stores the correspondence between different sensitive vectors and corresponding replacement texts. The replacement text is text with the same or similar semantics preset for the corresponding sensitive vector and without sensitive content. Extract the timbre features of the sensitive sentences, and input the timbre features and the replacement text into a pre-trained speech synthesis model to synthesize the replacement speech; After replacing the sensitive statements according to the replaced speech, the method further includes: The replacement text is segmented into words to obtain text segments. The text segments are then matched with a mouth opening and closing degree query database to obtain the mouth opening and closing degree. The mouth opening and closing degree query database stores the correspondence between different text segments and the corresponding mouth opening and closing degree. The process involves obtaining the image of the sensitive statement and determining the speaker in the image based on the statement identifier of the sensitive statement. The statement identifier is used to mark the image position of the speaker in the cloud video to be processed. In the statement screen, the speaker's mouth shape is modified according to the degree of mouth opening to obtain a replacement screen, and the statement screen is replaced according to the replacement screen; In the stated statement screen, the speaker's lip shape is modified according to the degree of lip opening to obtain a replacement screen, including: Obtain an image of the mouth region in the statement frame, and determine a mouth-related image based on the mouth region image; Vector transformation is performed on the image features of the mouth region image and the mouth-related image respectively to obtain image feature vector and related feature vector; The first matrix dimension is constructed based on the high-level semantic vector, the mid-level image vector, and the low-level appearance vector in the image feature vector; Calculate the feature association strength between the mouth region image and the mouth-related image, and construct a second matrix dimension based on the feature association strength; The third matrix dimension is determined based on the mouth opening degree, the temporal weight of the cloud video to be processed is obtained, and the fourth matrix dimension is constructed based on the temporal weight. A feature mapping matrix is constructed based on the first matrix dimension, the second matrix dimension, the third matrix dimension, and the fourth matrix dimension, and the feature mapping matrix is decoded to obtain the replacement features; A replacement image is generated based on the replacement features, and in the sentence screen, the mouth region image and the mouth-related image are replaced based on the replacement image to obtain the replacement screen.
2. The cloud video collaborative processing method as described in claim 1, characterized in that, Constructing a sensitivity vector based on the semantics of the speech-sensitive words and the sensitive sentences includes: Obtain the sensitivity type and sensitivity score of the speech sensitive words, and construct a sensitive word vector based on the sensitivity type and sensitivity score; Calculate the semantic similarity between the sensitive sentence semantics and the preset semantics, and determine the preset semantics corresponding to the maximum semantic similarity as the target semantics. The preset semantics is a pre-constructed semantic vector library, which contains a variety of common semantic categories and their corresponding vector representations. The semantic vector corresponding to the target semantic is determined as the sensitive semantic vector, and the sensitive semantic vector and the sensitive word vector are combined to obtain the sensitive vector.
3. The cloud video collaborative processing method as described in claim 1, characterized in that, The audio data is classified to obtain the dialogue audio, including: Zero-crossing detection is performed on the speech data, and silence filtering is performed on the speech data based on the zero-crossing detection results to obtain silence-filtered speech; The silence-filtered speech is subjected to frequency domain filtering to obtain frequency domain filtered speech, and the speech features of the frequency domain filtered speech are extracted. The speech features are classified to obtain the dialogue speech.
4. The cloud video collaborative processing method as described in claim 3, characterized in that, The silence-filtered speech is subjected to frequency domain filtering to obtain frequency domain filtered speech, including: The odd-numbered speech signals in the silence-filtered speech are flipped according to the first signal flipping function to obtain the odd-numbered flipped signal. The even-numbered speech signals in the silence-filtered speech are flipped according to the second signal flipping function to obtain even-numbered flipped signals. The odd-numbered flip signal and the even-numbered flip signal are combined to obtain the speech reconstruction signal, and the real part of the speech reconstruction signal is processed by fast Fourier transform to obtain the real part feature signal. The imaginary part of the reconstructed speech signal is subjected to discrete cosine transform to obtain the imaginary part feature signal, and the real part feature signal and the imaginary part feature signal are concatenated to obtain the concatenated feature signal; The spliced feature signal is subjected to an inverse real part fast Fourier transform to obtain an inverse transform feature signal, and the inverse transform feature signal is filtered according to a preset frequency domain signal to obtain the frequency domain filtered speech.
5. A cloud video collaborative processing system, characterized in that, The system includes: The voice extraction module is used to acquire cloud videos uploaded by users and extract voice data from the cloud videos. The speech recognition module is used to classify the speech data to obtain dialogue speech, and to perform speech recognition on the dialogue speech to obtain speech recognition results; The sensitivity detection module is used to determine the sensitive words in the dialogue based on the speech recognition results, and to identify the video sentences corresponding to the sensitive words as sensitive sentences. The vector construction module is used to perform semantic recognition on the sensitive sentences, obtain the semantics of the sensitive sentences, and construct sensitive vectors based on the sensitive words in the speech and the semantics of the sensitive sentences; The collaborative processing module is used to determine the replacement speech based on the sensitivity vector, and to replace the sensitive sentences in the cloud video to be processed based on the replacement speech to obtain the output cloud video; The collaborative processing module is further configured to: match the sensitive vector with the replacement database to obtain replacement text, wherein the replacement database stores the correspondence between different sensitive vectors and corresponding replacement texts, and the replacement text is text with the same or similar semantics preset for the corresponding sensitive vector and without sensitive content; Extract the timbre features of the sensitive sentences, and input the timbre features and the replacement text into a pre-trained speech synthesis model to synthesize the replacement speech; The collaborative processing module is further configured to: segment the replacement text to obtain text segments, match the text segments with the mouth opening and closing degree query database to obtain the mouth opening and closing degree, wherein the mouth opening and closing degree query database stores the correspondence between different text segments and the corresponding mouth opening and closing degree; The process involves obtaining the image of the sensitive statement and determining the speaker in the image based on the statement identifier of the sensitive statement. The statement identifier is used to mark the image position of the speaker in the cloud video to be processed. In the statement screen, the speaker's mouth shape is modified according to the degree of mouth opening to obtain a replacement screen, and the statement screen is replaced according to the replacement screen; The collaborative processing module is also used to: acquire a mouth region image in the statement screen, and determine a mouth-related image based on the mouth region image; Vector transformation is performed on the image features of the mouth region image and the mouth-related image respectively to obtain image feature vector and related feature vector; The first matrix dimension is constructed based on the high-level semantic vector, the mid-level image vector, and the low-level appearance vector in the image feature vector; Calculate the feature association strength between the mouth region image and the mouth-related image, and construct a second matrix dimension based on the feature association strength; The third matrix dimension is determined based on the mouth opening degree, the temporal weight of the cloud video to be processed is obtained, and the fourth matrix dimension is constructed based on the temporal weight. A feature mapping matrix is constructed based on the first matrix dimension, the second matrix dimension, the third matrix dimension, and the fourth matrix dimension, and the feature mapping matrix is decoded to obtain the replacement features; A replacement image is generated based on the replacement features, and in the sentence screen, the mouth region image and the mouth-related image are replaced based on the replacement image to obtain the replacement screen.
6. The cloud video collaborative processing system as described in claim 5, characterized in that, The vector construction module is also used for: Obtain the sensitivity type and sensitivity score of the speech sensitive words, and construct a sensitive word vector based on the sensitivity type and sensitivity score; Calculate the semantic similarity between the sensitive sentence semantics and the preset semantics, and determine the preset semantics corresponding to the maximum semantic similarity as the target semantics. The preset semantics is a pre-constructed semantic vector library, which contains a variety of common semantic categories and their corresponding vector representations. The semantic vector corresponding to the target semantic is determined as the sensitive semantic vector, and the sensitive semantic vector and the sensitive word vector are combined to obtain the sensitive vector.
7. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Method for identifying abnormal content in video stream and video stream processing system and method
CN108509827A
Audio sensitive information automatic shielding method and device, equipment and storage medium
CN116052635A