Conference role summary information generation method and device, electronic equipment and medium
By enhancing the audio and video of the meeting, identifying users, and performing emotion recognition, accurate and personalized summary information is generated, which solves the problem of low audio segmentation accuracy in existing technologies and improves meeting recording efficiency and user experience.
Patent Information
- Application Number
- CN202511152780.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-18
- Publication Date
- 2026-01-23
AI Technical Summary
Existing technologies suffer from low audio segmentation accuracy and redundant information when generating meeting role summary information, resulting in low accuracy of the generated personalized role summary information, which affects meeting recording efficiency and user experience.
By acquiring meeting audio and video recordings and participant information sets, audio and video enhancement processing is performed, participant facial information is identified, and combined with speech-to-text conversion and emotion recognition, a personalized summary information set for each role is generated and stored accordingly.
It improves the efficiency and accuracy of meeting minutes, reduces information redundancy and storage resource waste, and allows users to quickly and accurately understand the meeting content.
Smart Images

Figure CN121397178A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present disclosure relate to the technical field of computer technology, and particularly relate to a conference role summary information generation method and device, electronic equipment and medium. BACKGROUND
[0002] At present, conference speech-to-text technology is a technology for converting the speech of a conference participant into text in real time, which can help the conference participant quickly and accurately understand the conference content and reduce the problem of missing conference content. For the generation of conference role summary information, the commonly used method is as follows: extracting a voiceprint from conference audio to obtain user voiceprint information. Then, according to the user voiceprint information, the conference audio is segmented to obtain user role conference audio. Finally, speech recognition and text summarization are performed on the user role conference audio to obtain a role personalized summary information set.
[0003] However, in practice, it is found that when the above method is used to generate conference role summary information, the following technical problems often exist: since the audio is segmented only by the user voiceprint information, the influencing factors considered are relatively single, resulting in low accuracy of conference audio segmentation, and there is a large amount of redundant and repeated information and useless information in the conference audio. By converting the conference audio into text for summarization, the accuracy of the generated role personalized summary information is low, there is redundant and related recognition error summary information in the role personalized summary information, resulting in waste of storage resources and reduction of conference recording efficiency, affecting the quick review and understanding of conference content by conference participants who enter the conference later, and prolonging the time for users to understand the conference content.
[0004] The above information disclosed in this BACKGROUND section is only for the purpose of enhancing the understanding of the background of the present disclosure and, therefore, can include information that does not form the prior art known to those of ordinary skill in the art in the country. SUMMARY
[0005] The summary section is provided to introduce concepts briefly in a simplified form, which will be described in detail in the specific embodiments section below. The summary section is not intended to identify key or essential features of the claimed technology nor is it intended to be used to limit the scope of the claimed technology.
[0006] Some embodiments of the present disclosure propose a conference role summary information generation method, device, electronic equipment and medium to solve one or more of the technical problems mentioned in the background section above.
[0007] In a first aspect, some embodiments of the present disclosure provide a conference role summary information generation method, comprising: obtaining conference audio and video and a set of participant user information corresponding to the conference audio and video; performing audio and video enhancement processing on the conference audio and video to obtain enhanced conference audio and video; performing speech text conversion on the enhanced conference audio and video to obtain conference text information; performing participant user identification on the enhanced conference audio and video to obtain a set of participant user face information; performing user role matching identification on the enhanced conference audio and video according to the set of participant user face information, the set of participant user information, and the conference text information to obtain a set of role text paragraph information; performing audio and video segmentation on the enhanced conference audio and video according to the set of role text paragraph information to obtain a set of role conference audio and video; performing face and audio emotion identification on the set of role conference audio and video to obtain a set of role emotion information; generating a set of role personalized summary information of the set of role text paragraph information according to the set of role emotion information, and storing the set of role personalized summary information and the set of participant user information in association.
[0008] In a second aspect, some embodiments of the present disclosure provide a conference role summary information generation apparatus, comprising: an obtaining unit configured to obtain conference audio and video and a set of participant user information corresponding to the conference audio and video; an audio and video enhancement unit configured to perform audio and video enhancement processing on the conference audio and video to obtain enhanced conference audio and video; a speech text conversion unit configured to perform speech text conversion on the enhanced conference audio and video to obtain conference text information; a participant user identification unit configured to perform participant user identification on the enhanced conference audio and video to obtain a set of participant user face information; a user role matching identification unit configured to perform user role matching identification on the enhanced conference audio and video according to the set of participant user face information, the set of participant user information, and the conference text information to obtain a set of role text paragraph information; an audio and video segmentation unit configured to perform audio and video segmentation on the enhanced conference audio and video according to the set of role text paragraph information to obtain a set of role conference audio and video; a face and audio emotion identification unit configured to perform face and audio emotion identification on the set of role conference audio and video to obtain a set of role emotion information; and a generating unit configured to generate a set of role personalized summary information of the set of role text paragraph information according to the set of role emotion information, and store the set of role personalized summary information and the set of participant user information in association.
[0009] In a third aspect, some embodiments of the present disclosure provide an electronic device, comprising: one or more processors; a storage device having one or more programs stored thereon, when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any of the implementations of the first aspect.
[0010] Fourthly, some embodiments of this disclosure provide a computer-readable medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method as described in any implementation of the first aspect.
[0011] The above various embodiments of the present disclosure have the following beneficial effects: the conference role summary information generation method of some embodiments of the present disclosure can realize real-time voice-to-text, improve conference recording efficiency, reduce information redundancy and omission, generate accurate personalized summary information, and reduce waste of storage resources. Specifically, the reasons for wasting relevant storage resources and reducing conference recording efficiency, affecting users' quick conference and understanding of conference content, and reducing user experience are as follows: since audio segmentation is only performed by user voiceprint information, the influencing factors considered are relatively single, resulting in low accuracy of conference audio segmentation, and there is a large amount of redundant repeated information and useless information in the conference audio, only converting the conference audio into text for summary generation, resulting in low accuracy of the generated role personalized summary information, and there is redundant and associated recognition error summary information in the role personalized summary information, resulting in waste of storage resources and reduction of conference recording efficiency, affecting the quick review and understanding of conference content by users who enter the conference later, and prolonging the time for users to understand the conference content. Based on this, the conference role summary information generation method of some embodiments of the present disclosure can first obtain conference audio and video and a set of conference user information corresponding to the conference audio and video. Here, it is used for subsequent user role recognition and audio segmentation. Second, the conference audio and video is subjected to audio and video enhancement processing to obtain enhanced conference audio and video. Here, the quality of the conference audio and video can be improved, noise can be removed, and the data volume of the conference audio and video can be reduced. Third, the enhanced conference audio and video is subjected to speech-to-text conversion to obtain conference text information. Here, the accuracy of speech-to-text conversion can be improved, and subsequent user role matching recognition can be facilitated. Next, the enhanced conference audio and video is subjected to conference user recognition to obtain a set of conference user face information. Here, the accuracy of parameter user face information recognition can be improved, and the conference user face information can be used as auxiliary information to improve the accuracy of user role matching recognition. Subsequently, the enhanced conference audio and video is subjected to user role matching recognition according to the set of conference user face information, the set of conference user information, and the conference text information, to obtain a set of role text paragraph information. Here, user role matching recognition is performed through multi-modal data, which can improve the accuracy of user role matching recognition. Then, the enhanced conference audio and video is subjected to audio and video segmentation according to the set of role text paragraph information, to obtain a set of role conference audio and video. Here, the accuracy of audio and video segmentation can be improved. Then, the set of role conference audio and video is subjected to face and audio emotion recognition to obtain a set of role emotion information. Here, through the fusion of face and audio emotion recognition, the accuracy and comprehensiveness of emotion recognition can be improved, and the conference state information of the user can be accurately obtained. Finally, a set of role personalized summary information of the set of role text paragraph information is generated according to the set of role emotion information, and the set of role personalized summary information and the set of conference user information are associated and stored.Here, the accuracy and efficiency of the role personalized summary information can be improved, the efficiency of the conference record can be improved, the conference content can be traced back by the user, the user can quickly and accurately understand the conference content and the user trace back, and the user experience can be improved and the waste of storage resources can be reduced. Therefore, the conference role summary information generation method can realize real-time voice-to-text, improve the efficiency of conference recording, reduce information redundancy and omission, generate accurate personalized summary information, and reduce the waste of storage resources. BRIEF DESCRIPTION OF DRAWINGS
[0012] The above and other features, advantages, and aspects of embodiments of the present disclosure will become more apparent by describing in detail some embodiments thereof with reference to the attached drawings. The same or similar components have the same or similar reference numbers throughout the drawings. It should be understood that the drawings are schematic and elements and features are not necessarily drawn to scale.
[0013] Figure 1 is a flowchart of some embodiments of the conference role summary information generation method according to the present disclosure;
[0014] Figure 2 is a structural schematic diagram of some embodiments of the conference role summary information generation apparatus according to the present disclosure;
[0015] Figure 3 is a structural schematic diagram of an electronic device suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION
[0016] Embodiments of the present disclosure will be described in detail below with reference to the drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms, and should not be interpreted as being limited to the embodiments set forth herein. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes, and are not intended to limit the scope of protection of the present disclosure.
[0017] In addition, it should be further noted that only parts related to the present invention are shown in the drawings for ease of description. The embodiments in the present disclosure and the features in the embodiments can be combined with each other without conflict.
[0018] It should be noted that the concepts of "first", "second", etc. mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not intended to limit the order or interdependence of the functions performed by these devices, modules or units.
[0019] It should be noted that the modification of "one", "multiple" mentioned in the present disclosure is illustrative but not restrictive, and those skilled in the art should understand that unless otherwise explicitly indicated in the context, it should be understood as "one or more".
[0020] The names of the messages or information exchanged between the plurality of devices in the embodiments of the present disclosure are only for illustrative purposes, and are not intended to limit the scope of the messages or information.
[0021] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.
[0022] Figure 1 Flow 100 of some embodiments of a conference role summary information generation method according to the present disclosure is shown. The conference role summary information generation method includes the following steps:
[0023] Step 101, obtaining conference audio and video and conference audio and video corresponding participant user information set.
[0024] In some embodiments, the execution subject (such as an electronic device) of the conference role summary information generation method described above can obtain the conference audio and video and the participant user information set corresponding to the conference audio and video through wired connection or wireless connection. Wherein, the conference audio and video can be the audio and video recorded during the online conference. The participant user information in the participant user information set can be the information of the user participating in the online conference.
[0025] Step 102, performing audio and video enhancement processing on the conference audio and video to obtain enhanced conference audio and video.
[0026] In some embodiments, the execution subject can perform audio and video enhancement processing on the conference audio and video to obtain enhanced conference audio and video. Wherein, the enhanced conference audio and video can be audio and video that denoises the noise existing in the conference audio in the conference audio and video and enhances the quality of the video frames in the conference video.
[0027] In some optional implementations of some embodiments, the audio and video enhancement processing on the conference audio and video to obtain the enhanced conference audio and video can include the following steps:
[0028] First, extract the conference audio from the conference audio and video to obtain the conference audio.
[0029] Secondly, a short-time Fourier transform is performed on the conference audio to obtain a conference speech spectrogram. The conference speech spectrogram can be a spectrogram representing the conference audio in a time-frequency domain. The conference speech spectrogram can be denoted as F*T*2. T represents a time step, F represents a frequency band number, and 2 represents that a complex spectrum of the short-time Fourier transform is decomposed into two real components, i.e., a real part and an imaginary part of each video point, also known as a channel dimension.
[0030] Thirdly, amplitude masking feature extraction and phase feature extraction are performed on the conference speech spectrogram in parallel to obtain a speech amplitude feature map set and a speech phase feature map set. The speech amplitude feature map in the speech amplitude feature map set can represent amplitude information of the conference audio. The speech phase feature map in the speech phase feature map set can represent phase information of the conference audio. The speech amplitude feature map and the speech phase feature map have the same size as the conference speech spectrogram, i.e., T*F, but have different channel dimensions. In practice, the execution subject can input the conference speech spectrogram into convolution layers with convolution kernels of 1*7 and 7*1 in sequence to obtain the speech amplitude feature map set. The conference speech spectrogram is input into convolution layers with convolution kernels of 5*3 and 25*1 in parallel in sequence to obtain the speech phase feature map set.
[0031] Fourthly, speech noise reduction is performed on the speech amplitude feature map set and the speech phase feature map set to obtain a denoised conference speech spectrogram. The speech noise reduction can be speech denoising using a spectral subtraction method.
[0032] Fifthly, spectral speech enhancement is performed on the denoised conference speech spectrogram to obtain an enhanced conference speech spectrogram. The spectral speech enhancement can be further denoising and enhancement of residual noise in the denoised conference speech using a spectral subtraction method.
[0033] Sixthly, video enhancement is performed on conference video in the conference audio-video to obtain an enhanced conference video. The video enhancement can be video enhancement processing using a wavelet transform algorithm.
[0034] Seventhly, the enhanced conference speech spectrogram and the enhanced conference video are determined as an enhanced conference audio-video. In practice, the execution subject can first perform inverse transformation on the enhanced conference speech spectrogram to obtain an enhanced conference audio in a time domain. Then, the enhanced conference audio and the enhanced conference video are determined as the enhanced conference audio-video.
[0035] Optionally, the speech noise reduction on the speech amplitude feature map set and the speech phase feature map set to obtain the denoised conference speech can include the following steps:
[0036] In a first step, based on the speech amplitude feature map set and the speech phase feature map set, the following speech noise reduction steps are performed:
[0037] In sub-step 1, the speech amplitude feature map set is input into a first speech frequency transformation network to obtain a first speech harmonic amplitude feature map set. The first speech harmonic amplitude feature map in the first speech harmonic amplitude feature map set can be extracted from the harmonic in the speech amplitude feature map to capture the global correlation along the frequency axis.
[0038] In sub-step 2, the first speech harmonic amplitude feature map set is sequentially input into a plurality of amplitude convolution feature extraction layers to obtain a second speech harmonic amplitude feature map set. The second speech harmonic amplitude feature map in the second speech harmonic amplitude feature map set can represent the local video correlation of the speech harmonic amplitude feature map. The amplitude convolution feature extraction layer in the plurality of amplitude convolution feature extraction layers can include a convolution layer with a convolution kernel of 5*5, 25*1 and 5*5.
[0039] In sub-step 3, the second speech harmonic amplitude feature map set is input into a second speech frequency transformation network to obtain a third speech harmonic amplitude feature map set. The third speech harmonic amplitude feature map in the third speech harmonic amplitude feature map set can represent the global correlation of the conference audio along the frequency axis and the high-level feature information of the amplitude prediction of the conference audio. The second speech frequency transformation network can be a transformation network with the same network structure as the first speech frequency transformation network, but different parameters and input / output feature sizes of each network layer.
[0040] In sub-step 4, the speech phase feature map set is sequentially input into a plurality of phase convolution feature extraction networks to obtain a speech phase mask feature map set. The speech phase mask feature map in the speech phase mask feature map set can represent the long-range time domain correlation of the speech phase feature map. The plurality of phase convolution feature extraction networks can include convolution layers with a convolution kernel of 5*3 and 25*1, and a global normalization layer before each convolution layer.
[0041] Sub-step 5, performing a double-flow feature exchange processing on the third speech harmonic amplitude feature map set and the speech phase mask feature map set to obtain a speech amplitude exchange feature map set and a speech phase exchange feature map set. The speech amplitude exchange feature map in the speech amplitude exchange feature map set can represent information obtained by interaction to assist in predicting the speech amplitude. The speech phase exchange feature map in the speech phase exchange feature map set can represent information obtained by interaction to assist in predicting the phase. In practice, the execution subject can first input the third speech harmonic amplitude feature map set and the speech phase mask feature map set into a convolution layer with a convolution kernel of 1*1 to adjust the channel to obtain a speech harmonic amplitude channel feature map set and a speech phase mask channel feature map set. The channel number of the speech harmonic amplitude channel feature map set is the same as that of the speech phase mask feature map set. The channel number of the speech phase mask channel feature map set is the same as that of the third speech harmonic amplitude feature map set. Then, the speech harmonic amplitude channel feature map set and the speech phase mask channel feature map set are input into a Tanh (Hyperbolic Tangent, hyperbolic tangent function) activation function to obtain a speech harmonic amplitude normalized feature map set and a speech phase mask normalized feature map set. Finally, the speech harmonic amplitude normalized feature map set and the speech phase mask feature map set are multiplied element by element to obtain the speech amplitude exchange feature map set, and the third speech harmonic amplitude feature map set and the speech phase mask normalized feature map set are multiplied element by element to obtain the speech phase exchange feature map set.
[0042] Sub-step 6, in response to determining that the number of times of execution of the speech noise reduction step is greater than or equal to a preset execution number threshold, performing fusion denoising on the speech amplitude exchange feature map set and the speech phase exchange feature map set to obtain a denoised conference speech spectrum map. The preset execution number can be a maximum value of the preset execution number. For example, the preset execution number threshold can be 3.
[0043] In practice, the execution subject can first input the voice amplitude exchange feature map set into a convolution layer with a convolution kernel of 1*1, a bidirectional long short-term memory neural network, and three full connection layers in sequence to obtain a one-dimensional amplitude feature map set. The convolution layer with a convolution kernel of 1*1 reduces the channel number of the voice amplitude exchange feature map set to 8 and reshapes it into a one-dimensional feature map. The dimension of the feature map output by the convolution layer is T*(F*8). The rated activation function of the first two full connection layers in the three full connection layers is ReLU (Rectified Linear Unit), and the activation function of the last full connection layer is a Sigmoid activation function. Then, the voice phase exchange feature map set is input into a convolution layer with a convolution kernel of 1*1 to reduce the channel number to 2 to obtain a phase complex value feature map. After that, the amplitude features included in the phase complex value feature map are normalized to obtain a phase prediction feature map set containing only phase information. Finally, the absolute value of the conference voice spectrum graph, the one-dimensional amplitude feature map set, and the phase prediction feature map set are multiplied element by element to obtain a denoised conference voice spectrum graph.
[0044] In the second step, in response to determining that the number of times of execution is less than the preset threshold of the number of times of execution, the voice amplitude exchange feature map set and the voice phase exchange feature map set are determined as a voice amplitude feature map set and a voice phase feature map set, respectively, and the sum of the number of times of execution and a preset value is determined as the number of times of execution, so as to execute the voice denoising step again. The preset value can be a preset value. For example, the preset value can be 1.
[0045] Optionally, the inputting of the voice amplitude feature map set into the first voice frequency transformation network to obtain a first voice harmonic amplitude feature map set can include the following steps:
[0046] In the first step, the voice amplitude feature map set is input into a channel convolution network to obtain a voice amplitude channel feature map set. The voice amplitude channel feature map in the voice amplitude channel feature map set is a feature map with a channel number of 5. The channel convolution network can be a network for channel dimension reduction, including a convolution layer with a convolution kernel of 1*1, a batch normalization layer, and a ReLU activation function.
[0047] In the second step, the voice amplitude channel feature map set is subjected to feature dimension conversion to obtain a voice amplitude channel feature map set after dimension conversion. The feature dimension conversion can be a one-dimensional feature dimension conversion.
[0048] In the third step, the dimensionally converted speech amplitude channel feature map set is input into a one-dimensional convolutional network to obtain a speech amplitude attention feature map. The one-dimensional convolutional network can include a convolutional layer with a convolution kernel of 9, a batch normalization layer, and a ReLU activation function.
[0049] In the fourth step, the speech amplitude attention feature map and the speech amplitude feature map set are multiplied point by point to obtain a semantic amplitude weight feature map set.
[0050] In the fifth step, the semantic amplitude weight feature map set is mapped and sliced to obtain a speech amplitude frequency transformation feature map set. The speech amplitude frequency transformation feature map in the speech amplitude frequency transformation feature map set can be a feature map containing information of the frequency band of the semantic amplitude weight feature map set. In practice, the execution subject can use a fully connected layer to perform a linear transformation on the input semantic amplitude weight feature map set along the time axis and combine and map the features of different frequency points, thereby capturing long-range correlations in the frequency domain. The time slice can be F*C. C can represent the number of channels.
[0051] In the sixth step, each speech amplitude frequency transformation feature map in the speech amplitude frequency transformation feature map set is feature-mapped with the corresponding speech amplitude feature map in the speech amplitude feature map set to obtain a spliced speech amplitude feature map set.
[0052] In the seventh step, the spliced speech amplitude feature map set is subjected to channel dimension reduction to obtain a first speech harmonic amplitude feature map set. The channel dimension reduction can be channel reduction by sequentially inputting into a 1*1 convolutional layer, a batch normalization layer, and a ReLU activation function.
[0053] In step 103, the enhanced conference audio and video are subjected to speech-to-text conversion to obtain conference text information.
[0054] In some embodiments, the execution subject can perform speech-to-text conversion on the enhanced conference audio and video to obtain conference text information. The conference text information can be information about the audio in the enhanced conference audio and video displayed in the form of text.
[0055] In some optional implementations of some embodiments, the speech-to-text conversion on the enhanced conference audio and video to obtain conference text information can include the following steps:
[0056] In the first step, the enhanced conference audio is framed and windowed to obtain a short-time energy set and a short-time zero-crossing rate set. The short-time energy in the short-time energy set represents the energy distribution of the amplitude of the enhanced conference audio in time. The short-time zero-crossing rate in the short-time zero-crossing rate set represents the number of zero-crossing values of the enhanced conference audio. The short-time energy can be used to distinguish vowels and consonants in the audio. The short-time zero-crossing rate can be used to distinguish voiced and unvoiced sounds in the audio.
[0057] In the second step, the short-time energy set and the short-time zero-crossing rate set are used to perform adaptive time-domain feature segmentation on the enhanced conference audio to obtain a conference audio segmented syllable sequence. The conference audio segmented syllable in the conference audio segmented syllable sequence is a segmented syllable. It should be noted that adaptive time-domain feature segmentation can accurately divide the ambiguous boundaries between syllables caused by silent segments (silence, noise).
[0058] As an example, the execution subject can first determine the product of 0.25, a vowel control factor, and the mean of the set of short-time energies of the speech as the maximum short-time energy. The vowel control factor can be a factor for controlling the energy coverage range of vowels. The value range of the vowel control factor can be [0.8, 1]. Second, determine the sum of the mean of the set of short-time energies of the speech and the product of 0.2 and the maximum short-time energy, and the product of the consonant energy control factor, as the minimum short-time energy. The consonant energy control factor can be a factor for controlling the energy coverage range of consonants. The value range of the consonant energy control factor can be [0.3, 0.5]. Third, determine the product of the voicing zero-crossing rate control factor and the mean of the set of short-time zero-crossing rates of the speech as the minimum zero-crossing rate. The voicing zero-crossing rate control factor can be a factor for controlling the range determined as voiced and unvoiced. The value range of the voicing zero-crossing rate control factor can be [0.6, 0.8]. Subsequently, determine the speech frames greater than or equal to the maximum short-time energy as vowel speech frames, determine the speech frames greater than or equal to the minimum short-time energy and less than the maximum short-time energy as consonant frames, determine the speech frames greater than or equal to the minimum zero-crossing rate as unvoiced frames, and determine the speech frames less than the minimum zero-crossing rate as voiced frames, and superimpose the vowel frames, the consonant frames, the unvoiced frames, and the voiced frames to obtain a candidate syllable sequence. The candidate syllable sequence can include unvoiced consonants, voiced vowels, unvoiced vowels, and voiced consonants. Next, generate an adaptive speech segmentation fitness segmentation function. The fitness segmentation function can be the minimum value of the square of the difference between the artificially labeled partial real boundary segments and the predicted boundary segments. Then, using an evolutionary algorithm combined with an elitist strategy, the candidate syllable sequence is used as the initial population to perform adaptive time-domain feature segmentation on the enhanced conference audio to obtain a conference speech segmented syllable sequence.
[0059] Third, perform Mel-frequency cepstrum recognition on the conference speech segmented syllable sequence to obtain a set of speech syllable feature maps. The speech syllable feature maps in the set of speech syllable feature maps can represent Mel-frequency cepstrum coefficients.
[0060] Fourth, perform convolution feature extraction on the set of speech syllable feature maps to obtain a set of speech syllable convolution feature maps. The convolution feature extraction can be performed using eight convolution networks. The convolution network can include a convolution layer with a convolution kernel of 5*5, a max-pooling layer with a pooling kernel of 2*2, and a dropout layer.
[0061] In the fifth step, the regional difference phoneme feature set is extracted from the phoneme convolution feature set, and a phoneme regional difference feature set is obtained. The phoneme regional difference feature in the phoneme regional difference feature set can represent the difference information of the phonemes in different regions. In practice, the execution subject can use the multi-head attention mechanism layer to extract the regional difference phoneme feature from the phoneme convolution feature set, and obtain the phoneme regional difference feature set. It should be noted that in the multi-head attention mechanism, multiple group attention weights are used, each group of attention weights can learn different semantic information, and each group of attention weights will generate a context vector. Then, the obtained multiple context vectors are spliced, and a linear transformation is performed to obtain the phoneme regional difference feature set. For the phoneme pronunciation difference of different dialects in the enhanced conference audio and video, the multi-head attention mechanism uses the query vector, the key vector and the value vector to represent different subspaces, and connects the feature position information of different phonemes, enriching the phonetic feature diversity and phoneme regional difference.
[0062] In the sixth step, the spectrum of the phoneme regional difference feature set is enhanced, and a continuous speech local feature set is obtained. The continuous speech local feature in the continuous speech local feature set can represent the characters and tones in the conference audio. The spectrum enhancement can be spectrum enhancement through frequency and time mask.
[0063] In the seventh step, the continuous speech local feature set is subjected to convolution downsampling processing, and a down-sampled speech local feature set is obtained.
[0064] In the eighth step, the down-sampled speech local feature set is subjected to linear regularization processing, and a regularized speech local feature set is obtained. The linear regularization processing can be a processing of inputting to a linear layer and then inputting to a regularization layer.
[0065] In the ninth step, the regularized speech local feature map set is input into a grapheme syllable dependent recognition model to obtain conference text information. The grapheme syllable dependent recognition model can be a model for recognizing the position relationship and long sequence relationship between syllables in the extracted regularized speech local feature map to realize the recognition of continuous speech in the conference audio. The grapheme syllable dependent recognition model can include a first feedforward network, a multi-head self-attention network, a syllable convolution network, a second feedforward network, and a post-normalization layer. The first feedforward network can include a post-normalization layer, a first linear layer, a Swish activation layer, a first regularization layer, a second linear layer, a second regularization layer, and a pre-norm residual unit network. The pre-norm residual unit can be a residual unit between activation functions. The second feedforward network and the first feedforward network have the same network structure, but different network parameters and input and output feature sizes. The multi-head self-attention network can include a post-normalization layer, a multi-head attention relative position embedding based on a Transformer-XL (Attention Language Model Beyond Fixed-Length Context) model, and a regularization layer network. The syllable convolution network can include a post-normalization layer, a point convolution, a linear gating mechanism layer, a one-dimensional deep convolution layer, a batch normalization layer, a Swish activation function, and a regularization layer. The linear gating mechanism layer can include a gating mechanism layer composed of a point-by-point convolution and a linear gating unit. The one-dimensional deep convolution layer can be a convolution layer on sequence data, using a separate convolution on each channel and performing element-wise convolution operations.
[0066] In step 104, the enhanced conference audio and video are subjected to participant user recognition to obtain a set of participant user face information.
[0067] In some embodiments, the execution subject can perform participant user recognition on the enhanced conference audio and video to obtain a set of participant user face information. The participant user face information in the set of participant user face information can be used to identify the identity information of the participant user. The participant user face information can include, but is not limited to, at least one of the following: facial feature position information and contour information, facial feature proportion information, and skin texture information. The participant user recognition can be user recognition by a face recognition algorithm.
[0068] In step 105, according to the set of participant user face information, the set of participant user information, and the conference text information, user role matching recognition is performed on the enhanced conference audio and video to obtain a set of role text paragraph information.
[0069] In some embodiments, the execution subject can perform user role matching recognition on the enhanced conference audio and video according to the set of conference participant user facial information, the set of conference participant user information, and the conference text information, to obtain a set of role text paragraph information. The role text paragraph information in the set of role text paragraph information can be text information of conference speech of the conference participant user after irrelevant audio to the field of the conference audio and video is removed.
[0070] In the process of solving the technical problems mentioned in the background by adopting the technical solutions, the following technical problems often occur: due to the mixing of user identity information and text content information in the conference audio and video, it is difficult to accurately extract user voiceprint information from the conference audio and video, and the text content information has the problem of unstable clarity and consistency, resulting in low accuracy of user role recognition of the conference audio and video, a large amount of error redundancy data in the role text paragraph information, low data quality of the role text paragraph information, and waste of a large amount of storage resources. In view of the above technical problems, the conventional solution is generally as follows: by variational auto-encoding, the enhanced conference audio and video is disentangled to obtain user voiceprint information and text content information. Then, by using the user voiceprint information and the text content information, the enhanced conference audio and video is recognized for user role matching to obtain role text paragraph information. However, the above conventional solution still has the following problems: since the disentanglement of user identity information and text content information by variational auto-encoding depends on the division of the latent space determined by expert experience, and the training complexity of variational auto-encoding is high and the convergence speed is slow, the conference audio and video cannot be recognized for real-time user role matching, resulting in a long recognition time for user role matching, which is not suitable for real-time recognition scenarios of conference audio and video. Considering the shortcomings of the above conventional solution and combining the advantages of the conference speaker recognition technology possessed by the company where the inventors work, the inventors have decided to adopt the following solution:
[0071] In some optional implementations of some embodiments, the user role matching recognition on the enhanced conference audio and video according to the set of conference participant user facial information, the set of conference participant user information, and the conference text information to obtain a set of role text paragraph information can include the following steps:
[0072] First, a set of speech mel-spectrograms for the enhanced conference audio and video is generated. The speech mel-spectrogram in the set of speech mel-spectrograms can be a frequency spectrum graph with a mel frequency scale. The mel frequency scale can be a frequency scale simulating human auditory characteristics. In practice, the generation can be performed by an optimization library. The optimization library can include but is not limited to at least one of the following: Librosa library, TorchAudio library.
[0073] Secondly, input the above set of speech mel-spectrogram into the user voiceprint extraction coding network included in the trained speech-text disentanglement model to obtain a set of user voiceprint embedding vectors, wherein the trained speech-text disentanglement model further includes a speech-text coding network, a speech-text disentanglement decoding model and a user identification classifier. The user voiceprint extraction coding network includes a plurality of speech voiceprint convolutional networks, a plurality of voiceprint feature dimension weighting networks, a weighted feature fusion layer, a statistical pooling layer and a fully connected layer. The speech-text disentanglement model can be a deep neural network model that disentangles the speech content and user voiceprint features of the input speech mel-spectrogram, identifies the user through the separated user voiceprint features, and outputs the participant user identification information. The user voiceprint extraction coding network can be a deep neural network model for extracting user-specific acoustic features in the input set of speech mel-spectrogram. The plurality of speech voiceprint convolutional networks and the plurality of voiceprint feature dimension weighting networks have a one-to-one correspondence, i.e. the voiceprint feature dimension weighting network is located after the speech voiceprint convolutional network. The speech voiceprint convolutional network can be a network including a convolutional layer, a ReLU activation function and a batch normalization layer connected in series. The voiceprint feature dimension weighting network can be a deep neural network that learns the attention weight of the feature dimension of the feature vector output by the speech voiceprint convolutional network. The voiceprint feature dimension weighting network can be a deep neural network that first inputs the voiceprint feature vector output by the speech voiceprint convolutional network into a global average pooling layer to calculate the attention weight of each dimension, a first fully connected layer including a ReLU activation function to enhance the non-linear ability of the feature, a second fully connected layer including a Sigmoid activation function to learn the weight of the dimension to obtain a set of dimension weight values, then performs weighted summation on the set of dimension weight values and the voiceprint feature vector to obtain a user embedding feature vector and output. The weighted feature fusion layer can be a deep neural network that aggregates the four user embedding feature vectors output by the first four speech voiceprint convolutional networks and the first four voiceprint feature dimension weighting networks. The feature vector output by the weighted feature fusion layer is input again into the fifth speech voiceprint convolutional network and the fifth voiceprint feature dimension weighting network. The statistical pooling layer can be a pooling layer that obtains a speech-level feature vector by statistically pooling the mean and variance of the feature vector output by the fifth voiceprint feature dimension weighting network.
[0074] The voice-text disentanglement model is trained by the following steps: a text content reconstruction criterion and an identity invariant criterion. The text content reconstruction criterion can guide the voice-text disentanglement model to reconstruct the audio content of the input enhanced post-meeting audio-video in the training stage, so that the voice-text disentanglement model learns and extracts the voiceprint information of the participating user, and at the same time learns the strategy of separating the voiceprint features and speech content information of the participating user. The text content reconstruction criterion can make the voice-text disentanglement model extract the first speech mel-spectrogram of the first meeting audio and the second speech mel-spectrogram of the second meeting audio of the same participating user at the same time, extract the first speech semantic feature vector of the first speech mel-spectrogram, extract the second user voiceprint embedding vector of the second speech mel-spectrogram, then input the first speech semantic feature vector and the second user voiceprint embedding vector into the voice-text disentanglement decoding model to obtain the reconstructed first and second speech mel-spectrograms and input them into the voice-text encoding network to obtain the fused speech semantic feature vector, so as to minimize the error criterion of the first speech semantic feature vector and the fused speech semantic feature vector. The identity invariant criterion can be a criterion for reducing the intra-class distance of the same participating user between different videos. The joint reconstruction loss function of the voice-text disentanglement model can be a loss function composed of the reconstruction loss function part of the voice-text disentanglement decoding model, the KL (Kullback Leibler) divergence loss function part of the voice-text encoding network, the cross-entropy loss function part of the voice-text disentanglement model, the text content reconstruction loss function part of the text content reconstruction criterion, the loss function for measuring the intra-class variance of the same participating user under different audio, i.e. the participating user identity loss function part, and the weight values of each loss function part.
[0075] In the third step, the set of voice mel-spectrograms is input into the trained voice text encoding network to obtain a set of voice semantic feature vectors. The voice text encoding network includes a convolution filter set and a plurality of voice text convolution residual networks. The user voiceprint extraction encoding network and the voice text encoding network are executed in parallel. The voice text encoding network can be a deep neural network that separates the voice content and user identity information in the set of input voice mel-spectrograms. The convolution filter set can be a filter for capturing long-time information of the input set of voice mel-spectrograms. The convolution filter set can be a deep neural network that includes eight one-dimensional convolution filters of different scales, convolves the input set of voice mel-spectrograms, concatenates the output, and inputs the concatenated output into a 1*1 one-dimensional convolution layer. The one-dimensional convolution can be a convolution layer in which the convolution kernel gradually increases from 1 to 8, and each convolution output represents the information of the frequency spectrum at different time scales. The plurality of voice text convolution residual networks can be three voice text convolution residual networks. The voice text convolution residual network can be a network that includes the outputs of two convolution networks connected in series, concatenates the features output by the convolution filter set, obtains concatenated feature vectors, and inputs the concatenated feature vectors into two convolution networks with average pooling structures for downsampling and concatenation. The convolution network can be a deep neural network that includes convolution layers, ReLU activation functions, and target instance normalization layers connected in series. The target instance normalization layer can be an instance normalization function that does not involve affine transformation, can preserve the user's sentence text information, and can also normalize the user identity information to decouple and separate the voice text content and user identity information.
[0076] Fourthly, input the user voiceprint embedding vector set and the speech semantic feature vector set into the trained speech text disentanglement decoding model to obtain a reconstructed speech mel-spectrogram set. The speech text disentanglement decoding model can be a deep neural network that reconstructs the user voiceprint embedding vector set through the speech semantic feature vector set. The speech text disentanglement decoding model can be a deep neural network that inputs the input speech semantic feature vector set and the user voiceprint embedding vector set into a first disentanglement network to obtain a disentangled feature vector, splices the disentangled feature vector with the speech semantic feature vector set, inputs the user voiceprint embedding vector set into a second disentanglement network for up-sampling processing, splices the user voiceprint embedding vector set with the speech semantic feature vector set, and performs three times to obtain the reconstructed speech mel-spectrogram set. The first disentanglement network can include two convolutional layers connected in series, a ReLU activation function, and an adaptive instance normalization (AdaIN) layer. The second disentanglement network can be a network that adds a pixel reorganization layer (PixelShuffle) to restore the dimension of the feature up-sampled in the first disentanglement network before the second adaptive instance normalization layer. The user voiceprint embedding vector set can be an embedding vector set that is converted through an affine layer and participates in operation at each adaptive instance normalization layer in the speech text disentanglement decoding model.
[0077] Fifthly, input the reconstructed speech mel-spectrogram set into the trained user identification classifier to obtain conference participant user identification information. The user identification classifier can be a classifier that identifies the identity of the input reconstructed speech mel-spectrogram set. The user identification classifier can be a deep neural network including a linear layer and a Softmax layer connected in series. The conference participant user identification information can be information representing the identity of the conference participant.
[0078] Sixthly, according to the conference participant user identification information, perform user role segmentation on the enhanced conference audio included in the enhanced conference audio-video to obtain a segmented audio information set. The segmented audio information in the segmented audio information set can be information of all audio of the same conference participant in the enhanced conference audio-video.
[0079] Seventhly, perform keyword extraction on the conference text information to obtain a conference keyword set. The keyword extraction and the user role segmentation are performed in parallel. The conference keyword in the conference keyword set can be a text token representing the overall semantic information of the conference text information.
[0080] In the eighth step, the conference text information is clustered according to the conference keyword set and the participant information set to obtain a text clustering paragraph set. The text clustering paragraph in the text clustering paragraph set can be a text paragraph of all speeches of a participant in the conference.
[0081] As an example, the execution subject can first perform keyword extraction on each conference text sentence in the conference text sentence set included in the conference text information to obtain a sentence keyword group set. Second, the same keyword proportion value of each sentence keyword group in the sentence keyword group set and the conference keyword set is determined to obtain a keyword proportion value set. Then, the conference text sentences corresponding to at least one keyword proportion value greater than or equal to a preset proportion threshold value are filtered from the keyword proportion value set to obtain a target conference text sentence set. The preset proportion threshold value can be a preset value. For example, the preset proportion threshold value can be 0.5. Finally, the target conference text sentence set is clustered by the participant information set to obtain a text clustering paragraph set.
[0082] In the ninth step, the lip feature information in the set of face information of the participant, the set of segmented audio information, and the set of text clustering paragraphs are fused and adjusted to obtain a set of role text paragraph information, and the set of role text paragraph information and the set of participant information are stored in association. The lip feature information can represent whether the participant is speaking. In practice, the execution subject can first correct the set of segmented audio information based on the lip feature information to obtain a set of corrected audio information. Each piece of corrected audio information in the set of corrected audio information can be the audio information of the participant corresponding to the same audio determined based on the lip feature information. The participant corresponding to the same audio determined based on the user voiceprint feature recognition can be determined. If the determined participants are different, the participant determined based on the lip feature information is determined as the final participant audio information of the same audio. Then, the set of corrected audio information is subjected to text recognition to obtain a set of corrected conference text information. Then, at least one text clustering paragraph with the same text statement but different participants in the set of text clustering paragraphs is selected from the set of text clustering paragraphs. Then, the participant information set of the corrected conference text information corresponding to the at least one text clustering paragraph is determined as the target participant information set corresponding to the at least one text clustering paragraph to obtain a set of modified text clustering paragraphs. Finally, the set of modified text clustering paragraphs and the target text clustering paragraph set are fused and adjusted based on the participant information set corresponding to the set of modified text clustering paragraphs and the target text clustering paragraph set to obtain a set of role text paragraph information, and the set of role text paragraph information and the set of participant information are stored in association. The target text clustering paragraph set can be a set of paragraphs with the same text and participant information in the set of text clustering paragraphs and the set of corrected conference text information.
[0083] The technical solution and related content thereof serve as one of the invention points of the embodiments of the present disclosure, and solve the technical problem mentioned in the background art, i.e., due to the mixing of user identity information and text content information in the conference audio and video, it is difficult to accurately extract user voiceprint information from the conference audio and video, and the text content information has the problems of instability in clarity and consistency, resulting in low accuracy of user role recognition on the conference audio and video, a large amount of error redundancy data in the role text paragraph information, low data quality of the role text paragraph information, and waste of a large amount of storage resources. The factors leading to a large amount of error redundancy in the role text paragraph information set and low content quality are often as follows: due to the mixing of user identity information and text content information in the conference audio and video, it is difficult to accurately extract user voiceprint information from the conference audio and video, and the text content information has the problems of instability in clarity and consistency, resulting in low accuracy of user role recognition on the conference audio and video, a large amount of error redundancy data in the role text paragraph information, low data quality of the role text paragraph information, and waste of a large amount of storage resources. If the above factors are solved, the effects of reducing error redundancy data in the role text paragraph information set, improving low content quality of the role text paragraph, and reducing waste of storage resources can be achieved. To achieve this effect, the present disclosure first inputs the generated speech mel spectrum graph into the user voiceprint extraction coding network included in the speech text disentanglement model, the user voiceprint extraction coding network extracts multi-scale voiceprint features, the introduced voiceprint feature dimension weighting network can focus on frequency bands with strong discrimination, the weighted feature fusion layer can aggregate multi-level features to avoid information loss, and capture and represent the unique acoustic features of the conference user. Second, the speech mel spectrum graph is input into the speech text coding network, which can not only accurately retain the speech content information of the conference user, but also realize the normalization of the identity information of the conference user, which can effectively decouple and separate the speech content and the identity information of the conference user. Subsequently, through the speech text disentanglement decoding model, the conference user is adaptively normalized, and the reconstructed feature vector can be more consistent with the initial input audio information in the process of reconstruction. Then, the reconstructed speech mel spectrum graph set is input into the user recognition classifier, and the reconstructed speech mel spectrum graph set forced model uses the disentangled pure voiceprint features to improve the recognition of the conference user. The speech text disentanglement model is trained through the text content reconstruction criterion and the identity invariant criterion to convert the frame-level features into sentence-level vectors, represent the global voiceprint features of the conference user, and the identity invariant criterion can reduce the intra-class difference and ensure that the voiceprint vectors in different sentences of the same user are similar.Then, through the conference user identification information, the enhanced conference audio and video are subjected to multi-aspect user role matching recognition of text, audio, and lip feature information in the face information of the conference user, so that the text paragraph information of each conference user can be accurately recognized, and the redundant content inconsistent with the conference theme can be removed, so that the quality of the role text paragraph information can be improved, and the waste of storage resources can be reduced.
[0084] In step 106, the enhanced conference audio and video are subjected to audio and video segmentation according to the role text paragraph information set, to obtain a role conference audio and video set.
[0085] In some embodiments, the above-mentioned execution subject can perform audio and video segmentation on the above-mentioned enhanced conference audio and video according to the above-mentioned role text paragraph information set, to obtain a role conference audio and video set. The role conference audio and video in the role conference audio and video set can be the audio and video of a user participating in the conference.
[0086] As an example, the above-mentioned execution subject can match the above-mentioned role text paragraph information set with the above-mentioned enhanced conference audio and video, to perform audio and video segmentation on the enhanced conference audio and video, to obtain a role conference audio and video set.
[0087] In step 107, face audio emotion recognition is performed on the role conference audio and video set, to obtain a role emotion information set.
[0088] In some embodiments, the above-mentioned execution subject can perform face audio emotion recognition on the above-mentioned role conference audio and video set, to obtain a role emotion information set. The role emotion information in the role emotion information set can be information about the emotional change of a user participating in the conference. The role emotion information can include but is not limited to at least one of the following: voice emotion based on speech speed and tone, facial emotion.
[0089] In some optional implementations of some embodiments, the above-mentioned face audio emotion recognition on the above-mentioned role conference audio and video set, to obtain a role emotion information set, can include the following steps:
[0090] First, for each role conference audio and video in the above-mentioned role conference audio and video set, the following emotion recognition steps are performed:
[0091] Substep 1: emotion feature extraction is performed on the above-mentioned role conference audio and video, to obtain a voice emotion feature map set. The voice emotion feature map can represent the mel-spectrogram, the root mean square of energy, the chroma distribution, the zero-crossing rate, and the mel-frequency cepstral coefficient of the role conference audio. The emotion feature extraction can be performed by using the Librosa audio library.
[0092] Sub-step 2, the variance and mean based feature stacking is performed on the above-mentioned speech emotion feature map set, and a one-dimensional speech emotion feature vector is obtained. The one-dimensional speech emotion feature vector can include 160 mel frequency cepstral coefficient features, 2 root mean square features of energy, 2 zero-crossing rate features, 256 mel spectrum features, and 24 chroma distribution features.
[0093] Sub-step 3, the one-dimensional speech emotion feature vector is input into a bidirectional long short-term memory neural network to obtain a speech forward emotion time sequence feature vector and a speech reverse emotion time sequence feature vector.
[0094] Sub-step 4, the speech forward emotion time sequence feature vector and the speech reverse emotion time sequence feature vector are spliced to obtain a speech emotion time sequence feature vector. The speech emotion time sequence feature vector can represent the time sequence information of the speech.
[0095] Sub-step 5, the speech emotion time sequence feature vector is input into a self-attention mechanism layer to obtain a speech emotion weight feature vector.
[0096] Sub-step 6, the speech emotion weight feature vector and the speech emotion time sequence feature vector are input into a sentiment recognition capsule network to obtain a speech emotion time sequence feature vector. The sentiment recognition capsule network includes a sentiment recognition convolution layer, a main capsule network, and a category capsule network. The sentiment recognition capsule network is a one-dimensional capsule network that uses vectorization and dynamic routing algorithm to extract features between speech spatial levels of the input speech emotion weight feature vector and the speech emotion time sequence feature vector. The convolution layer can be a convolution layer with 64 convolution kernels of 2*2 and a step of 1. The main capsule layer can include 32 capsule networks, each of which includes a filter of 9*1 size and a capsule network with a stride of 1. The category capsule network can include a preset number of emotion categories and a capsule network of size 9.
[0097] Sub-step 7, dynamic body emotion recognition is performed on the role conference video included in the role conference audio and video to obtain a dynamic body emotion feature vector. The dynamic body emotion feature vector can represent the emotion information determined by the user's body movements. The dynamic body emotion recognition can be performed by a dual-flow network model.
[0098] Sub-step 8, facial expression recognition is performed on the role conference video to obtain a facial expression feature vector. The facial expression recognition can be performed by a three-dimensional convolutional neural network model that simultaneously extracts spatial and temporal expression feature information. The three-dimensional convolutional neural network model can include 7 three-dimensional convolutional layers, one pooling layer after the first, second, third, fifth, and seventh layers, and three fully connected layers.
[0099] Sub-step 9, text sentiment recognition is performed on the role text paragraph information corresponding to the role audio and video of the above-mentioned conference to obtain a text sentiment feature vector. The text sentiment recognition can be performed by using a BERT (Bidirectional Encoder Representations from Transformers) model.
[0100] Sub-step 10, orthogonal constraint fusion sentiment recognition is performed on the voice emotion time sequence feature vector, the dynamic body emotion feature vector, the facial expression feature vector, and the text emotion feature vector to obtain role emotion information.
[0101] In the process of adopting the technical solutions to the technical problems mentioned in the background art, the following technical problems often occur: due to the heterogeneity of the emotion information of each modality in multi-modal emotion recognition, the feature distribution gap and information redundancy are caused, the data amount of the feature vector extracted by emotion recognition is increased, the emotion recognition efficiency is low, the recognition time is long, and the recognition accuracy is low. In view of the above technical problems, the conventional solution is generally as follows: the correlation between each emotion feature vector is determined by using a correlation analysis algorithm, and then each emotion feature vector is recognized after feature fusion according to the correlation to obtain role emotion information. However, the above conventional solution still has the following problems: since only the correlation between each emotion vector is considered, the distribution gap and information redundancy between each emotion vector are not considered, the data amount of each emotion feature vector is large, the emotion recognition accuracy is low, the recognition efficiency is low, and the recognition time is long, which further leads to low accuracy of the generated role personalized summary information. Considering the shortcomings of the above conventional solution and combining the advantages / technical status of the multi-modal emotion recognition technology possessed by the company where the inventors work, we decided to adopt the following solution:
[0102] Optionally, the orthogonal constraint fusion emotion recognition of the voice emotion time sequence feature vector, the dynamic body emotion feature vector, the facial expression feature vector, and the text emotion feature vector to obtain the role emotion information can include the following steps:
[0103] Firstly, the voice emotion time sequence feature vector, the dynamic body emotion feature vector, the facial expression feature vector, and the text emotion feature vector are grouped and combined to obtain emotion feature vector grouping information set. The emotion feature vector grouping information in the emotion feature vector grouping information set can be grouping information obtained by combining any two emotion feature vectors.
[0104] Secondly, for each piece of the above emotional feature vector grouping information, the following feature redundancy removal step is performed:
[0105] Sub-step 1: Perform fusion mapping on each emotional feature vector included in the above emotional feature vector grouping information to obtain an emotional fusion feature vector. The emotional fusion feature vector is a feature vector obtained by fusing emotional feature vectors mapped to the same subspace. The fusion mapping can be performed by a three-layer MLP with dimensions of 2048, 1024, and 1024.
[0106] Sub-step 2: Generate a positive and negative sample set according to the above each emotional feature vector and the above emotional fusion feature vector. The positive samples in the positive and negative sample set can be emotional feature vectors and emotional fusion feature vectors under the same emotional category. The negative samples can be samples other than the positive samples.
[0107] Sub-step 3: Determine the emotional shared feature vector of the above each emotional feature vector and the above emotional fusion feature vector according to the above positive and negative sample set by using an emotional shared feature contrast learning algorithm. The emotional shared feature vector can be a feature vector shared by the above each emotional feature vector. The emotional shared feature contrast learning algorithm can be
[0108]
[0109] wherein, LF con represents the contrast learning loss function of the emotional shared feature contrast learning algorithm. I represents a preset emotional category set, and the preset emotional category in the preset emotional category set can be a category of user emotion preset in advance. For example, the preset emotional category set can include but is not limited to at least one of the following: happy, angry, and passionate. represents one emotional feature vector under the i-th emotional category in the emotional feature vector grouping information. represents another emotional feature vector under the i-th emotional category in the emotional feature vector grouping information. represents the emotional fusion feature vector. (·, ·) represents the mutual information of the emotional feature vector, which is obtained by the InfoNCE loss function. Z represents the positive and negative sample set. |·| represents the L1 norm.
[0110] In practice, the above positive and negative sample set is input into the above emotional shared feature contrast learning algorithm, and a minimization solution is performed to obtain the emotional shared feature vector.
[0111] Sub-step 4, according to the above-mentioned emotion sharing feature vector, determine the unique emotion feature vector set of each of the above-mentioned emotion feature vectors. Among them, the unique emotion feature vector in the unique emotion feature vector set can be the feature vector that the emotion feature vector alone has.
[0112] As an example, the above-mentioned execution subject can subtract each emotion feature vector from the above-mentioned emotion sharing feature vector to obtain the unique emotion feature vector set.
[0113] Sub-step 5, generate a feature redundancy orthogonal constraint function of the above-mentioned emotion sharing feature vector and the above-mentioned unique emotion feature vector set. Among them, the feature redundancy orthogonal constraint function can be a function for the above-mentioned emotion sharing feature vector and the above-mentioned unique emotion feature vector set to have no correlation relationship and not affect each other, so as to reduce the redundancy information between the feature vectors. The feature redundancy orthogonal constraint function can be:
[0114]
[0115] Among them, L ort represents the feature redundancy orthogonal constraint function. represents the unique emotion feature vector of one emotion feature vector under the i-th emotion category in the emotion feature vector grouping information. represents the unique emotion feature vector of another emotion feature vector under the i-th emotion category in the emotion feature vector grouping information. 2 represents the square of the dot product of the feature vector vector, which is used to measure the similarity of two feature vectors.
[0116] Sub-step 6, minimize the above-mentioned feature redundancy orthogonal constraint function to obtain the emotion sharing redundancy removed feature vector and the unique emotion redundancy removed feature vector set. Among them, the emotion sharing redundancy removed feature vector can be the feature vector after removing the redundancy information of the emotion sharing feature vector. The unique emotion redundancy removed feature vector in the unique emotion redundancy removed feature vector set can be the feature vector after removing the redundancy information of the unique emotion feature vector.
[0117] Third step, perform feature association graph construction on the obtained each emotion sharing redundancy removed feature vector and each unique emotion redundancy removed feature vector set to obtain a feature association graph set. Among them, the feature association graph in the feature association graph set can be a graph form representing the association relationship between the feature vectors in different groups.
[0118] In the fourth step, inter-group feature redundancy is removed from the feature correlation graph set to obtain a shared sentiment feature vector set and a unique sentiment feature vector group set. The shared sentiment feature vector can be a shared sentiment feature vector after removing the redundancy information between groups. The unique sentiment feature vector in the unique sentiment feature vector group set can be a unique sentiment feature vector after removing the redundancy information between groups. In practice, the execution subject can use a community discovery algorithm to remove the inter-group feature redundancy from the feature correlation graph set to obtain the shared sentiment feature vector set and the unique sentiment feature vector group set.
[0119] In the fifth step, fusion sentiment recognition is performed on the shared sentiment feature vector set and the unique sentiment feature vector group set to obtain role sentiment information. The fusion sentiment recognition can be fusion sentiment recognition in which a three-layer MLP (Multi-Layer Perceptron) is first used for feature fusion, and then input into a sentiment classifier for sentiment classification. The dimension of the first layer MLP can be the sum of the dimensions of the shared sentiment feature vector and the unique sentiment feature vector set. The dimension of the second layer MLP is 1024. The dimension of the third layer MLP is the number of categories of sentiment classification.
[0120] The technical solution and related content thereof serve as one inventive point of the embodiments of the present disclosure, and solve the technical problem mentioned in the background art, i.e., the feature distribution gap and information redundancy caused by the heterogeneous emotion information of each modality in multi-modal emotion recognition, which increases the data volume of the feature vector extracted by emotion recognition, results in low emotion recognition efficiency, long recognition time, and low recognition accuracy. The factors that result in low emotion recognition efficiency and long recognition time, and further result in low accuracy of generated role personalized summary information are usually as follows: the feature distribution gap and information redundancy caused by the heterogeneous emotion information of each modality in multi-modal emotion recognition, which increases the data volume of the feature vector extracted by emotion recognition, results in low emotion recognition efficiency, long recognition time, and low recognition accuracy. If the above factors are solved, the emotion recognition efficiency can be improved, the recognition time can be shortened, and the accuracy of role personalized summary information can be improved. To achieve this effect, the present disclosure first groups and combines the speech emotion feature vector, the dynamic body emotion feature vector, the facial expression feature vector, and the text emotion feature vector, and then determines the shared feature vector and the unique feature vector in the grouped information set by using an emotion shared feature comparison learning algorithm, so as to reduce the distribution gap of the feature vector. Then, by minimizing the feature redundancy orthogonal constraint function, the redundant information between the features can be accurately identified and removed. Finally, the redundant information between groups is removed from the obtained feature vectors after grouping, and through the combination of inter-group and intra-group, the redundant information between the feature vectors can be more accurately identified and removed, the data volume of the feature vector can be reduced, the emotion recognition accuracy and efficiency can be improved, the recognition time can be shortened, and the accuracy of the role personalized summary information can be improved.
[0121] In step 108, the role personalized summary information set of the role text paragraph information set is generated according to the role emotion information set, and the role personalized summary information set and the participant user information set are stored in association.
[0122] In some embodiments, the above execution subject can generate the role personalized summary information set of the role text paragraph information set according to the above role emotion information set, and store the above role personalized summary information set and the above participant user information set in association. The role personalized summary information in the role personalized summary information set can be the summary information of the text information generated by the audio of the user. The above association storage can be storing the role personalized summary information set with the participant user information as an index, so that the participant user can quickly backtrack the information of the conference audio and video, shorten the backtracking and memory time of the participant user, and improve the user experience.
[0123] As an example, the above execution subject can first filter at least one role emotional information with a positive excitement emotional category from the above role emotional information set. Then, the text information corresponding to the above at least one role emotional information is highlighted to obtain a set of role text paragraph information with highlighting. Finally, the set of role text paragraph information with highlighting is summarized to obtain a set of role personalized summary information.
[0124] Further reference Figure 2 , as an implementation of the method shown in the above figures, the disclosure provides some embodiments of a conference role summary information generation device, which corresponds to the method embodiments shown in Figure 1 , the conference role summary information generation device can be applied to various electronic devices.
[0125] As shown in Figure 2 , a conference role summary information generation device 200 includes an acquisition unit 201, an audio and video enhancement unit 202, a speech text conversion unit 203, a conference user identification unit 204, a user role matching identification unit 205, an audio and video segmentation unit 206, a face audio emotion recognition unit 207, and a generation unit 208. Wherein, the acquisition unit 201 is configured to: acquire conference audio and video and a set of conference user information corresponding to the conference audio and video. The audio and video enhancement unit 202 is configured to: perform audio and video enhancement processing on the conference audio and video to obtain enhanced conference audio and video. The speech text conversion unit 203 is configured to: perform speech text conversion on the enhanced conference audio and video to obtain conference text information. The conference user identification unit 204 is configured to: perform conference user identification on the enhanced conference audio and video to obtain a set of conference user face information. The user role matching identification unit 205 is configured to: according to the set of conference user face information, the set of conference user information and the conference text information, perform user role matching identification on the enhanced conference audio and video to obtain a set of role text paragraph information. The audio and video segmentation unit 206 is configured to: according to the set of role text paragraph information, perform audio and video segmentation on the enhanced conference audio and video to obtain a set of role conference audio and video. The face audio emotion recognition unit 207 is configured to: perform face audio emotion recognition on the set of role conference audio and video to obtain a set of role emotional information. The generation unit 208 is configured to: generate a set of role personalized summary information of the set of role text paragraph information according to the set of role emotional information, and store the set of role personalized summary information and the set of conference user information in association.
[0126] It can be understood that the units recorded in the conference role summary information generation device 200 are referred to Figure 1The various steps in the described methods correspond. Thus, the operations, features, and benefits described above for the methods apply equally to the conference role summary information generation apparatus 200 and the elements incorporated within it, and are not repeated here.
[0127] Reference is made below to Figure 3 which shows a structural schematic diagram of an electronic device (e.g., electronic device) 300 suitable for use in implementing some embodiments of the present disclosure. Figure 3 The illustrated electronic device is merely one example. It should not be considered a limitation on the functioning and usefulness of embodiments of the present disclosure.
[0128] As shown in Figure 3 , the electronic device 300 can include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 301 that can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 302 or loaded into a random access memory (RAM) 303 from a storage device 308. Various programs and data required for operation of the electronic device 300 are also stored in the RAM 303. The processing device 301, the ROM 302, and the RAM 303 are connected to each other through a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0129] Generally, the following devices can be connected to the I / O interface 305: an input device 306 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 308 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 309. The communication device 309 can allow the electronic device 300 to communicate wirelessly or wired with other devices to exchange data. Although Figure 3 The electronic device 300 is shown with various devices, but it should be understood that not all of the illustrated devices are required to be implemented or present. More or fewer devices can alternatively be implemented or present. Figure 3 Each block shown in the flowchart of
[0130] In particular, the processes described above with reference to the flowcharts can be implemented as a computer software program according to some embodiments of the present disclosure. For example, some embodiments of the present disclosure include a computer program product comprising a computer program carried on a computer readable medium, the computer program comprising program code for performing the methods illustrated by the flowcharts. In some such embodiments, the computer program can be downloaded and installed from a network via the communication device 309, or installed from the storage device 308, or installed from the ROM 302. When the computer program is executed by the processing device 301, the above-mentioned functions defined in the methods of some embodiments of the present disclosure are performed.
[0131] It should be noted that the computer readable medium mentioned above in some embodiments of the present disclosure can be a computer readable signal medium or a computer readable storage medium or any combination of the two. The computer readable storage medium may, for example, be, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any combination of the above. More specific examples of the computer readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, the computer readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus or device. In some embodiments of the present disclosure, the computer readable signal medium can include a data signal carried in a baseband or as part of a carrier wave, in which the computer readable program code is carried. Such a propagated data signal can take many forms, including but not limited to, an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer readable signal medium can also be any computer readable medium that is not a computer readable storage medium and that can communicate, propagate or transport a program for use by or in connection with an instruction execution system, apparatus or device. The program code contained on the computer readable medium can be transmitted by any suitable medium, including but not limited to, wire, cable, RF (radio frequency), etc., or any suitable combination of the above.
[0132] In some embodiments, the client, server, or other computing devices can communicate information using any known or future developed end-to-end communications protocol, such as the Hyper Text Transfer Protocol (HTTP), and can be interconnected via any form or medium of digital data communication (for example, a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), the Internet, and peer-to-peer networks (for example, ad hoc peer-to-peer networks), as well as any current or future developed network.
[0133] The computer readable medium described above can be included within the electronic device described above; alternatively, the computer readable medium can exist as a standalone entity independent of the electronic device. The computer readable medium described above carries one or more programs which, when executed by the electronic device, cause the electronic device to perform steps 101-108.
[0134] Computer program code for carrying out operations of some embodiments of the present disclosure can be written in any of one or more programming languages, including object oriented programming languages such as Java, Smalltalk, C++, or conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network ("LAN") or a wide area network ("WAN"), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0135] The flow and block diagrams in the drawings represent possible architectural, functional, and operational architectures of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flow and block diagrams can represent a module, a segment, or a portion of code that comprises one or more executable instructions for implementing the specified logical function(s). It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations thereof, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or combinations of hardware and software.
[0136] The units described in some embodiments of the present disclosure can be implemented in the form of software, or can be implemented in the form of hardware. The described units can also be arranged in a processor, for example, it can be described that: a processor includes an acquisition unit, an audio and video enhancement unit, a speech text conversion unit, a participant user identification unit, a user role matching identification unit, an audio and video segmentation unit, a face audio emotion identification unit and a generation unit. Among them, the name of these units does not constitute a limitation to the unit itself in some cases, for example, the acquisition unit can also be described as "a unit for acquiring conference audio and video and participant user information set corresponding to the conference audio and video".
[0137] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, without limitation, example types of hardware logic components that can be used include: field programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip systems (SOCs), complex programmable logic devices (CPLDs), etc.
[0138] The above description is merely some of the preferred embodiments of the present disclosure and a description of the principles of the technology used. Those skilled in the art should understand that the scope of the application involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or equivalent features without departing from the above inventive concept. For example, the above features are replaced with the technical features disclosed in the embodiments of the present disclosure (but not limited to) having similar functions to form technical solutions.
Claims
1. A conference role summary information generation method, comprising: obtaining conference audio and video and a set of conference participant user information corresponding to the conference audio and video; performing audio and video enhancement processing on the conference audio and video to obtain enhanced conference audio and video; performing speech text conversion on the enhanced conference audio and video to obtain conference text information; performing participant user identification on the enhanced conference audio and video to obtain a set of participant user face information; performing user role matching identification on the enhanced conference audio and video according to the set of participant user face information, the set of conference participant user information, and the conference text information to obtain a set of role text paragraph information; performing audio and video segmentation on the enhanced conference audio and video according to the set of role text paragraph information to obtain a set of role conference audio and video; performing face and audio emotion recognition on the set of role conference audio and video to obtain a set of role emotion information; generating a set of role personalized summary information of the set of role text paragraph information according to the set of role emotion information, and storing the set of role personalized summary information and the set of conference participant user information in association.
2. The method of claim 1, wherein, The audio and video enhancement processing on the conference audio and video to obtain enhanced conference audio and video comprises: extracting audio from the conference audio and video to obtain conference audio; performing short-time Fourier transform on the conference audio to obtain conference speech spectrum graph; performing amplitude masking feature extraction and phase feature extraction on the conference speech spectrum graph in parallel to obtain a set of speech amplitude feature graphs and a set of speech phase feature graphs; performing speech noise reduction on the set of speech amplitude feature graphs and the set of speech phase feature graphs to obtain denoised conference speech spectrum graph; performing spectrum speech enhancement on the denoised conference speech spectrum graph to obtain enhanced conference speech spectrum graph; performing video enhancement on the conference video in the conference audio and video to obtain enhanced conference video; determining the enhanced conference speech spectrum graph and the enhanced conference video as enhanced conference audio and video.
3. The method of claim 2, wherein, The speech noise reduction on the set of speech amplitude feature graphs and the set of speech phase feature graphs to obtain denoised conference speech spectrum graph comprises: based on the set of speech amplitude feature graphs and the set of speech phase feature graphs, performing the following speech noise reduction steps: inputting the set of speech amplitude feature graphs into a first speech frequency transformation network to obtain a set of first speech harmonic amplitude feature graphs; inputting the set of first speech harmonic amplitude feature graphs into a plurality of amplitude convolution feature extraction layers in sequence to obtain a set of second speech harmonic amplitude feature graphs; inputting the set of second speech harmonic amplitude feature graphs into a second speech frequency transformation network to obtain a set of third speech harmonic amplitude feature graphs; inputting the set of speech phase feature graphs into a plurality of phase convolution feature extraction networks in sequence to obtain a set of speech phase mask feature graphs; performing double-flow feature exchange processing on the set of third speech harmonic amplitude feature graphs and the set of speech phase mask feature graphs to obtain a set of speech amplitude exchange feature graphs and a set of speech phase exchange feature graphs; In response to determining that the number of times of executing the voice denoising step is greater than or equal to a preset execution threshold, performing fusion denoising on the voice amplitude exchange feature map set and the voice phase exchange feature map set to obtain a denoised conference voice spectrum map; In response to determining that the number of times of executing is less than the preset execution threshold, determining the voice amplitude exchange feature map set and the voice phase exchange feature map set as a voice amplitude feature map set and a voice phase feature map set respectively, and determining the sum of the number of times of executing and a preset value as the number of times of executing, so as to execute the voice denoising step again.
4. The method of claim 2, wherein, The inputting of the voice amplitude feature map set into the first voice frequency transformation network to obtain a first voice harmonic amplitude feature map set comprises: The inputting of the voice amplitude feature map set into a channel convolution network to obtain a voice amplitude channel feature map set; The feature dimension conversion of the voice amplitude channel feature map set to obtain a dimension-converted voice amplitude channel feature map set; The inputting of the dimension-converted voice amplitude channel feature map set into a one-dimensional convolution network to obtain a voice amplitude attention feature map; The point-by-point multiplication of the voice amplitude attention feature map and the voice amplitude feature map set to obtain a semantic amplitude weight feature map set; The mapping and slicing processing of the semantic amplitude weight feature map set to obtain a voice amplitude frequency transformation feature map set; The feature map splicing of each voice amplitude frequency transformation feature map in the voice amplitude frequency transformation feature map set and the corresponding voice amplitude feature map in the voice amplitude feature map set to obtain a spliced voice amplitude feature map set; The channel number dimension reduction of the spliced voice amplitude feature map set to obtain the first voice harmonic amplitude feature map set.
5. The method of claim 1, wherein, The voice text conversion of the enhanced conference audio and video to obtain conference text information comprises: The frame windowing processing of the enhanced conference audio included in the enhanced conference audio and video to obtain a voice short-time energy set and a voice short-time zero-crossing rate set; The adaptive time domain feature segmentation of the enhanced conference audio according to the voice short-time energy set and the voice short-time zero-crossing rate set to obtain a conference voice segmented syllable sequence; The Mel frequency cepstrum recognition of the conference voice segmented syllable sequence to obtain a voice syllable feature map set; The convolution feature extraction of the voice syllable feature map set to obtain a voice syllable convolution feature map set; The regional difference syllable feature extraction of the voice syllable convolution feature map set to obtain a voice syllable regional difference feature map set; The frequency spectrum enhancement of the voice syllable regional difference feature map set to obtain a continuous voice local feature map set; The convolution down-sampling processing of the continuous voice local feature map set to obtain a down-sampled voice local feature map set; The linear regularization processing of the down-sampled voice local feature map set to obtain a regularized voice local feature map set; The inputting of the regularized voice local feature map set into a grapheme syllable dependency recognition model to obtain conference text information.
6. The method of claim 1, wherein, The face and audio emotion recognition of the role conference audio and video set to obtain a role emotion information set comprises: For each role conference audio and video in the role conference audio and video set, the following emotion recognition steps are executed: extracting an emotional feature of the role conference audio and video to obtain a speech emotional feature graph set; stacking the speech emotional feature graph set based on variance and mean to obtain a one-dimensional speech emotional feature vector; inputting the one-dimensional speech emotional feature vector into a bidirectional long short-term memory neural network to obtain a speech positive emotional time sequence feature vector and a speech reverse emotional time sequence feature vector; concatenating the speech positive emotional time sequence feature vector and the speech reverse emotional time sequence feature vector to obtain a speech emotional time sequence feature vector; inputting the speech emotional time sequence feature vector into a self-attention mechanism layer to obtain a speech emotional weight feature vector; inputting the speech emotional weight feature vector and the speech emotional time sequence feature vector into an emotional recognition capsule network to obtain a speech emotional time sequence feature vector; performing dynamic body emotional recognition on a role conference video included in the role conference audio and video to obtain a dynamic body emotional feature vector; performing facial expression recognition on the role conference video to obtain a facial expression feature vector; performing text emotional recognition on role text paragraph information corresponding to the role conference audio and video to obtain a text emotional feature vector; performing orthogonal constraint fusion emotional recognition on the speech emotional time sequence feature vector, the dynamic body emotional feature vector, the facial expression feature vector, and the text emotional feature vector to obtain role emotional information. 7.A conference role summary information generation apparatus, comprising: an acquisition unit configured to acquire conference audio and video and a set of participant user information corresponding to the conference audio and video; an audio and video enhancement unit configured to perform audio and video enhancement processing on the conference audio and video to obtain enhanced conference audio and video; a speech text conversion unit configured to perform speech text conversion on the enhanced conference audio and video to obtain conference text information; a participant user identification unit configured to perform participant user identification on the enhanced conference audio and video to obtain a set of participant user face information; a user role matching identification unit configured to perform user role matching identification on the enhanced conference audio and video according to the set of participant user face information, the set of participant user information, and the conference text information to obtain a set of role text paragraph information; an audio and video segmentation unit configured to perform audio and video segmentation on the enhanced conference audio and video according to the set of role text paragraph information to obtain a set of role conference audio and video; a face audio emotional recognition unit configured to perform face audio emotional recognition on the set of role conference audio and video to obtain a set of role emotional information; a generation unit configured to generate a set of role personalized summary information of the set of role text paragraph information according to the set of role emotional information, and to store the set of role personalized summary information and the set of participant user information in association. 8.An electronic device, comprising: one or more processors; a storage device having one or more programs stored thereon, when the one or more programs are executed by the one or more processors, the one or more processors implement the method of any one of claims 1-6.
9. A computer readable medium having stored thereon a computer program, wherein, The computer program, which is executed by a processor, implements the method as claimed in any one of claims 1 to 6.