Speech recognition and transcription method and system based on multi-modal fusion and sentiment analysis

The speech recognition transcription method based on multimodal fusion and sentiment analysis solves the problems of dialect diversity and neglect of emotional information, achieves higher recognition accuracy and robustness, and is suitable for a variety of practical application scenarios.

CN120808788APending Publication Date: 2025-10-17山西益通电网保护自动化有限责任公司
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202511004798.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-21
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing speech recognition technology is difficult to adapt to the diversity of dialects, and ignores emotional information, resulting in insufficient recognition accuracy and robustness. It is especially difficult to achieve ideal recognition and transcription effects in complex and changeable practical application scenarios.

Method used

By acquiring the target speech signal and synchronized visual information and text context information, segmentation and recognition processing are performed, speech and text feature vectors are extracted, and multimodal fusion is performed. Emotional feature labels are generated in combination with sentiment analysis to optimize and correct the initial transcribed text.

Benefits of technology

It improves the accuracy and quality of speech recognition transcription, can more accurately identify speech content and optimize transcribed text, and enhances the system's ability to perceive and process complex emotional scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120808788A_ABST
    Figure CN120808788A_ABST
Patent Text Reader

Abstract

The invention relates to a speech recognition and transcription method and system based on multi-modal fusion and sentiment analysis, and relates to the field of speech recognizing.The speech recognition and transcription method comprises the steps that a target speech signal and auxiliary modal information of synchronous visual information and text context information are obtained firstly, and the speech signal is segmented and recognized to obtain speech feature vectors; the method comprises the following steps: extracting text context information to obtain a text auxiliary feature vector, carrying out multi-modal fusion on the text auxiliary feature vector and the text auxiliary feature vector to generate fusion feature representation so as to carry out voice transcription to obtain an initial transcription text, and carrying out sentiment analysis according to visual information and the initial transcription text to generate a sentiment feature tag; and finally, optimizing and correcting the initial transliteration text based on the label to obtain a target transliteration text, thereby solving the technical problems that the speech recognition transliteration is difficult to adapt to dialect diversity and the recognition accuracy and robustness are insufficient due to neglect of emotion information, and improving the recognition accuracy and robustness through fusion of multi-modal information and emotion analysis. The voice content can be recognized more accurately, the transcription text can be optimized, and the accuracy and quality of voice recognition transcription are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of speech recognition, and in particular to a speech recognition transcription method and system based on multi-modal fusion and emotion analysis. BACKGROUND

[0002] With the rapid development of information technology and the increasing demand for work communication efficiency, speech, as one of the most direct and natural communication methods, has an increasing demand for recognition and transcription technology. Traditional speech recognition technology mainly relies on a single modality of speech signal. However, in complex and variable practical application scenarios such as conference recording, remote interview, multimedia teaching, etc., it is often difficult to achieve ideal recognition and transcription effect by simply relying on speech signal.

[0003] On the one hand, the existence of dialects brings great challenges to speech recognition. Due to geographical differences, there are many types of dialects, and the speech features, vocabulary usage and even grammar structures between different dialects are significantly different. Traditional speech recognition systems are often trained based on standard Mandarin, although some systems with dialect speech recognition capabilities have appeared, but it is still difficult to adapt to the complexity and diversity of dialect speech, resulting in a significant decrease in recognition accuracy in dialect environment. On the other hand, speech signals often carry rich emotional information, such as the speaker's tone, intonation, and speech rate, which can reflect their emotional state, and these emotional information is often ignored in existing speech recognition systems, which leads to misinterpretation and ambiguity of dialect semantics. However, in many practical application scenarios such as conference communication, customer service, psychological counseling, etc., accurately capturing and understanding the emotional state of the speaker is crucial to correctly understand and transcribe the speech content. SUMMARY

[0004] The present application provides a speech recognition transcription method and system based on multi-modal fusion and emotion analysis to solve the technical problems of existing technology that speech recognition transcription is difficult to adapt to dialect diversity and ignores emotional information, resulting in insufficient recognition accuracy and robustness.

[0005] The technical solution of the present application to solve the above technical problems is as follows: In a first aspect, the present invention provides a speech recognition transcription method based on multimodal fusion and sentiment analysis, comprising: obtaining a target speech signal, and synchronously collecting auxiliary modal information corresponding to the target speech signal, the auxiliary modal information including visual information and text context information; performing segmentation and recognition processing on the target speech signal to obtain a speech feature vector, and performing feature extraction on the text context information to obtain a text auxiliary feature vector; performing multimodal fusion processing on the speech feature vector and the text auxiliary feature vector to generate a fusion feature representation; performing speech transcription based on the fusion feature representation to obtain an initial transcription text; performing sentiment analysis based on the visual information and the initial transcription text to generate a sentiment feature label; performing text optimization and correction on the initial transcription text based on the sentiment feature label to obtain a target transcription text.

[0006] In a second aspect, the present invention provides a speech recognition and transcription system based on multimodal fusion and sentiment analysis, the system comprising: an information acquisition module for acquiring a target speech signal and synchronously acquiring auxiliary modal information corresponding to the target speech signal, the auxiliary modal information comprising visual information and text context information; a signal segmentation module for performing segmentation and recognition processing on the target speech signal to obtain a speech feature vector, and performing feature extraction on the text context information to obtain a text auxiliary feature vector; a fusion processing module for performing multimodal fusion processing on the speech feature vector and the text auxiliary feature vector to generate a fusion feature representation; a speech transcription module for performing speech transcription based on the fusion feature representation to obtain an initial transcription text; a sentiment analysis module for performing sentiment analysis based on the visual information and the initial transcription text to generate a sentiment feature label; an optimization and correction module for performing text optimization and correction on the initial transcription text based on the sentiment feature label to obtain a target transcription text.

[0007] The beneficial effects of the present invention are as follows: by first acquiring the target speech signal and the auxiliary modal information of the synchronized visual information and text context information, the speech signal is segmented and recognized to obtain a speech feature vector, the text context information is extracted to obtain a text auxiliary feature vector, and then the two are multimodally fused to generate a fused feature representation for speech transcription to obtain an initial transcribed text, and then sentiment analysis is performed based on the visual information and the initial transcribed text to generate a sentiment feature label, and finally the initial transcribed text is optimized and corrected based on the label to obtain the target transcribed text. Through the fusion of multimodal information and sentiment analysis, the speech content can be identified more accurately and the transcribed text can be optimized, thereby improving the accuracy and quality of speech recognition transcription. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Figure 1 A flow chart of the speech recognition transcription method based on multimodal fusion and sentiment analysis provided by the present invention.

[0009] Figure 2 A structural schematic diagram of a speech recognition transcription system based on multi-modal fusion and emotion analysis provided by the present application.

[0010] Legend: information acquisition module 11, signal segmentation module 12, fusion processing module 13, speech transcription module 14, emotion analysis module 15, optimization correction module 16. DETAILED DESCRIPTION

[0011] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0012] In the description of the present application, the terms "first", "second" are used only for the purpose of description, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined with "first", "second" can explicitly or implicitly include one or more of the features. In the description of the present application, the meaning of "multiple" is two or more, unless otherwise specifically limited.

[0013] In the description of the present application, the term "for example" is used to indicate "as an example, illustration or explanation". Any embodiment described as "for example" in the present application is not necessarily interpreted as more preferred or more advantageous than other embodiments. The following description is given in order to enable any person skilled in the art to implement and use the present application. In the following description, details are listed for the purpose of explanation. It should be understood that those skilled in the art can realize the present application without using these specific details. In other examples, well-known structures and processes will not be described in detail to avoid unnecessary details making the description of the present application obscure. Therefore, the present application is not intended to be limited to the shown embodiments, but is consistent with the broadest scope in accordance with the principles and characteristics disclosed.

[0014] Embodiment one: As shown in the figure, the present application provides a speech recognition transcription method based on multi-modal fusion and emotion analysis, comprising: Figure 1 S10: obtaining a target speech signal, and synchronously acquiring auxiliary modal information corresponding to the target speech signal, the auxiliary modal information including visual information and text context information. S10: obtaining a target speech signal, and synchronously acquiring auxiliary modal information corresponding to the target speech signal, the auxiliary modal information including visual information and text context information.

[0015] For example, voice recognition transcription is widely used in daily life and many industry fields. For example, in a business meeting, participants discuss project planning, market strategy, and other topics. A voice recognition system can convert the speeches of all parties into text records in real time, which facilitates the subsequent preparation of meeting minutes and the tracing of decision-making basis. In an online education classroom, teachers explain knowledge points and answer students' questions. Voice recognition technology can quickly convert the teaching speech into text, which is convenient for students to review key content after class and provides convenience for making course text materials. In a medical setting, doctors ask patients to describe their condition and explain their diagnosis approach. Voice recognition transcription helps quickly generate medical records, improving medical efficiency and the standardization of medical records. In addition, in news interviews, customer service communication and other scenarios, voice recognition transcription also plays a key role in efficiently converting voice information into editable and searchable text, greatly facilitating information processing, storage and dissemination.

[0016] Dialects, as unique carriers of specific regional culture and language, have both challenges and important roles in voice recognition transcription. On the one hand, dialects have significant differences in pronunciation, vocabulary, and grammar from Mandarin. The complex phoneme combination, special tone variation, and unique vocabulary expression increase the difficulty of accurate recognition and transcription by voice recognition systems. If the system is not optimized for specific dialects, it may lead to high error rates and poor transcription quality. On the other hand, dialects contain rich regional cultural information and social historical value. Accurate recognition and transcription of dialects can help protect and pass on local culture, preserve and disseminate dialect materials, and meet the actual needs of public services and folklore research in dialect areas, improving information processing efficiency and accuracy and providing strong support for related work.

[0017] The present scheme is based on the transcription of dialect speech to improve recognition and information processing efficiency. Specifically, the target voice signal is first obtained. This signal is the core data to be processed and carries the specific voice content conveyed by the speaker, with specific frequency, amplitude, and duration characteristics.

[0018] Meanwhile, to improve the accuracy and reliability of speech recognition transcription, auxiliary modal information corresponding to the target speech signal needs to be collected synchronously, where the auxiliary modal information includes visual information and text context information. The visual information is usually collected by means of a camera or the like, and it can capture non-verbal cues such as facial expressions and body movements of the speaker. For example, when the speaker expresses a happy emotion, his face may show a smile and his body movements may be more lively; while expressing anger, the facial muscles are tense and the body movements may be more intense. These visual information can be used as important auxiliary basis to help better understand the emotions and semantics contained in the speech. The text context information refers to the text content related to the target speech signal, which can be the text appearing before or after the speech, such as the speech record before the meeting agenda in a meeting scenario. By analyzing these text contexts, the topic background and professional term usage of the speech can be determined, thereby providing more semantic clues for speech recognition.

[0019] The comprehensive use of visual information and text context information as auxiliary modal information in cooperation with the target speech signal can significantly make up for the limitations of single speech signal in semantic understanding, effectively reduce recognition errors caused by environmental noise, accent differences and other factors, and further improve the accuracy and robustness of subsequent speech recognition transcription.

[0020] S20: performing segmentation and recognition processing on the target speech signal to obtain a speech feature vector, and performing feature extraction on the text context information to obtain a text auxiliary feature vector.

[0021] Preferably, further segmentation and recognition processing of the obtained target speech signal is a key link. According to the acoustic characteristics of the speech signal, such as the features of the silent section and the phoneme boundary, an endpoint detection algorithm is used to segment the continuous target speech signal into a plurality of short-time speech segments, each segment usually having a duration of tens of milliseconds. Such segmentation can better capture the local features of the speech signal in the time dimension.

[0022] Subsequently, feature extraction is performed on each short-time speech segment after segmentation. Common features include Mel-frequency cepstral coefficients (MFCC), which simulate the human ear's perception characteristics of different frequencies and can effectively represent the spectral information of the speech signal; and linear prediction coefficients (LPC), which are based on the speech signal generation model and extract the vocal tract characteristic parameters of the speech through linear prediction analysis. Through these feature extraction methods, each short-time speech segment is converted into a fixed-dimensional speech feature vector, which contains the acoustic feature information of the segment and provides basic data for subsequent speech recognition.

[0023] Meanwhile, for the synchronous collected text context information, natural language processing techniques are used for feature extraction. For example, the bag-of-words model is used to convert the words in the text into vector representation, the frequency of each word appearing in the text is counted, and a word-frequency matrix is constructed, each row or column of which can be regarded as a text auxiliary feature vector; or word embedding techniques such as Word2Vec, GloVe, etc. are used to map words to low-dimensional continuous vector space, so that semantically similar words are closer in vector space, thereby capturing the semantic relationship between words in the text and generating text auxiliary feature vectors that can reflect the semantic information of the text.

[0024] By performing segmentation and recognition processing on the target speech signal to obtain a speech feature vector, and performing feature extraction on the text context information to obtain a text auxiliary feature vector, effective feature representations can be provided for subsequent multi-modal fusion from the acoustic level of the speech and the semantic level of the text, which helps to improve the accuracy and robustness of speech recognition transcription, enabling the system to better understand the relationship between the speech signal and the text context, and thus more accurately complete the speech-to-text transcription task.

[0025] S30: Perform multi-modal fusion processing on the speech feature vector and the text auxiliary feature vector to generate a fusion feature representation.

[0026] Further, multi-modal fusion processing of the speech feature vector and the text auxiliary feature vector is an important step to improve transcription quality. The speech feature vector contains the acoustic characteristics of the target speech signal, such as pitch, intensity, timbre, etc. These characteristics are presented in numerical form, reflecting the specific information of the speech at the physical level; the text auxiliary feature vector captures the semantic and grammatical structured features in the text context information, such as the vectors obtained through word embedding techniques that can reflect the semantic association and similarity between words. Multi-modal fusion processing aims to integrate the feature information of these two different modalities to fully utilize their respective advantages.

[0027] Common fusion methods include early fusion, mid-fusion, and late fusion. Early fusion directly concatenates the speech feature vector and the text auxiliary feature vector at the feature level to form a longer feature vector, such as concatenating a 128-dimensional speech feature vector and a 300-dimensional text auxiliary feature vector into a 428-dimensional fusion feature vector. This approach is simple and direct, but may have problems such as high feature dimension, large difference in distribution of different modal features, etc. Mid-fusion is to fuse in the middle layer of the model, such as in a neural network model, the speech feature vector and the text auxiliary feature vector are input into different sub-networks for processing, and the outputs of the two sub-networks are fused in the middle layer. This way can preserve the independence of different modal features to some extent, while also allowing them to complement each other. Late fusion is to fuse at the decision level, first get the respective recognition results based on the speech feature vector and the text auxiliary feature vector, and then fuse these results according to certain strategies such as weighted average, voting mechanism, etc.

[0028] By generating fusion feature representation through multi-modal fusion processing, the acoustic information of speech and the semantic information of text can be integrated to make up for the shortcomings of single modal features. For example, when the presence of noise in speech causes some speech features to be difficult to extract accurately, the text auxiliary feature vector can provide semantic complementation to help the model better understand the speech content; conversely, when the text context information is ambiguous, the acoustic information in the speech feature vector can provide additional clues to improve the accuracy and robustness of speech recognition transcription, making the generated fusion feature representation more comprehensive and accurate in reflecting the true meaning of the target speech signal.

[0029] S40: Perform speech transcription based on the fusion feature representation to obtain an initial transcription text.

[0030] In detail, after completing the multi-modal fusion of the speech feature vector and the text auxiliary feature vector and generating the fusion feature representation, performing speech transcription based on the fusion feature representation is a key step to obtain the text result. The fusion feature representation combines the acoustic information of speech and the semantic information of text, and has more comprehensive and rich feature representation capability.

[0031] In the speech transcription process, deep learning models such as recurrent neural network (RNN) and its variants long short-term memory network (LSTM), gated recurrent unit (GRU), or attention mechanism-based Transformer model can be used. Taking the LSTM model as an example, it can process sequence data and effectively capture the temporal dependency in the fusion feature representation.

[0032] The model takes the fusion feature representation as input, and in the training phase, a large amount of speech-text data with labels is used to train the model, adjust the model parameters, and make the model learn the mapping relationship from the fusion feature to the correct text. In the inference phase, the model receives the fusion feature representation and sequentially predicts the characters or words corresponding to each time step. For example, when processing a fusion feature representation containing the speech content "today the weather is very good", the model will gradually predict the characters "today" "day" "day" "weather" "very" "good" according to the information in the fusion feature in time sequence, and finally combine them into the initial transcription text.

[0033] Compared with transcription based on only a single speech feature, transcription based on a fusion feature representation can significantly improve the accuracy of transcription. Because the fusion feature representation not only contains the acoustic characteristics of the speech, but also combines the semantic information of the text context, when the speech signal has noise, accent, and other interference, the semantic information can be used as an aid to help the model more accurately identify the speech content; at the same time, for some semantically ambiguous speech segments, acoustic information can also provide additional clues to reduce transcription errors. In addition, the fusion feature representation can also enhance the model's ability to handle complex language phenomena such as polysemy, ellipsis, etc., thereby generating an initial transcription text that is closer to the original speech meaning, providing a more reliable basis for subsequent text optimization and correction.

[0034] S50: performing sentiment analysis on the visual information and the initial transcription text to generate a sentiment feature label.

[0035] Optionally, sentiment analysis on the given visual information and initial transcription text to generate a sentiment feature label is a key link in the fusion of multi-modal data. Visual information usually covers images, video frames, etc., which contain rich expressions, postures, scenes, etc.; the initial transcription text is the initial text content obtained after transcribing non-text information such as speech, which carries semantic information.

[0036] In the sentiment analysis process, first, the feature words that can reflect the sentiment are extracted from the visual information, for example, through facial expression recognition technology, the feature words related to expressions such as "smile", "frown", and "anger" can be identified. These feature words are directly related to different emotional states, such as "smile" is usually associated with positive and happy emotions, and "frown" and "anger" are usually associated with negative and angry emotions.

[0037] At the same time, from the perspective of scene elements, if the visual information presents a lively party scene, it may contain feature words such as "joy" and "lively", indicating a positive emotional atmosphere; while if it is a solemn scene, it may extract feature words such as "depression", indicating a negative emotional tendency.

[0038] For the initial transcribed text, natural language processing techniques can extract words that directly express emotions such as "happy", "sad", "excited", "depressed" as feature words. Also, by analyzing the tone, rhetorical devices, and other elements that indirectly express emotions in the text, such as exclamatory sentences, rhetorical questions, and other sentence patterns that may intensify emotional expression, corresponding emotional feature words can be extracted.

[0039] Based on the comprehensive visual information and the feature words of the initial transcribed text, machine learning or deep learning algorithms are used for emotion classification and label generation. For example, if the visual information recognizes a "smile" expression and the initial transcribed text contains feature words such as "happy" and "excited", the emotional feature labels "positive" and "pleasure" can be generated. If the visual information presents an "angry" expression and the text contains feature words such as "angry" and "discontent", the emotional feature labels "negative" and "angry" can be generated.

[0040] This multi-modal emotion analysis technique for generating emotional feature labels can fully utilize the advantages of visual and textual data of different modalities, complement and verify each other, and more comprehensively and accurately capture emotional information in the data, providing strong support for subsequent emotional understanding and emotional interaction applications, significantly improving the accuracy and reliability of emotion analysis, and enhancing the system's perception and processing capabilities for complex emotional scenarios.

[0041] S60: Based on the emotional feature labels, the initial transcribed text is optimized and corrected to obtain the target transcribed text.

[0042] Specifically, based on the generated emotional feature labels, the initial transcribed text is optimized and corrected to obtain the target transcribed text, aiming to improve the quality of the text and the accuracy of emotional expression. Emotional feature labels are a comprehensive extraction of the emotions contained in the initial transcribed text and its associated visual information, such as "positive", "negative", "angry", "joy", "sadness", etc. These labels contain information about the overall emotional tendency and emotional intensity of the text.

[0043] In the text optimization correction process, first, the initial transcribed text is analyzed at the semantic level according to the emotional feature label. If the emotional feature label is "positive", and there are words or expressions in the initial transcribed text that are contrary to positive emotions, such as "terrible" and "disappointing", these words or expressions will be considered as objects that need to be corrected. At the same time, attention will also be paid to whether the tone and wording of the text are consistent with the emotional feature label. If the emotional feature label is "angry", but the text expression is relatively flat and mild, it will be adjusted to a more angry expression. In addition, the emotional feature label will also be used to check and optimize the logical coherence of the text. If the emotional feature label indicates that the text should express strong emotions, but the text content has logical jumps or contradictions, resulting in a lack of smooth emotional expression, the text will be logically combed and adjusted. For example, when expressing angry emotions in the initial transcribed text, the cause of the event is first described, then a piece of unrelated content is suddenly inserted, and then the expression of angry emotions is returned. At this time, the irrelevant content will be deleted, making the text logic more coherent and the emotional expression more fluent.

[0044] By performing text optimization correction on the initial transcribed text based on the emotional feature label, the quality of the target transcribed text can be significantly improved. On the one hand, the emotional expression of the text is more accurate and vivid, and can better convey the emotional connotation of the original information; on the other hand, the logical coherence and readability of the text are enhanced, making it easier for readers or subsequent processing systems to understand the text content, providing more reliable and high-quality text data basis for subsequent text analysis, emotional interaction and other applications, effectively improving the accuracy and effectiveness of the entire data processing process.

[0045] In a preferred embodiment, the target speech signal is subjected to segmentation and recognition processing to obtain a speech feature vector, comprising: The target speech signal is subjected to dialect recognition to determine the dialect type of the target speech signal.

[0046] According to the dialect type, the corresponding dialect speech segmenter and dialect vocabulary feature encoding table are extracted from the dialect management repository, and the dialect vocabulary feature encoding table comprises a plurality of dialect vocabularies and corresponding feature encoding values.

[0047] According to the dialect speech segmenter, the target speech signal is subjected to word-level segmentation to obtain a plurality of target dialect vocabularies.

[0048] Based on the dialect vocabulary feature encoding table, the plurality of target dialect vocabularies are subjected to encoding matching to obtain a plurality of target feature encoding values, which constitute the speech feature vector.

[0049] Preferably, the target speech signal is subjected to dialect recognition to determine the dialect type to which the target speech signal belongs before being subjected to segmentation and recognition to obtain the speech feature vector. This process involves comparing the acoustic features of the target speech signal, such as pitch, duration, intensity, and specific pronunciation patterns, with the pre-stored dialect features in the dialect feature library to accurately determine the dialect type, such as Cantonese or Sichuan-Chongqing dialect.

[0050] After determining the dialect type, the corresponding dialect speech segmenter and dialect vocabulary feature encoding table are extracted from the dialect management repository based on the dialect type. The dialect speech segmenter is constructed as follows: based on a sample dialect speech set and a sample dialect segmentation vocabulary set, the sample dialect speech set covers a large number of dialect speech data in different scenarios, different speakers, and different contents; the sample dialect segmentation vocabulary set contains a representative set of vocabulary in the dialect. Machine learning algorithms such as Hidden Markov Model (HMM) or Recurrent Neural Network (RNN) in deep learning and its variants (such as LSTM, GRU) are used to train the sample dialect speech set to learn the boundary features of the vocabulary in the dialect speech, including the rules of speech pauses, tone changes, and pronunciation duration, and to conduct supervised learning in combination with the sample dialect segmentation vocabulary set, so that the model can accurately identify the start and end positions of the dialect vocabulary, and further construct the dialect speech segmenter. The dialect vocabulary feature encoding table contains a plurality of dialect vocabulary and their corresponding feature encoding values, which are the results of abstracting and quantifying the dialect vocabulary in multiple dimensions such as acoustics and semantics. For example, the pronunciation frequency, tone pattern, and vocabulary semantic category of a dialect vocabulary are encoded.

[0051] Subsequently, the target speech signal is subjected to word-level segmentation using the extracted dialect speech segmenter. The dialect speech segmenter analyzes the target speech signal frame by frame based on the learned dialect vocabulary boundary features to identify the start and end positions of each dialect vocabulary, thereby segmenting the target speech signal into a plurality of target dialect vocabularies. For example, if the target speech signal is a dialogue containing Sichuan-Chongqing dialect, the dialect speech segmenter can accurately segment dialect vocabularies such as "wannian" and "ba shi" from the speech signal.

[0052] Finally, the multiple target dialect words segmented are encoded and matched based on the dialect word feature encoding table. The acoustic features of each target dialect word are compared with the word features in the dialect word feature encoding table to find the most matched dialect word and obtain its corresponding feature encoding value. The feature encoding values corresponding to all target dialect words are combined in a certain order to form a speech feature vector. For example, if the segmented target dialect word is "wàide", its corresponding feature encoding value in the dialect word feature encoding table is [0.3, 0.7, 0.1] (this is only an example encoding value), and after combining the encoding values of all target dialect words, a complete speech feature vector is formed.

[0053] Through the above series of operations, the target speech signal can be effectively segmented and recognized to obtain a speech feature vector that accurately reflects the features of the target speech. On the one hand, the use of dialect recognition and dialect speech segmenter can accurately process speech signals of different dialects, overcoming the limitations of speech processing systems in dialect processing and improving the accuracy and adaptability of speech processing. On the other hand, based on the encoding matching of the dialect word feature encoding table, the speech signal can be converted into an encoding vector with clear semantic and acoustic features, providing high-quality feature input for subsequent speech recognition, semantic understanding, speech synthesis, etc., which helps to improve the performance and effect of the entire speech processing system.

[0054] In a preferred embodiment, the text context information is feature extracted to obtain a text auxiliary feature vector, including: According to the dialect type, a corresponding dialect word semantic annotation table is retrieved, and the dialect word semantic annotation table includes multiple dialect words and corresponding multiple semantic identifiers, each semantic identifier having a semantic feature encoding value.

[0055] Context key words are extracted from the text context information to construct a context key word set.

[0056] Based on the context key word set and the dialect word semantic annotation table, a target semantic identifier that conforms to the current context is matched for each target dialect word, and a corresponding semantic feature encoding value is obtained.

[0057] The semantic feature encoding values of each target dialect word are arranged in the same time sequence as the multiple target feature encoding values to form the text auxiliary feature vector.

[0058] Specifically, extracting features from text context information to obtain auxiliary feature vectors is a key step in enhancing semantic understanding and improving processing accuracy. First, based on the identified dialect type, the corresponding dialect vocabulary semantic annotation table is retrieved from the dialect knowledge base. This dialect vocabulary semantic annotation table is a constructed dataset containing multiple dialect words and multiple semantic identifiers corresponding to each word. Each semantic identifier is assigned a specific semantic feature encoding value. For example, in the Cantonese dialect vocabulary semantic annotation table, the word "行" (travel) may have multiple semantic identifiers, such as "to walk," "can," and "industry." Each semantic identifier corresponds to a different semantic feature encoding value. For example, the encoding value of the semantic identifier "to walk" is [0.1, 0.5, 0.3] (this is just an example encoding value), and the encoding value of the semantic identifier "can" is [0.3, 0.2, 0.6], etc.

[0059] Next, contextual keywords are extracted from the identified text context information to construct a contextual keyword set. This process is achieved through natural language processing techniques, such as word frequency statistics and keyword extraction algorithms (such as the TF-IDF algorithm). For example, if the text context information is "It's so convenient to go shopping in Nidu," the keyword extraction algorithm can extract contextual keywords such as "Nidu" (here), "Xingjie" (shopping), and "convenient" to construct a contextual keyword set.

[0060] Then, based on the constructed context keyword set and the retrieved dialect vocabulary semantic annotation table, the target dialect vocabulary previously segmented and identified is matched with a target semantic identifier that matches the current context and obtains the corresponding semantic feature encoding value. This process requires comprehensive consideration of the target dialect vocabulary's position in the text, the semantic associations of the contextual key words, and the semantic information in the dialect vocabulary semantic annotation table. For example, for the target dialect vocabulary "行," combined with the context keyword "街," it can be determined that the semantic meaning of "行" in the current context is more inclined to "走" (walk). Therefore, it is matched with the target semantic identifier "走" (walk) and the corresponding semantic feature encoding value [0.1, 0.5, 0.3] is obtained.

[0061] Finally, the semantic feature encoding values of each target dialect word are arranged in the same time sequence as the obtained multiple target feature encoding values. Since the speech signal is processed in stages, each target dialect word has a corresponding time position in the speech signal, so arranging the semantic feature encoding values in the same time sequence can ensure the consistency of the text auxiliary feature vector and the speech feature vector in the time dimension. For example, if the first target feature encoding value in the speech feature vector corresponds to the target dialect word at time point t1, then the first semantic feature encoding value in the text auxiliary feature vector also corresponds to the target dialect word at time point t1. After arranging the semantic feature encoding values of all target dialect words in this way, the text auxiliary feature vector is formed.

[0062] The text auxiliary feature vector is obtained through the above steps. On the one hand, the text context information is fully utilized, which provides more context clues for the semantic understanding of dialect words, effectively solves the problem of polysemy of dialect words, and can more accurately determine the specific meaning of each target dialect word in the current context, thereby improving the accuracy of semantic labeling. On the other hand, the semantic feature encoding values are arranged in the same time sequence as the target feature encoding values to form the text auxiliary feature vector, so that the text auxiliary feature vector and the speech feature vector remain consistent in structure and time sequence, which facilitates subsequent feature fusion and model processing, and helps to improve the performance of the entire speech and text integrated processing system. For example, in speech recognition, semantic understanding, machine translation and other tasks, more accurate and comprehensive feature input can be provided, thereby significantly improving the accuracy and robustness of the system.

[0063] In a preferred embodiment, the speech feature vector and the text auxiliary feature vector are subjected to multi-modal fusion processing to generate a fusion feature representation, comprising: According to the speech feature vector or the text auxiliary feature vector, a vector index array is determined.

[0064] The first vector index is obtained by traversing the vector index array.

[0065] According to the first vector index, feature extraction is performed in the speech feature vector and the text auxiliary feature vector to obtain a first speech feature and a first text auxiliary feature.

[0066] The first speech feature and the first text auxiliary feature are spliced and combined to obtain a first fusion feature encoding.

[0067] The first fusion feature encoding is added to the fusion feature representation.

[0068] Further, to fully integrate the speech and text information and improve the comprehensive understanding ability of the target speech signal and its associated text, the speech feature vector and the text auxiliary feature vector need to be processed by multi-modal fusion to generate a fusion feature representation.

[0069] Here, the speech feature vector and the text auxiliary feature vector have the same length and are one-to-one correspondence, for example, the first position of the vector stores the speech code of a certain dialect vocabulary (the code is the result of quantitatively representing the acoustic features of the dialect vocabulary in the speech signal, such as a combination of numerical values containing pitch, duration, spectral features, etc.) and the semantic code of the vocabulary (the code is a digital abstraction of the semantics of the dialect vocabulary in a specific context, obtained based on the dialect vocabulary semantic annotation table), the second position stores the speech code and semantic code of the next dialect vocabulary, and so on. Based on this feature, the vector index array is determined according to the speech feature vector or the text auxiliary feature vector. The vector index array is an ordered index set, and the number of elements is the same as the length of the vector. Each index corresponds to a position in the vector, for example, the vector index array can be [0, 1, 2,..., n-1] (assuming the length of the vector is n), which is used to access the features in the vector one by one.

[0070] Subsequently, the first vector index is obtained by traversing the vector index array. For example, in the first traversal, the first vector index obtained is 0, which corresponds to the first position in the vector.

[0071] Further, according to the first vector index obtained, feature extraction is performed in the speech feature vector and the text auxiliary feature vector. The first speech feature is extracted from the first position of the speech feature vector, which is the speech code of the dialect vocabulary stored in that position; at the same time, the first text auxiliary feature is extracted from the first position of the text auxiliary feature vector, which is the semantic code of the dialect vocabulary stored in that position. The first speech feature and the first text auxiliary feature are combined by concatenation. The concatenation operation is to connect two feature vectors in dimension, for example, if the first speech feature is an m-dimensional vector and the first text auxiliary feature is a k-dimensional vector, the first fusion feature code obtained after concatenation is an (m+k) -dimensional vector, which contains both speech and semantic information.

[0072] Finally, the first fusion feature code obtained is added to the fusion feature representation. The fusion feature representation is a set for storing all fused features. As the vector index array is traversed, the fusion feature code corresponding to each position is constantly added to it, and finally a complete fusion feature representation is formed. For example, when the entire vector index array is traversed, the fusion feature representation contains the speech and semantic fusion features of all dialect vocabularies.

[0073] Through the above-mentioned multimodal fusion processing, the generated fused feature representation, on the one hand, fully utilizes the information of two different modalities, speech and text, organically combining the acoustic features of speech with the semantic features of text, overcoming the limitations of single-modal information in expression, and can more comprehensively and accurately describe the characteristics of the target speech signal and its associated text. For example, in the task of dialect speech recognition, relying solely on speech features may not be able to accurately distinguish dialect words with similar pronunciations but different semantics. However, after integrating text auxiliary features, the speech features can be corrected and supplemented based on semantic information, thereby improving recognition accuracy. On the other hand, the generated fused feature representation provides subsequent machine learning or deep learning models with richer and more discriminative feature inputs, helping to improve the model's performance in tasks such as speech understanding, semantic analysis, and emotion recognition, and enhancing the system's ability to process complex multimodal information.

[0074] In a preferred embodiment, performing speech transcription based on the fused feature representation to obtain an initial transcription text includes: A text conversion model corresponding to the dialect type is extracted, where the text conversion model is used to convert the fusion feature encoding of the dialect vocabulary into standard Mandarin vocabulary.

[0075] Each fusion feature code in the fusion feature representation is sequentially input into the text conversion model to obtain a plurality of initial text words.

[0076] The multiple initial text words are aggregated to obtain the initial transcribed text.

[0077] Specifically, to accurately convert speech containing dialect information into standard Mandarin text, speech transcription is performed based on a fused feature representation to produce an initial transcribed text. First, a text conversion model corresponding to the identified dialect type is extracted from a model library. This text conversion model is trained using a large amount of dialect and standard Mandarin corpus. Its core function is to convert the fused feature encoding of dialect vocabulary into standard Mandarin vocabulary. This fused feature encoding is derived through a multimodal fusion process, combining the phonetic features of dialect vocabulary (such as quantized representations of acoustic features such as pitch, duration, and spectrum) and semantic features (digital abstractions of semantics derived from a semantic annotation table for dialect vocabulary). For example, for the Sichuan-Chongqing dialect word "ba shi," its fused feature encoding encompasses both the word's pronunciation characteristics in the Sichuan-Chongqing dialect and its semantic information in a specific context. The trained text conversion model can recognize this specific fused feature encoding and accurately convert it into the corresponding standard Mandarin word "comfortable."

[0078] Then, each fusion feature code in the fusion feature representation is input into the extracted text conversion model in turn. The model processes each fusion feature code, analyzes and maps the features through an internal neural network structure (such as a recurrent neural network, a Transformer, etc. in deep learning), and outputs the corresponding initial text vocabulary. For example, the fusion feature representation sequentially contains fusion feature codes corresponding to dialect words such as “basi” and “yaode”. After inputting these codes into the model in turn, the model outputs standard Mandarin words such as “shufu” and “keyi” as initial text vocabulary.

[0079] Finally, the obtained multiple initial text vocabularies are summarized. The summary operation is to arrange and combine the initial text vocabularies obtained by conversion in turn according to the original order of the fusion feature codes in the fusion feature representation, to form a complete text sequence, i.e. an initial transcribed text. For example, if the fusion feature codes correspond to dialect words such as “basi”, “yaode”, and “anyi” in turn, the initial text vocabularies obtained after conversion are “shufu”, “keyi”, and “qini”, and after summarizing these vocabularies in order, the initial transcribed text obtained is “shufu keyi qini”. In actual application, appropriate space or punctuation processing will be performed according to semantic and grammatical rules, so that it is more in line with the habit of written expression.

[0080] Through the above speech transcription operation based on the fusion feature representation, the accuracy and standardization of speech transcription can be significantly improved. On the one hand, the text conversion model makes full use of the speech and semantic information contained in the fusion feature codes, overcomes the semantic ambiguity problem that may occur when relying solely on speech features for transcription, and improves the accuracy of dialect vocabulary conversion. For example, for some dialect words with similar pronunciation but different semantics, the model can accurately distinguish and convert them according to semantic features. On the other hand, the generated initial transcribed text is in standard Mandarin form, which is convenient for subsequent text processing, analysis and application, such as text classification, sentiment analysis, machine translation, etc. in natural language processing tasks, providing convenience for cross-dialect speech information exchange and processing, and enhancing the universality and practicality of the system.

[0081] In a preferred embodiment, the construction step of the text conversion model comprises: obtaining a pre-trained language model as a base model, the pre-trained language model having natural language understanding and generation capabilities.

[0082] constructing a sample fusion feature code set based on the dialect vocabulary feature code table and the dialect vocabulary semantic annotation table, and annotating the sample fusion feature code set with standard Mandarin vocabulary to obtain a sample standard vocabulary set.

[0083] The pre-trained language model is fine-tuned using the sample fusion feature encoding set as input features and the sample standard vocabulary set as supervision labels to obtain the text conversion model.

[0084] Optionally, you can first obtain a pre-trained language model as a base model. Pre-trained language models are trained on massive amounts of text data and possess powerful natural language understanding and generation capabilities. For example, large language models like Doubao and DeepSeek acquire a wealth of linguistic knowledge during training, including vocabulary semantics, grammatical structure, and contextual relationships. They can deeply understand and analyze input text and generate text that conforms to linguistic rules. These models provide a solid linguistic foundation and powerful feature extraction capabilities for subsequent dialect vocabulary conversion tasks.

[0085] Next, a sample fusion feature code set is constructed based on the dialect vocabulary feature coding table and the dialect vocabulary semantic annotation table. Standard Mandarin vocabulary is annotated on this set to produce a sample standard vocabulary set. The dialect vocabulary feature coding table contains multiple dialect words and their corresponding feature code values. These code values ​​are the result of quantifying the acoustic characteristics of dialect words in speech signals, such as a numerical combination of pitch, duration, and spectral characteristics. The dialect vocabulary semantic annotation table contains dialect words and their corresponding multiple semantic identifiers and semantic feature code values, which are used to describe the semantic connotations of dialect words in different contexts. By fusing the phonetic feature codes and semantic feature codes of dialect words, a sample fusion feature code set is constructed. For example, for the Cantonese word "帅仔," its phonetic feature code may be a specific set of numerical values, while its semantic feature code also has corresponding numerical values ​​based on its semantic meaning in different contexts (such as "帅哥"). Fusing these two sets of codes yields a sample fusion feature code for the word. Then, these sample fusion feature codes are used to annotate standard Mandarin vocabulary, that is, to determine the standard Mandarin vocabulary corresponding to each dialect vocabulary, such as the standard Mandarin vocabulary corresponding to "漂亮仔" is "帅哥". All annotated standard Mandarin vocabulary constitutes the sample standard vocabulary set.

[0086] Finally, the constructed sample fusion feature encoding set is taken as the input feature, and the sample standard vocabulary set is taken as the supervised label to fine-tune the pre-trained language model. Fine-tuning is a process of adjusting and optimizing parameters of the pre-trained language model for specific dialect vocabulary conversion tasks. During the training process, the sample fusion feature encoding set is input into the pre-trained language model, and the model processes and analyzes the input feature according to its internal structure and parameters, and outputs the predicted vocabulary. Then, the predicted vocabulary is compared with the supervised label in the sample standard vocabulary set, the loss function is calculated, and the parameters of the model are adjusted through the back propagation algorithm, so that the output of the model is as close as possible to the supervised label. After multiple rounds of iterative training, the model gradually learns the mapping relationship between the dialect vocabulary fusion feature encoding and the standard Mandarin vocabulary, and finally obtains the text conversion model.

[0087] The text conversion model obtained through the above construction steps, on the one hand, utilizes the powerful language understanding and generation capability of the pre-trained language model, which can quickly adapt to the dialect vocabulary conversion task and reduce the large amount of time and computing resources required for training the model from scratch. On the other hand, the sample fusion feature encoding set and the sample standard vocabulary set constructed based on the dialect vocabulary feature encoding table and the dialect vocabulary semantic annotation table provide the model with rich and accurate training data, so that the model can accurately learn the corresponding relationship between the dialect vocabulary and the standard Mandarin vocabulary, and improve the accuracy and reliability of the dialect vocabulary conversion. In practical applications, the text conversion model can effectively convert speech containing dialect information into standard Mandarin text, promoting information exchange and understanding between different dialect regions.

[0088] In a preferred embodiment, sentiment analysis is performed according to the visual information and the initial transcribed text to generate a sentiment feature label, including: Image preprocessing is performed on the visual information to extract facial expression features and body movement features.

[0089] Based on the facial expression features and body movement features, a sentiment state is identified to obtain a visual sentiment identifier.

[0090] Text sentiment analysis is performed on the initial transcribed text to extract sentiment tendencies in the text semantics to obtain a text sentiment identifier.

[0091] The visual sentiment identifier and the text sentiment identifier are fused to generate the sentiment feature label.

[0092] To comprehensively and accurately capture the emotions expressed by the target object, exemplary, the visual information and the initial transcribed text are used to generate emotion feature labels. First, image preprocessing is performed on the visual information to extract facial expression features and body movement features. The image preprocessing process includes multiple steps, such as image denoising, grayscale conversion, normalization, etc., aiming to improve image quality, reduce noise interference, and make subsequent feature extraction more accurate. When extracting facial expression features, computer vision techniques are used, such as face detection algorithms based on deep learning (e.g., MTCNN) to locate the face region, and then feature extraction networks (e.g., convolutional neural networks CNN) are used to analyze the position changes of facial key points, such as the raising or lowering of eyebrows, the widening or narrowing of eyes, the opening or closing of mouth, etc. The combination of these key point changes can reflect different facial expressions, such as happiness, sadness, anger, etc. For example, when the eyebrows are raised, the eyes are wide open, and the corners of the mouth are raised, it may indicate a happy emotion. For body movement features, target detection and pose estimation techniques are used to identify the body parts (such as head, limbs, torso) of the target object and their position and posture changes. For example, arm waving, body leaning, footstep moving, etc. actions may contain emotional information, such as waving hands may indicate welcome or goodbye, body leaning forward may indicate attention or excitement.

[0093] Then, based on the extracted facial expression features and body movement features, the emotional state is identified to obtain visual emotion labels. This process is achieved with machine learning or deep learning models, such as support vector machines (SVM), recurrent neural networks (RNN) and their variants (such as LSTM, GRU), etc. The model learns from a large amount of facial expression and body movement data labeled with emotional states, establishing a mapping relationship between features and emotional states. For example, a trained model can determine which of the happy, sad, angry, surprised, etc. emotional states the target object is currently in according to the combination of facial expression features and body movement features, and express the emotional state with specific labels, such as using numerical codes "1" for happy, "2" for sad, etc. These labels are visual emotion labels.

[0094] Next, sentiment analysis is performed on the initial transcribed text to extract the sentiment within the text's semantics and generate a text sentiment marker. Sentiment analysis primarily utilizes natural language processing techniques. The initial transcribed text undergoes preprocessing operations such as word segmentation, part-of-speech tagging, and named entity recognition to break the text down into basic linguistic units and annotate their parts of speech and semantic information. Next, sentiment lexicons or deep learning models (such as the Transformer-based pre-trained language model BERT) are used to analyze the sentiment polarity (e.g., positive, negative, neutral) and intensity of the words in the text. For example, words like "like," "happy," and "beautiful" typically convey positive sentiment, while words like "hate," "sad," and "bad" convey negative sentiment. By comprehensively analyzing the sentiment polarity and intensity of all words in the text, the overall sentiment of the text is determined and represented by specific markers, such as "+" for positive sentiment, "-" for negative sentiment, and "0" for neutral sentiment. These markers are referred to as text sentiment markers.

[0095] Finally, the visual and textual emotion labels are fused and judged to generate an emotional feature label. This fusion judgment process can employ various strategies, such as weighted averaging and decision fusion. For example, different weights can be assigned to the visual and textual emotion labels based on their importance in expressing emotion. A weighted average is then calculated to generate a comprehensive emotion score, and the final emotional feature label is determined based on the score range. Alternatively, preliminary emotional judgments can be made based on the visual and textual emotion labels separately, and then fused using certain rules (such as majority voting) to determine the final emotional feature label. For example, if the visual emotion label is happy (coded "1") and the textual emotion label is positive ("+"), the fusion judgment may generate an emotional feature label such as "positive happy."

[0096] The emotional feature labels generated through the aforementioned sentiment analysis process based on visual information and initial transcriptions combine information from both visual and textual modalities, overcoming the emotional expression limitations of single-modal information and enabling a more comprehensive and accurate capture of the target subject's emotional state. For example, in some cases, the target subject may appear to be smiling (visual information indicates happiness), while the textual expression reveals negativity. By integrating information from both modalities, their true emotions can be more accurately judged. Furthermore, the generated emotional feature labels provide important feature input for subsequent tasks such as sentiment classification and prediction, helping to improve the performance and accuracy of sentiment analysis systems and providing strong support for applications in areas such as intelligent interaction, public opinion monitoring, and psychological analysis.

[0097] In a preferred embodiment, based on the emotional feature label, the initial transcription text is subjected to text optimization correction to obtain a target transcription text, including: According to the emotional feature label, a corresponding emotional correction rule set is determined, and the emotional correction rule set includes word expression optimization rules under different emotional states.

[0098] Based on the emotional correction rule set, the words in the initial transcription text are subjected to emotional adaptability detection, and words to be adjusted that do not match the current emotional state are identified.

[0099] The words to be adjusted are subjected to emotional guidance word replacement to generate adjusted words after emotional correction, and the corresponding words in the initial transcription text are replaced by the adjusted words to obtain the target transcription text.

[0100] Specifically, in order to make the initial transcription text more suitable for emotional expression needs and improve the quality of the text and the emotional fit, the initial transcription text needs to be subjected to text optimization correction based on the emotional feature label to obtain a target transcription text. First, according to the generated emotional feature label, a corresponding emotional correction rule set is determined. The emotional feature label is an abstract representation of the emotional state of the target object, such as "positive happy", "negative sad", "angry excited", etc. The emotional correction rule set is pre-constructed and contains word expression optimization rules under different emotional states. These rules are based on a large number of language samples and emotional analysis research and specify which words are appropriate and which words may not match the current emotional state under different emotional states. For example, under the emotional state of "positive happy", the rules may specify that words with positive emotional color such as "wonderful", "excellent", "ecstatic" should be used first; while under the emotional state of "negative sad", words that are too happy and positive should be avoided, and words such as "painful", "disappointed", "sorrowful" are preferred.

[0101] Then, based on the determined emotional correction rule set, the words in the initial transcription text are subjected to emotional adaptability detection. This process compares each word in the initial transcription text with the emotional correction rule set to determine whether it meets the word expression requirements under the current emotional state. For example, if the emotional feature label is "positive happy" and the word "terrible" appears in the initial transcription text, according to the emotional correction rule set, "terrible" is a negative emotional word and does not match the current positive happy emotional state, so the word is identified as a word to be adjusted.

[0102] Then, the identified to-be-adjusted words are subjected to emotion-oriented word replacement. The emotion-oriented word replacement refers to selecting appropriate words from an emotion correction rule set or a pre-constructed word library to replace the to-be-adjusted words according to the current emotion state. The word library stores a large number of words with different emotional colors and classifies them according to emotion types. For example, for the to-be-adjusted word "terrible", in the "positive happy" emotion state, positive words such as "excellent" and "wonderful" can be selected from the word library to replace the to-be-adjusted word, generating the emotion-corrected adjusted word.

[0103] Finally, the corresponding words in the initial transcribed text are replaced by the adjusted words to obtain the target transcribed text. The replacement operation is to accurately replace the to-be-adjusted words in the initial transcribed text with the adjusted words, ensuring that the replaced text is coherent and correct in grammar and semantics. For example, the initial transcribed text is "this performance is terrible, and people are very disappointed", and after emotion-oriented word replacement, the target transcribed text becomes "this performance is excellent, and people are very excited".

[0104] Through the above text optimization correction operation based on the emotion feature label, on the one hand, the target transcribed text can more accurately express the emotion state of the target object, enhancing the emotional appeal and expressiveness of the text. For example, in application scenarios such as emotional dialogue systems and film and television subtitle generation, the optimized and corrected text can better convey the emotions of the characters, improving the user experience. On the other hand, the quality and readability of the text are improved, avoiding semantic ambiguity or inaccurate emotional expression caused by improper use of words, and providing a more reliable data foundation for subsequent text processing and analysis tasks.

[0105] In a preferred embodiment, the emotion adaptation detection of the words in the initial transcribed text based on the emotion correction rule set identifies to-be-adjusted words that do not match the current emotion state, including: Extracting word emotion adaptation detection rules from the emotion correction rule set, the word emotion adaptation detection rules including word expression adaptation rules and word expression disabling rules corresponding to the current emotion feature label.

[0106] According to the word expression adaptation rules, the words in the initial transcribed text are scanned word by word to identify abnormal words that do not conform to the adaptation rules.

[0107] According to the word expression disabling rules, the words in the initial transcribed text are screened word by word to identify conflict words that conform to the disabling rules.

[0108] The abnormal words and the conflict words are combined to constitute the to-be-adjusted words.

[0109] Optionally, first, according to the generated emotional feature label "excitement / positivity", the corresponding lexical emotional adaptability detection rules are extracted from the emotional correction rule set, including lexical expression adaptation rules and lexical expression prohibition rules. Among them, the adaptation rules clearly specify the lexical features that should be used in the excited emotional state, such as affirmative words "really", "very", "especially", which can strengthen the certainty of the sentence; positive words "very good", "not bad", "satisfied", which are used to directly express positive evaluation; and action words "as soon as possible", "immediately", "start", which embody the attitude of positive action. And the prohibition rules list the lexical features that should be avoided in the excited emotional state, such as uncertain words "maybe", "perhaps", "probably", which will weaken the certainty of the sentence; hesitant words "try", "see", "consider", which show hesitation in action; and conservative expressions "fair", "average", which lack positive emotional color.

[0110] Next, the initial transcribed text, such as "this thing really may be fair, we can try it", is detected word by word. In the adaptation rule detection phase, scanning finds that "maybe" does not meet the requirement of affirmation, "fair" does not meet the requirement of positive expression, and "try" does not meet the requirement of action, which are marked as abnormal words. In the prohibition rule detection phase, it is found that "maybe", "perhaps" meet the prohibition rule of uncertain words, "try" meets the prohibition rule of hesitant words, and "fair" meets the prohibition rule of conservative expressions, which are marked as conflict words.

[0111] Further, the abnormal words and conflict words are combined to form a list of words to be adjusted ["maybe", "fair", "perhaps", "try"]. Finally, according to the emotional guidance lexical replacement strategy, these words to be adjusted are replaced with words that meet the excited emotional state, such as "very good" replacing "fair", "as soon as possible" replacing "try", and "maybe" and "perhaps" are similar in function, so one of them can be selected for replacement. The optimized target transcribed text is "this thing really very good, we start as soon as possible!".

[0112] This process ensures the high consistency of the transcribed text with the original speech signal in emotional expression through fine emotional adaptability detection and optimization, significantly improves the emotional appeal and expressiveness of the text, and enhances the quality and readability of the text, providing a more reliable data foundation for subsequent text processing and analysis tasks.

[0113] The speech recognition and transcription method based on multi-modal fusion and emotional analysis provided by the embodiment of the application has at least the following technical effects: 1. By dialect recognition to determine the dialect type, using the corresponding dialect speech segmenter and vocabulary feature coding table to perform word-level segmentation and coding matching on the target speech signal, and with the help of the dialect vocabulary semantic annotation table, the semantic feature coding value of the text context information is extracted, the text auxiliary feature vector is constructed, and finally the accurate conversion of dialect vocabulary to standard Chinese vocabulary is realized, effectively solving the dialect speech transcription problem and improving the transcription accuracy.

[0114] 2. The speech feature vector and the text auxiliary feature vector are processed by multi-modal fusion, the vector index array is determined, the features are extracted and spliced to generate a fusion feature representation, which fully utilizes the speech and text context information and mines the correlation between different modal information, providing a more comprehensive and rich feature basis for subsequent speech transcription and improving the performance of the transcription system.

[0115] 3. Based on visual information and initial transcription text, sentiment analysis is performed to generate sentiment feature labels, and according to the labels, a sentiment correction rule set is determined to detect and replace the words in the initial transcription text according to the sentiment adaptability, so that the target transcription text is more suitable for emotional expression needs, enhancing the emotional appeal and expressiveness of the text, improving the text quality and readability, and providing more reliable data for subsequent text processing and analysis.

[0116] Embodiment two: As Figure 2 shown, based on the same inventive concept as the speech recognition and transcription method based on multi-modal fusion and sentiment analysis provided in embodiment one, the present embodiment also provides a speech recognition and transcription system based on multi-modal fusion and sentiment analysis, which comprises: An information acquisition module 11 is configured to acquire a target speech signal and synchronously acquire auxiliary modal information corresponding to the target speech signal, wherein the auxiliary modal information includes visual information and text context information.

[0117] A signal segmentation module 12 is configured to perform segmentation and recognition processing on the target speech signal to obtain a speech feature vector, and perform feature extraction on the text context information to obtain a text auxiliary feature vector.

[0118] A fusion processing module 13 is configured to perform multi-modal fusion processing on the speech feature vector and the text auxiliary feature vector to generate a fusion feature representation.

[0119] A speech transcription module 14 is configured to perform speech transcription based on the fusion feature representation to obtain an initial transcription text.

[0120] A sentiment analysis module 15 is configured to perform sentiment analysis according to the visual information and the initial transcription text to generate sentiment feature labels.

[0121] The optimization correction module 16 is configured to perform text optimization correction on the initial transcribed text based on the emotion feature label to obtain a target transcribed text.

[0122] Further, the signal segmentation module 12 is further configured to perform the following steps: The dialect recognition is performed on the target voice signal to determine a dialect type of the target voice signal; the corresponding dialect speech segmenter and dialect vocabulary feature coding table are extracted from the dialect management repository according to the dialect type, the dialect vocabulary feature coding table including a plurality of dialect vocabularies and corresponding feature coding values; the target voice signal is segmented at the word level according to the dialect speech segmenter to obtain a plurality of target dialect vocabularies; and the plurality of target dialect vocabularies are matched and coded based on the dialect vocabulary feature coding table to obtain a plurality of target feature coding values, which constitute the voice feature vector.

[0123] Further, the signal segmentation module 12 is further configured to perform the following steps: The corresponding dialect vocabulary semantic labeling table is called according to the dialect type, the dialect vocabulary semantic labeling table including a plurality of dialect vocabularies and corresponding a plurality of semantic identifiers, each of the semantic identifiers having a semantic feature coding value; the context key vocabularies are extracted from the text context information to construct a context key vocabulary set; the target semantic identifier conforming to the current context is matched for each of the target dialect vocabularies based on the context key vocabulary set and the dialect vocabulary semantic labeling table, and the corresponding semantic feature coding value is obtained; and the semantic feature coding values of each of the target dialect vocabularies are arranged in the same time sequence as the plurality of target feature coding values to constitute the text auxiliary feature vector.

[0124] Further, the fusion processing module 13 is further configured to perform the following steps: The vector index array is determined according to the voice feature vector or the text auxiliary feature vector; the first vector index is obtained by traversing the vector index array; the first voice feature and the first text auxiliary feature are obtained by feature extraction in the voice feature vector and the text auxiliary feature vector according to the first vector index; the first fusion feature coding is obtained by splicing and combining the first voice feature and the first text auxiliary feature; and the first fusion feature coding is added to the fusion feature representation.

[0125] Further, the voice transcribing module 14 is further configured to perform the following steps: extract a text conversion model corresponding to the dialect type, the text conversion model being used to convert the fusion feature code of the dialect vocabulary into a standard Mandarin vocabulary; input each fusion feature code in the fusion feature representation into the text conversion model in sequence to obtain a plurality of initial text vocabularies; and aggregate the plurality of initial text vocabularies to obtain the initial transcribed text.

[0126] Further, the speech transcribing module 14 is further configured to perform the following steps: obtain a pre-trained language model as a base model, the pre-trained language model having natural language understanding and generation capabilities; construct a sample fusion feature code set based on a dialect vocabulary feature code table and a dialect vocabulary semantic annotation table, and perform standard Mandarin vocabulary annotation on the sample fusion feature code set to obtain a sample standard vocabulary set; and fine-tune the pre-trained language model by taking the sample fusion feature code set as input features and the sample standard vocabulary set as supervision labels to obtain the text conversion model.

[0127] Further, the sentiment analysis module 15 is further configured to perform the following steps: perform image preprocessing on the visual information to extract facial expression features and body movement features; identify a sentiment state based on the facial expression features and the body movement features to obtain a visual sentiment identifier; perform text sentiment analysis on the initial transcribed text to extract a sentiment tendency in a text semantic to obtain a text sentiment identifier; and fuse the visual sentiment identifier and the text sentiment identifier to generate the sentiment feature label.

[0128] Further, the optimization correction module 16 is further configured to perform the following steps: determine a corresponding sentiment correction rule set according to the sentiment feature label, the sentiment correction rule set including vocabulary expression optimization rules under different sentiment states; perform sentiment adaptability detection on the vocabularies in the initial transcribed text based on the sentiment correction rule set to identify to-be-adjusted vocabularies that do not match a current sentiment state; perform sentiment-oriented vocabulary replacement on the to-be-adjusted vocabularies to generate adjusted vocabularies after sentiment correction; and replace corresponding vocabularies in the initial transcribed text with the adjusted vocabularies to obtain the target transcribed text.

[0129] Further, the optimization correction module 16 is further configured to perform the following steps: extract a lexical sentiment adaptability detection rule from the sentiment correction rule set, the lexical sentiment adaptability detection rule including a lexical expression adaptation rule and a lexical expression disabling rule corresponding to the current sentiment feature label; performing word-by-word scanning on the lexical expressions in the initial transcribed text according to the lexical expression adaptation rule to identify abnormal lexical expressions that do not conform to the adaptation rule; performing word-by-word screening on the lexical expressions in the initial transcribed text according to the lexical expression disabling rule to identify conflict lexical expressions that conform to the disabling rule; and combining the abnormal lexical expressions and the conflict lexical expressions to form the to-be-adjusted lexical expressions.

[0130] The present specification describes the method for speech recognition and transcription based on multi-modal fusion and sentiment analysis in detail. Based on the speech recognition and transcription system based on multi-modal fusion and sentiment analysis in the embodiments, the system disclosed in the embodiments is described simply because it corresponds to the method disclosed in the embodiments, and the relevant part can be found in the method part.

[0131] The above description of disclosed embodiments enables a person skilled in the art to implement or use the present application. Various modifications to the embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A speech recognition transcription method based on multimodal fusion and sentiment analysis, characterized in that: include: Acquire a target speech signal and simultaneously collect auxiliary modal information corresponding to the target speech signal, the auxiliary modal information including visual information and textual context information; Performing segmentation and recognition processing on the target speech signal to obtain a speech feature vector, and performing feature extraction on the text context information to obtain a text auxiliary feature vector; Performing multimodal fusion processing on the speech feature vector and the text auxiliary feature vector to generate a fused feature representation; Perform speech transcription based on the fused feature representation to obtain an initial transcription text; Performing sentiment analysis based on the visual information and the initial transcribed text to generate a sentiment feature label; Based on the emotional feature label, the initial transcription text is optimized and corrected to obtain a target transcription text.

2. The method according to claim 1, characterized in that The target speech signal is segmented and recognized to obtain a speech feature vector, including: Performing dialect recognition on the target speech signal to determine the dialect type of the target speech signal; Extracting a corresponding dialect speech segmenter and a dialect vocabulary feature coding table from a dialect management repository according to the dialect type, wherein the dialect vocabulary feature coding table includes a plurality of dialect words and corresponding feature coding values; performing word-level segmentation on the target speech signal according to the dialect speech segmenter to obtain a plurality of target dialect words; The plurality of target dialect words are coded and matched based on the dialect vocabulary feature coding table to obtain a plurality of target feature coding values ​​to form the speech feature vector.

3. The method according to claim 2, characterized in that Performing feature extraction on the text context information to obtain a text auxiliary feature vector includes: Retrieving a corresponding dialect vocabulary semantics annotation table according to the dialect type, wherein the dialect vocabulary semantics annotation table includes a plurality of dialect words and a plurality of corresponding semantic identifiers, each of the semantic identifiers having a semantic feature coding value; Extracting context key words from the text context information to construct a context keyword set; Based on the context keyword set and the dialect vocabulary semantic annotation table, matching the target semantic identifier that conforms to the current context for each target dialect vocabulary, and obtaining the corresponding semantic feature coding value; The semantic feature coding values ​​of each of the target dialect words are arranged in the same time sequence as the multiple target feature coding values ​​to form the text auxiliary feature vector.

4. The method according to claim 1, wherein Performing multimodal fusion processing on the speech feature vector and the text auxiliary feature vector to generate a fusion feature representation includes: Determine a vector index array according to the speech feature vector or the text auxiliary feature vector; Traverse the vector index array to obtain the first vector index; Performing feature extraction in the speech feature vector and the text auxiliary feature vector according to the first vector index to obtain a first speech feature and a first text auxiliary feature; Concatenate the first speech feature and the first text auxiliary feature to obtain a first fusion feature code; The first fused feature encoding is added to the fused feature representation.

5. The method according to claim 2, characterized in that Performing speech transcription based on the fused feature representation to obtain an initial transcription text includes: Extracting a text conversion model corresponding to the dialect type, wherein the text conversion model is used to convert the fusion feature encoding of the dialect vocabulary into standard Mandarin vocabulary; Inputting each fusion feature code in the fusion feature representation into the text conversion model in sequence to obtain a plurality of initial text words; The multiple initial text words are aggregated to obtain the initial transcribed text.

6. The method according to claim 5, characterized in that The steps of constructing the text conversion model include: Obtaining a pre-trained language model as a base model, wherein the pre-trained language model has natural language understanding and generation capabilities; Constructing a sample fusion feature coding set based on the dialect vocabulary feature coding table and the dialect vocabulary semantic annotation table, and performing standard Mandarin vocabulary annotation on the sample fusion feature coding set to obtain a sample standard vocabulary set; The pre-trained language model is fine-tuned using the sample fusion feature encoding set as input features and the sample standard vocabulary set as supervision labels to obtain the text conversion model.

7. The method according to claim 1, characterized in that Performing sentiment analysis based on the visual information and the initial transcribed text to generate sentiment feature labels includes: Performing image preprocessing on the visual information to extract facial expression features and body movement features; Based on the facial expression features and body movement features, identifying the emotional state and obtaining a visual emotion identifier; Performing text sentiment analysis on the initial transcribed text to extract the sentiment tendency in the text semantics and obtain a text sentiment identifier; The visual emotion identifier and the text emotion identifier are fused and judged to generate the emotion feature label.

8. The method according to claim 1, characterized in that Based on the emotional feature label, the initial transcription text is optimized and corrected to obtain a target transcription text, including: Determining a corresponding emotion correction rule set according to the emotion feature label, wherein the emotion correction rule set includes vocabulary expression optimization rules under different emotion states; Performing sentiment adaptability testing on the vocabulary in the initial transcription text based on the sentiment correction rule set, and identifying the vocabulary to be adjusted that does not match the current sentiment state; Performing sentiment-oriented vocabulary replacement on the vocabulary to be adjusted to generate sentiment-corrected adjusted vocabulary, and replacing corresponding vocabulary in the initial transcription text with the adjusted vocabulary to obtain the target transcription text.

9. The method according to claim 8, characterized in that Performing sentiment adaptability testing on the vocabulary in the initial transcription based on the sentiment correction rule set to identify the vocabulary to be adjusted that does not match the current sentiment state, including: Extracting vocabulary emotion adaptability detection rules from the emotion correction rule set, wherein the vocabulary emotion adaptability detection rules include vocabulary expression adaptation rules and vocabulary expression prohibition rules corresponding to the current emotion feature tag; Scanning the vocabulary in the initial transcription text word by word according to the vocabulary expression adaptation rule to identify abnormal vocabulary that does not conform to the adaptation rule; Screening the words in the initial transcription text word by word according to the word expression prohibition rule to identify conflicting words that meet the prohibition rule; The abnormal vocabulary and the conflicting vocabulary are merged to form the vocabulary to be adjusted.

10. A speech recognition transcription system based on multimodal fusion and sentiment analysis, characterized by: A system for implementing the speech recognition transcription method based on multimodal fusion and sentiment analysis according to any one of claims 1 to 9, comprising: An information acquisition module is used to acquire a target speech signal and simultaneously acquire auxiliary modal information corresponding to the target speech signal, wherein the auxiliary modal information includes visual information and textual context information; A signal segmentation module is used to perform segmentation and recognition processing on the target speech signal to obtain a speech feature vector, and to perform feature extraction on the text context information to obtain a text auxiliary feature vector; A fusion processing module, configured to perform multimodal fusion processing on the speech feature vector and the text auxiliary feature vector to generate a fusion feature representation; A speech transcription module, configured to perform speech transcription based on the fused feature representation to obtain an initial transcription text; A sentiment analysis module, configured to perform sentiment analysis based on the visual information and the initial transcribed text to generate a sentiment feature label; The optimization and correction module is used to perform text optimization and correction on the initial transcription text based on the emotional feature label to obtain a target transcription text.

Citation Information

Cited By

  • Pickup pen voice content semantic analysis method fused with industry knowledge base

    CN121191507A

  • Multi-language voice content recognition method and system

    CN121662048A

  • A method and system for multilingual speech content recognition

    CN121662048B