A multimodal speech recognition error correction method and system
By using a multimodal speech recognition error correction method and combining multi-dimensional associated data, a verification benchmark and associated influence coefficient are constructed, which solves the problem of speech transcription deviation in traditional methods, and achieves the accuracy and reliability of speech error correction, adapting to the needs of different scenarios and individual users.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-18
- Publication Date
- 2026-03-27
AI Technical Summary
Traditional speech recognition error correction methods fail to fully integrate multi-dimensional related information, resulting in speech transcription errors. They are difficult to adapt to the differentiated needs of different scenarios and individual users, affecting the accuracy and reliability of error correction.
By using multimodal speech recognition error correction methods, combining multimodal correlation data such as pronunciation duration, syllable interval, intonation changes, accent type, language habits, noise level, and scene type, a verification benchmark is constructed between standard scenes and adapted profile scenes. The correlation influence coefficient is extracted to judge and calibrate the degree of deviation of speech-to-text, and to perform accurate error correction.
It enables speech-to-text transcription to conform to general language standards while also catering to individual expression habits and scenario needs, improving the accuracy and reliability of error correction, and adapting to the needs of different speech application scenarios.
Smart Images

Figure CN121354572B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of speech recognition, more particularly, it relates to a multi-modal speech recognition error correction method and system. BACKGROUND
[0002] Traditional speech recognition error correction methods mostly rely on single text semantic analysis or only focus on limited speech feature data, and fail to fully integrate multi-dimensional associated information in the speech expression process. The generation of speech transcription deviation in actual scenarios is not only caused by semantic specification inconsistency. The actual speech transcription deviation is influenced by sound features such as pronunciation duration and syllable interval, environmental features such as noise level and scene type, and individual features of the user such as accent type and language habit. The fragmented processing of multi-modal information makes it difficult for traditional methods to fully cover the causes of deviation. Speech expression needs differ in different scenarios, and traditional methods lack differentiated consideration of differences between general standard scenarios and user personalized adaptation scenarios. Only relying on general language specification for error correction can easily ignore the expression habits of specific users, resulting in a deviation between the corrected text and the user's true intention, making it difficult to control the degree of deviation in the correction process, and ultimately affecting the accuracy and reliability of the transcribed text. SUMMARY
[0003] In view of the deficiencies of the prior art, the purpose of the present application is to provide a multi-modal speech recognition error correction method and system.
[0004] To achieve the above-mentioned purpose, the present application provides the following technical solutions:
[0005] A multi-modal speech recognition error correction method, the method comprising the following steps:
[0006] According to the target speech transcription text to be corrected and the corresponding multi-modal associated data, a first error correction value of the target speech transcription text in the standard scenario and the adaptation scenario is obtained.
[0007] If the different data dimension features in the historical multi-modal associated data show a non-adaptation variation state, a first associated influence coefficient of the non-adaptation variation state affecting the recognition deviation is extracted.
[0008] If the different data dimension features in the historical multi-modal associated data show an adaptation stable state, a second associated influence coefficient of the cross-dimension information adaptation conversion affecting the recognition deviation is extracted, and a third associated influence coefficient of the same-dimension internal feature adaptation conversion affecting the recognition deviation is extracted.
[0009] According to the first associated influence coefficient, the second associated influence coefficient and the third associated influence coefficient, the deviation degree of the target speech transcription text in the data adaptation state and the dimension conversion influence is judged to obtain a second error correction value.
[0010] The target comparison verification benchmark for target speech transcription text error correction verification is selected from the comprehensive comparison verification benchmark, and the target comparison verification benchmark, the first error correction value and the second error correction value are processed to obtain actual error correction values one and two. The target speech transcription text is corrected according to the actual error correction value one or the actual error correction value two to obtain a correct recognition result.
[0011] Preferably, the multi-modal association data includes pronunciation duration, syllable interval, intonation change, accent type, language habit, noise level, scene type and dialogue scene.
[0012] Preferably, according to the target speech transcription text to be corrected and the corresponding multi-modal association data, the first error correction value of the target speech transcription text in the standard scene and the adaptive portrait scene is obtained, specifically including the following steps:
[0013] The semantic structure between the text core ideographic element and the semantic association is obtained by performing semantic splitting processing on the target speech transcription text; the multi-dimensional feature system is constructed by separating the sound features, environment features related to the use environment and individual features of the user in the multi-modal association data;
[0014] The standard scene verification benchmark is built according to the semantic structure and the multi-dimensional feature system, the standard scene verification benchmark integrates the general language expression specification, the voice transmission rule of the conventional environment and the semantic logic criterion, the semantic deviation items are obtained by comparing the semantic structure with the standard scene verification benchmark, the number and influence degree of the semantic deviation items are determined to form the standard scene deviation information;
[0015] The adaptive portrait scene verification benchmark is constructed according to the individual features, the language expression habit of the user, the expression preference of the specific scene and the common semantic combination mode of the multi-dimensional feature system, the adaptive deviation items that the text conforms to the general specification and is contrary to the expression habit of the user are identified after matching the semantic structure with the adaptive portrait scene verification benchmark, the necessity of correction and the associated influence range of the adaptive deviation items are judged to obtain the adaptive portrait scene deviation information;
[0016] The deviation weight distribution rule is established based on the standard scene deviation information and the adaptive portrait scene deviation information, the influence weights of the standard scene deviation information and the adaptive portrait scene deviation information are determined according to the scene use priority and the specification requirement, and the standard scene deviation information and the adaptive portrait scene deviation information are fused according to the weight distribution rule to obtain the first error correction value reflecting the scene adaptability and the semantic accuracy.
[0017] Preferably, the first associated influence coefficient of the non-adaptive change state influence recognition deviation is extracted, specifically including the following steps:
[0018] The historical transcription text one is selected from all historical speech transcription texts;
[0019] Extract the recognition error rate data of the dependent data dimension from the historical transcribed text;
[0020] Based on the degree of misfit variation in the recognition deviation rate data, the first correlation influence coefficient of the misfit variation state on the recognition deviation is obtained.
[0021] Preferably, the second correlation coefficient affecting the recognition bias of cross-dimensional information adaptation and transformation is extracted, specifically including the following steps:
[0022] Select historical transcribed text two from all historical speech-to-text transcripts;
[0023] Extract the recognition error rate data from the core data dimension and auxiliary data dimension of the historical transcribed text 2;
[0024] Based on the stable value of the data adaptation of the recognition deviation rate data 2, the second correlation influence coefficient of cross-dimensional information adaptation transformation affecting the recognition deviation is obtained.
[0025] Preferably, the third correlation influence coefficient of the influence of feature adaptation and transformation on the recognition bias within the same dimension is extracted, specifically including the following steps:
[0026] Extract the recognition error rate data from the historical transcribed text (Part 2) in terms of dependent data dimension, auxiliary data dimension, and core data dimension (Part 3);
[0027] Based on the stable value of data adaptation of the recognition deviation rate data 3, the third correlation influence coefficient of the influence of feature adaptation transformation on recognition deviation in the same dimension is obtained.
[0028] Preferably, the second error correction value is obtained by judging the degree of deviation of the target speech-to-text in terms of data adaptation status and dimension transformation influence based on the first correlation influence coefficient, the second correlation influence coefficient, and the third correlation influence coefficient, specifically including the following steps:
[0029] The first, second, and third correlation influence coefficients are subjected to characteristic judgment to obtain coefficient characteristic status information; wherein, the coefficient characteristic status information includes the deviation influence type, scope of action, and transmission path;
[0030] The coefficient association weight information is obtained based on the semantic structure of the target speech-to-text and the feature system of multimodal association data; wherein, the coefficient association weight information includes semantic deviation, data adaptation status, and dimensional transformation association strength;
[0031] Establish a coefficient influence transmission link based on coefficient characteristic information and coefficient correlation weight information; determine the deviation impact of the first correlation influence coefficient, second correlation influence coefficient, and third correlation influence coefficient based on the coefficient influence transmission link to obtain the coefficient linkage impact result;
[0032] construct a deviation degree quantization standard based on the coefficient linkage influence result; wherein, the deviation degree quantization standard combines the semantic tolerance range of general language expression, the reasonable fluctuation interval of multi-modal data adaptation, and the normal connection requirement of dimension conversion, compares the coefficient linkage influence result with the deviation degree quantization standard to obtain a quantization index corresponding to the deviation degree level;
[0033] According to the quantization index, the coefficient linkage influence result is numerically converted to obtain a numerical conversion result, and the numerical conversion result is adjusted by fusing the coefficient correlation weight information to obtain a second error correction value of the data adaptation state deviation and the dimension conversion influence degree.
[0034] Preferably, the target comparison verification benchmark, the first error correction value and the second error correction value are processed to obtain actual error correction value one and actual error correction value two, specifically including the following steps:
[0035] The benchmark verification specification is extracted from the target comparison verification benchmark; wherein, the benchmark verification specification includes the judgment standard of semantic accuracy, the compliance range of data adaptation, and the connection requirement of dimension conversion;
[0036] After comparing the second error correction value with the benchmark verification specification, a deviation component exceeding the benchmark compliance range in the second error correction value is obtained, and a deviation correction factor is formed after judging the influence amplitude of the deviation component;
[0037] According to the deviation correction factor, the second error correction value is calibrated to obtain a second calibrated error correction value; the correlation fit degree of the scene adaptation deviation to which the first error correction value belongs and the data dimension deviation to which the second calibrated error correction value belongs is judged to obtain a complementary feature;
[0038] Taking the scene adaptation priority as the guide, the first error correction value is taken as the leading component, the second calibrated error correction value is taken as the deviation supplementary component, and the actual error correction value one is obtained by component superposition;
[0039] Taking the data dimension deviation as the core, the second calibrated error correction value is taken as the verification component, and the first error correction value is taken as the scene adaptation constraint component, and the actual error correction value two is obtained by component balance and complementary cooperation.
[0040] Preferably, the target speech transcription text is corrected according to the actual error correction value one or the actual error correction value two to obtain a correct recognition result, specifically including the following steps:
[0041] According to the actual error correction value one or the actual error correction value two, a corresponding scheme of a multi-modal error correction strategy library is matched to obtain a target error correction scheme;
[0042] According to the target error correction scheme, the deviation content of the target speech transcription text is corrected to obtain a correct recognition result.
[0043] A multi-modal speech recognition error correction system comprises:
[0044] A processing module obtains a first error correction value of a target speech transcription text in a standard scenario and an adaptive scenario according to the target speech transcription text and corresponding multi-modal correlation data to be corrected;
[0045] A first extraction module extracts a first correlation influence coefficient of a non-adaptive variation state affecting recognition deviation if different data dimension features in historical multi-modal correlation data are in a non-adaptive variation state;
[0046] A second extraction module extracts a second correlation influence coefficient of a cross-dimension information adaptive conversion affecting recognition deviation and a third correlation influence coefficient of a same-dimension internal feature adaptive conversion affecting recognition deviation if different data dimension features in historical multi-modal correlation data are in an adaptive stable state;
[0047] A judgment module judges a deviation degree of the target speech transcription text in a data adaptive state and a dimension conversion influence to obtain a second error correction value according to the first correlation influence coefficient, the second correlation influence coefficient and the third correlation influence coefficient;
[0048] An output module selects a target comparison and verification benchmark for error correction verification of the target speech transcription text from a comprehensive comparison and verification benchmark, processes the target comparison and verification benchmark, the first error correction value and the second error correction value to obtain actual error correction values one and two, and performs error correction processing on the target speech transcription text according to the actual error correction values one and two to obtain a correct recognition result.
[0049] Compared with the prior art, the present application has the following beneficial effects:
[0050] The standard scene in the application guarantees that the transcribed text conforms to the general language specification and semantic logic, avoiding obvious expression deviation that is contrary to the public cognition. The portrait scene fully considers the individual expression habits of the user and the communication preferences of the specific scene, so that the corrected text not only conforms to the specification, but also meets the actual communication needs, that is, quantifying the influence of different data states on recognition deviation. For different states of historical multi-modal associated data, the corresponding associated influence coefficients are extracted, and the second correction value can accurately reflect the comprehensive influence of data adaptation state and dimension conversion on deviation, so that the correction process is no longer dependent on fuzzy experience judgment, but based on clear quantitative data, improving the accuracy of correction. The target comparison verification benchmark, the first correction value and the second correction value are processed to obtain actual correction value one and actual correction value two, and different actual correction values match the corresponding correction scheme, so that the method can flexibly adapt to the needs of different scenes such as restaurant ordering and office reporting, and improve the comprehensiveness and reliability of speech recognition correction. The scheme integrates multi-dimensional features of multi-modal associated data, covering multiple deviation causes of semantics, data adaptation and dimension conversion, and forms a complete correction link from scene verification, coefficient quantization to multi-benchmark collaborative processing. Thus, the scheme finally obtains a correct recognition result that is more in line with the actual speech expression. BRIEF DESCRIPTION OF DRAWINGS
[0051] Fig. 1 A step schematic diagram of a multi-modal speech recognition correction method is provided for the application;
[0052] Fig. 2 A module schematic diagram of a multi-modal speech recognition correction system is provided for the application. DETAILED DESCRIPTION
[0053] In order to make the above-mentioned purposes, features and advantages of the application more obvious and easy to understand, the specific embodiments of the application will be described in detail below with reference to the accompanying drawings.
[0054] In the following description, many specific details are set forth in order to provide a thorough understanding of the application, but the application can also be implemented in other ways different from the description, and those skilled in the art can make similar generalizations without departing from the connotation of the application, so the application is not limited to the specific embodiments disclosed below.
[0055] Secondly, "one embodiment" or "embodiment" referred to herein means that a specific feature, structure or characteristic can be included in at least one implementation of the application. "In one embodiment" appearing in different places in the specification does not mean the same embodiment, nor is it an embodiment that is independent of or selectively excludes other embodiments.
[0056] Reference Figs. 1-2 is shown.
[0057] The embodiment further illustrates a multi-modal speech recognition error correction method and system of the present application.
[0058] A multi-modal speech recognition error correction method, comprising the following steps:
[0059] According to the target speech transcription text to be corrected and the corresponding multi-modal associated data, the first error correction value of the target speech transcription text in the standard scene and the adaptive image scene is obtained.
[0060] If the different data dimension features in the historical multi-modal associated data show a non-adaptive variation state, the first associated influence coefficient of the non-adaptive variation state affecting the recognition bias is extracted.
[0061] If the different data dimension features in the historical multi-modal associated data show an adaptive stable state, the second associated influence coefficient of the cross-dimension information adaptive conversion affecting the recognition bias is extracted, and the third associated influence coefficient of the same-dimension internal feature adaptive conversion affecting the recognition bias is extracted.
[0062] According to the first associated influence coefficient, the second associated influence coefficient, and the third associated influence coefficient, the second error correction value of the target speech transcription text in the data adaptive state and the dimension conversion influence bias degree is obtained.
[0063] The target comparison and verification benchmark for target speech transcription text error correction verification is selected from the comprehensive comparison and verification benchmark.
[0064] First of all, it needs to be clear that the comprehensive comparison and verification benchmark is a set of multi-scene, multi-dimension verification rules. The comprehensive comparison and verification benchmark includes the semantic accuracy judgment standard of different speech application scenes, the compliance range of multi-modal data adaptation, and the connection requirement content of data dimension conversion. The target comparison and verification benchmark is selected from the comprehensive comparison and verification benchmark to match the current target speech transcription text. In this way, the pertinence and accuracy of subsequent error correction verification are ensured.
[0065] The selection process is first carried out in combination with the multi-modal associated data corresponding to the target speech transcription text. If the scene type corresponding to the target speech transcription text is a noisy restaurant conversation scene, and the user's accent type is a certain local dialect accent, and the noise level is high noise, then the related benchmark content in the comprehensive comparison and verification benchmark for the restaurant conversation scene, the local dialect accent adaptation, and the speech transmission in the high noise environment will be preferentially included in the selection range.
[0066] The semantic structure features of the target speech transcription text are combined for selection. Assuming that the semantic core expression element of the target speech transcription text is the ordering demand, the content related to the dining scene semantic expression specification and the ordering class semantic logic criterion in the comprehensive comparison and verification benchmark is selected.
[0067] After the preliminary screening, the candidate reference content needs to be checked for adaptability. For example, the compliance range of the candidate reference on the syllable interval in a high noise environment matches the syllable interval feature corresponding to the current target speech transcription text. If the compliance range of the reference can cover the current target syllable interval fluctuation interval, the reference content is retained; if it cannot be covered, the part of the content is excluded.
[0068] After scene matching, semantic matching and adaptability checking, the reference content obtained is used as the target comparison verification reference for target speech transcription text error correction verification. The target comparison verification reference contains semantic accuracy judgment standards that are highly consistent with the current target, such as the standard form of dish name expression in the ordering scene, such as the reasonable fluctuation interval of the tone change in the compliance range of the data adaptation in the high noise scene of the restaurant, such as the transition requirement of the dimension conversion from the pronunciation duration dimension to the semantic expression integrity dimension. The transition connection rule content provides a reference basis for subsequent error correction value calibration and actual error correction processing.
[0069] The target comparison verification reference, the first error correction value and the second error correction value are processed to obtain actual error correction value one and actual error correction value two. The target speech transcription text is corrected according to the actual error correction value one or the actual error correction value two to obtain the correct recognition result.
[0070] The multi-modal associated data includes pronunciation duration, syllable interval, tone change, accent type, language habit, noise level, scene type and dialogue scene.
[0071] According to the target speech transcription text to be corrected and the corresponding multi-modal associated data, the first error correction value of the target speech transcription text in the standard scene and the adaptation image scene is obtained, which specifically includes the following steps:
[0072] The target speech transcription text is processed to obtain the semantic structure between the text core ideographic element and the semantic association. The multi-dimensional feature system is constructed by separating the sound features, environment features related to the use environment, and individual features of the user in the multi-modal associated data.
[0073] The standard scene checking reference is built according to the semantic structure and the multi-dimensional feature system. The standard scene checking reference integrates general language expression specifications, conventional environment speech transmission rules and semantic logic criteria. The semantic structure and the standard scene checking reference are compared element by element to obtain the semantic deviation item. After determining the number and influence degree of the semantic deviation item, the standard scene deviation information is formed.
[0074] According to the individual characteristics of the multi-dimensional feature system, the user language expression habit, the expression preference of the specific scene, and the commonly used semantic combination method, the adaptive image scene verification benchmark is constructed. After matching the semantic structure with the adaptive image scene verification benchmark, the adaptive deviation items of the text that conform to the general specification and are contrary to the user expression habit are identified. The correction necessity and the associated influence range of the adaptive deviation items are judged to obtain the adaptive image scene deviation information.
[0075] Based on the standard scene deviation information and the adaptive image scene deviation information, the deviation weight distribution rule is established. According to the scene use priority and the specification requirement, the influence weight of the standard scene deviation information and the adaptive image scene deviation information is determined. According to the weight distribution rule, the standard scene deviation information and the adaptive image scene deviation information are fused to obtain the first error correction value reflecting the scene adaptability and the semantic accuracy.
[0076] The target voice transcription text and the corresponding multi-modal associated data are processed. The pronunciation duration refers to the duration of each syllable or sentence, such as the pronunciation duration of a user saying "Hello" is 0.5 seconds longer than usual. The syllable interval is the pause time between syllables, such as a dialect user may add an interval between certain words. The intonation change is the ups and downs of the voice, such as the degree of intonation rising at the end of a question. The accent type is the dialect or regional voice characteristics of the user, such as the pronunciation deviation of Cantonese accent for Mandarin words. The language habit is the user's common expression sentence, such as someone is used to replacing "I" with "we". The noise level is the noise intensity of the environment where the voice is collected, such as the medium noise in a coffee shop. The scene type is the occasion where the voice occurs, such as the specific situation of communication in the office conversation scene is the customer consultation scene. During processing, the semantics of the target voice transcription text are first split to obtain the core meaning elements and the semantic associated structure, such as the core meaning of the transcription text "I want a coffee" is "ordering coffee", and the semantic association is the subject, action and object. The multi-modal associated data are constructed into a multi-dimensional feature system according to the sound characteristics (pronunciation duration, syllable interval, intonation change, accent type), environmental characteristics (noise level, scene type), individual characteristics (language habit, dialogue scene).
[0077] A standard scene verification benchmark is built and standard scene deviation information is obtained. The standard scene deviation information integrates general language expression specifications, regular environmental voice transmission rules and semantic logic criteria. For example, the standard expression of "coffee" in the general specification is "coffee" rather than "gah", and the voice transmission rule in the regular office scene is that the voice is clear when the noise level is lower than 20 decibels. The semantic structure of the text is compared with the benchmark element by element: for example, "gah" in the transcribed text does not conform to the general expression of "coffee", which is a semantic deviation item; at the same time, the noise level corresponding to the text is 30 decibels, which is higher than the 20 decibels of the regular office scene, resulting in a shortened pronunciation time, which is also a deviation item. The number of these semantic deviation items is counted to evaluate the influence degree of each deviation item, for example, the deviation of "gah" will cause ambiguous expression, and the influence degree is high, and finally the standard scene deviation information is integrated.
[0078] An adaptive portrait scene verification benchmark is built and adaptive portrait scene deviation information is obtained. For example, for a user who is used to saying "we", "we" is a reasonable subject in the adaptive portrait of the user; for a user who often goes to a coffee shop, "gah" may be a common simplified expression of the user. After matching the semantic structure with this benchmark, adaptive deviation items are identified: for example, the transcribed text "we want a gah" conforms to the semantic logic in the general specification, but "gah" is contrary to the user's habit of saying "coffee" completely in daily life, which is an adaptive deviation item. If the dialogue scene is a formal business order, the simplified expression of "gah" may cause misunderstanding, and correction is necessary, with a high necessity; if it is only a daily order between acquaintances, the correction necessity is low, and the associated influence range is determined to be "the completeness of the expression of the name of the article", so as to finally form the adaptive portrait scene deviation information.
[0079] The deviation information of the two scenes is fused to obtain a first correction value. First, a deviation weight distribution rule is established, and the weights are determined according to the scene use priority and specification requirements, for example, the weight of the standard scene deviation information is set to 0.7 and the weight of the adaptive portrait scene deviation information is set to 0.3 in a formal business scene; in a daily acquaintance scene, the weights of the two are adjusted to 0.4 and 0.6. Assuming that the quantitative value of the standard scene deviation information is S, the quantitative value of the adaptive portrait scene deviation information is P, the weights are W1 and W2 respectively, and the first correction value is SxW1+PxW2. For example, in a formal business scene, the quantitative value of the standard scene deviation information S is 8 and the quantitative value of the adaptive portrait scene deviation information P is 5, and the first correction value is obtained by substituting the data, which is 8x0.7+5x0.3=7.1, the first correction value reflects the deviation degree of the text in the scene adaptability and semantic accuracy, which provides a basis for subsequent correction.
[0080] A first associated influence coefficient of the non-adaptive variable state influence recognition deviation is extracted, which includes the following steps:
[0081] A historical transcribed text one is selected from all historical voice transcribed texts;
[0082] extracting recognition bias rate data one belonging to the dependent data dimension from the historical transcription text one;
[0083] obtaining a first correlation influence coefficient of the non-adaptation variation state affecting recognition bias according to the non-adaptation variation degree value of the recognition bias rate data one.
[0084] selecting historical transcription text one from all historical voice transcription texts. Historical voice transcription text refers to voice text data that has been transcribed and biased in the past. The selection of historical transcription text one needs to match the core features of the current target voice transcription text, such as the current target text corresponds to the scene of a noisy restaurant ordering, and the voice transcription text in the historical data that also belongs to the "noisy restaurant ordering" scene is selected as the historical transcription text one.
[0085] extracting recognition bias rate data one belonging to the dependent data dimension from the historical transcription text one. The dependent data dimension refers to the multi-modal correlation data dimension that directly affects the voice recognition bias, such as noise level, pronunciation duration, and syllable interval in a noisy restaurant scene, which are usually the core dependent data dimensions, and the change of the dependent data dimension directly affects the accuracy of transcription. The recognition bias rate data one refers to the proportion of transcription bias caused by these dependent data dimensions, such as the proportion of transcription bias caused by the shortening of pronunciation duration when the noise level of historical transcription text one exceeds 60 decibels, or the probability of semantic omission bias when the syllable interval exceeds 0.8 seconds. These specific proportion or probability values are the recognition bias rate data one.
[0086] obtaining a first correlation influence coefficient according to the non-adaptation variation degree value of the recognition bias rate data one. The non-adaptation variation degree value refers to the degree of deviation of the actual feature value of the dependent data dimension from its reasonable adaptation range, such as the reasonable adaptation range of noise level in a noisy restaurant scene is 40-60 decibels, if the noise level of a certain historical data reaches 75 decibels, then its non-adaptation variation degree value is (actual value-adaptation upper limit) / adaptation upper limit, so the non-adaptation variation degree value is (75-60) / 60=0.25.
[0087] The first correlation influence coefficient is calculated according to the recognition deviation rate data one and the non-adaptation variation degree value. Assuming that the recognition deviation rate data one is R1 and the non-adaptation variation degree value is D1, the calculation formula of the first correlation influence coefficient K1 can be represented as K1=R1×(1+D1). If the recognition deviation rate data one R1 corresponding to the noise level in the historical transcription text one is 0.3, that is, 30% deviation is caused by the noise level, and the non-adaptation variation degree value D1 is 0.25, then K1=0.3×(1+0.25)=0.375 is obtained by substituting the formula. The first correlation influence coefficient quantifies the influence degree of the non-adaptation variation state on the recognition deviation. If the first correlation influence coefficient is higher, it means that the non-adaptation variation of the dependent data dimension has a stronger driving effect on the transcription deviation.
[0088] The second correlation influence coefficient of the cross-dimension information adaptation conversion influence recognition deviation is extracted, specifically including the following steps:
[0089] Selecting a historical transcription text two from all historical voice transcription texts;
[0090] Extracting recognition deviation rate data two in the core data dimension and the auxiliary data dimension from the historical transcription text two;
[0091] According to the data adaptation stability value of the recognition deviation rate data two, the second correlation influence coefficient of the cross-dimension information adaptation conversion influence recognition deviation is obtained.
[0092] Selecting a historical transcription text two from all historical voice transcription texts, the selection basis is to match the current target voice transcription text scene, such as the current target is the voice transcription in the restaurant ordering scene, and the voice transcription text in the same restaurant ordering scene is selected from the historical data as the historical transcription text two, to ensure that the deviation data extracted subsequently has relevance with the target scene.
[0093] In the restaurant ordering scenario, the core data dimension is the multi-modal correlation data dimension that directly affects the semantic accuracy of voice transcription, such as pronunciation duration, tone change, and language habits. Pronunciation duration refers to the duration of each word spoken during ordering, such as when a user says "hamburger" with a too-short pronunciation duration, which can result in transcription omission. Tone change refers to the ups and downs of the voice, such as the rising tone of a question during ordering, which can be confusing if not obvious, resulting in a mixed statement and question in the transcription text. Language habits refer to the user's common expressions, such as some users' habit of referring to "coke" as "le", which directly affects the accuracy of the transcription. The auxiliary data dimension is a dimension that indirectly affects the feature transmission of the core dimension, such as the noise level and scene type in the restaurant. Noise level refers to the noise intensity of the environment, such as the high noise during peak hours in a restaurant, which can interfere with the recognition of pronunciation duration in the core dimension. Scene type refers to the specific occasion of ordering, such as the noisy environment of a fast-food restaurant, which can affect the clarity of voice collection. After determining the dimensions, the recognition bias rate data II is extracted from the historical transcription text II, which is the statistical proportion of the bias of the core dimension, such as the proportion of transcription bias caused by pronunciation duration anomalies and the bias proportion under the auxiliary dimension, such as the proportion of core dimension feature recognition bias caused by high noise level. These proportion values are the recognition bias rate data II.
[0094] The data adaptation stability value is an index that measures the degree of feature adaptation and matching of the core data dimension and the auxiliary data dimension. Assuming that the actual value of the feature of the core data dimension is C, and the reasonable adaptation range mean value of this dimension in the restaurant ordering scenario is C0; the actual value of the feature of the auxiliary data dimension is A, and the reasonable adaptation range mean value of this dimension in the scenario is A0, then the calculation formula of the data adaptation stability value S2 is S2 = 1 - |(C-C0) / C0-(A-A0) / A0|. Taking the restaurant ordering scenario as an example, the reasonable mean value C0 of the pronunciation duration in the core dimension is 0.8 seconds, and the actual pronunciation duration C of the user saying "French fries" in a certain historical data is 0.6 seconds. The reasonable mean value A0 of the noise level in the auxiliary dimension is 40 decibels, and the actual noise level A corresponding to this historical data is 55 decibels. Substituting the formula gives the data adaptation stability value as S2 = 1 - |(0.6-0.8) / 0.8-(55-40) / 40| = 1 - |-0.25-0.375| = 0.375. The closer the data adaptation stability value is to 1, the more stable the feature adaptation of the core and auxiliary dimensions.
[0095] If the recognition bias rate data two extracted from the historical transcription text two is R2, and the data adaptation stability value is S2, then the calculation formula of the second correlation influence coefficient K2 is K2=R2×(1-S2). For example, in the restaurant ordering scenario, the recognition bias rate data two R2 is 0.25, that is, 25% of the transcription bias is caused by the adaptation problem of the core and auxiliary dimensions, and K2=0.25×(1-0.375)=0.15625 is obtained by substituting the formula. This coefficient reflects the influence degree of cross-dimension information adaptation conversion on recognition bias, and the higher the coefficient, the worse the feature adaptation stability of the core and auxiliary dimensions.
[0096] The third correlation influence coefficient of the same dimension internal feature adaptation conversion affecting recognition bias is extracted, which specifically includes the following steps:
[0097] Extracting recognition bias rate data three in the dependent data dimension, the auxiliary data dimension, and the core data dimension from the historical transcription text two;
[0098] According to the data adaptation stability value of the recognition bias rate data three, the third correlation influence coefficient of the same dimension internal feature adaptation conversion affecting recognition bias is obtained.
[0099] Extracting recognition bias rate data three from the historical transcription text two, taking the noisy restaurant ordering scenario as an example, the syllable interval is selected as the dependent data dimension, the noise level is selected as the auxiliary data dimension, and the pronunciation duration is selected as the core data dimension. The recognition bias rate data three is the bias proportion corresponding to different sub-features of the three types of dimensions, for example, the proportion of semantic bias caused by short pronunciation duration (less than 0.5 seconds) in the core data dimension pronunciation duration, the bias proportion corresponding to normal pronunciation duration (0.5-1 second), and the recognition bias proportion caused by long interval (more than 1 second) in the dependent data dimension syllable interval. These specific proportion values together constitute the recognition bias rate data three.
[0100] The data adaptation stability value is an index for measuring the matching degree between different sub-features in the same dimension. Taking the core data dimension pronunciation duration as an example, it is assumed that the actual short pronunciation duration in the dimension is C1, and the reasonable adaptation range mean value is C10; the actual normal pronunciation duration is C2, and the reasonable adaptation range mean value is C20. The calculation formula of the data adaptation stability value S3 can be expressed as: S3 = 1 - |(C1-C10) / C10-(C2-C20) / C20|. For example, in the noisy restaurant ordering scene, the reasonable mean value C10 of the short pronunciation duration is 0.4 seconds, and the actual short pronunciation duration C1 of a certain historical data is 0.3 seconds; the reasonable mean value C20 of the normal pronunciation duration is 0.8 seconds, and the actual normal pronunciation duration C2 is 0.9 seconds, and S3 = 1 - |(0.3-0.4) / 0.4-(0.9-0.8) / 0.8| = 1 - |-0.25-0.125| = 0.625 is obtained by substituting the formula. The closer this value is to 1, the more stable the adaptation of the sub-features in the same dimension.
[0101] It is assumed that the recognition bias rate data three R3 extracted from the historical transcription text two is R3, and the data adaptation stability value is S3. The calculation formula of the third correlation influence coefficient K3 is K3 = R3 x (1-S3). For example, the recognition bias rate data three R3 is 0.25, that is, 25% of the bias is caused by the adaptation problem of the sub-features in the same dimension, and the data adaptation stability value S3 is 0.625. Substituting the formula obtains K3 = 0.25 x (1-0.625) = 0.09375. This coefficient quantifies the influence of the adaptation conversion of the sub-features in the same dimension on the recognition bias. The higher the coefficient, the worse the stability of the adaptation of the sub-features in the same dimension.
[0102] According to the first correlation influence coefficient, the second correlation influence coefficient, and the third correlation influence coefficient, the bias degree of the target voice transcription text in the data adaptation state and the dimension conversion influence is determined to obtain a second error correction value, which specifically includes the following steps:
[0103] The characteristics of the first correlation influence coefficient, the second correlation influence coefficient, and the third correlation influence coefficient are determined to obtain coefficient characteristic state information. The coefficient characteristic state information includes bias influence type, action range, and transmission path.
[0104] According to the semantic structure of the target voice transcription text and the feature system of the multi-modal correlation data, coefficient correlation weight information is obtained. The coefficient correlation weight information includes semantic bias, data adaptation state, and dimension conversion correlation strength.
[0105] According to the coefficient characteristic state information and the coefficient correlation weight information, a coefficient influence transmission link is established. According to the coefficient influence transmission link, the bias influence of the first correlation influence coefficient, the second correlation influence coefficient, and the third correlation influence coefficient is determined to obtain a coefficient linkage influence result.
[0106] construct a deviation degree quantization standard based on the coefficient linkage influence result; wherein the deviation degree quantization standard combines the semantic fault tolerance range of general language expression, the reasonable fluctuation interval of multi-modal data adaptation, and the normal connection requirement of dimension conversion, compares the coefficient linkage influence result with the deviation degree quantization standard to obtain a quantization index corresponding to the deviation degree level;
[0107] According to the quantization index, the coefficient linkage influence result is numerically converted to obtain a numerical conversion result, and the numerical conversion result is adjusted by fusing the coefficient correlation weight information to obtain a second error correction value of the data adaptation state deviation and the dimension conversion influence degree.
[0108] The first correlation influence coefficient, the second correlation influence coefficient, and the third correlation influence coefficient are subjected to characteristic judgment. The first correlation influence coefficient corresponds to the influence of a non-adaptation variation state, such as a sudden increase in noise level from 40 decibels to 60 decibels; the second correlation influence coefficient corresponds to the influence of cross-dimension adaptation conversion, such as the adaptation deviation of pronunciation duration (core dimension) and noise level (auxiliary dimension); and the third correlation influence coefficient corresponds to the influence of intra-dimension feature conversion, such as the adaptation deviation of different interval durations within the syllable interval (dependent dimension). After characteristic judgment, coefficient characteristic condition information is obtained. In terms of deviation influence type, the first correlation influence coefficient is an environmental mutation type deviation, the second correlation influence coefficient is a cross-dimension adaptation type deviation, and the third correlation influence coefficient is a same-dimension feature type deviation; in terms of action range, the first correlation influence coefficient affects the recognition accuracy of the entire ordering voice, the second correlation influence coefficient affects the information transmission of the core and auxiliary dimensions, and the third correlation influence coefficient affects the feature matching of a single dimension; in terms of transmission path, the first correlation influence coefficient interferes with pronunciation feature recognition through environmental noise, the second correlation influence coefficient amplifies the deviation through adaptation imbalance between the core and auxiliary dimensions, and the third correlation influence coefficient directly causes semantic deviation through mismatching of intra-dimension features.
[0109] In combination with the semantic structure of the restaurant ordering scene and the multi-modal correlation data feature system, such as pronunciation duration and noise level, the semantic deviation weight corresponds to the deviation degree of the transcribed text from the actual ideographic expression, such as the deviation weight of "hamburger" being transcribed as "han" being set to 0.4; the data adaptation state weight corresponds to the matching degree of multi-modal data, such as the adaptation state weight of noise level and pronunciation duration being set to 0.3; and the dimension conversion correlation strength weight corresponds to the information transmission effect of different dimensions, such as the conversion strength weight from pronunciation duration to semantic expression being set to 0.3.
[0110] The coefficient influence conduction link is constructed according to the coefficient characteristic condition information and the correlation weight information, and a linkage influence result is obtained. For example, in a restaurant ordering scenario, a non-adaptive change in noise level (first correlation influence coefficient) first interferes with the pronunciation duration, amplifies the deviation through the cross-dimensional adaptive deviation between pronunciation duration and noise level (second correlation influence coefficient), and finally superimposes the deviation of the same-dimensional feature of syllable interval (third correlation influence coefficient), thereby forming a complete deviation conduction path. The linkage influence of the three coefficients is judged through the link, for example, the influence degree of the first correlation influence coefficient is 0.3, the influence degree of the second correlation influence coefficient is 0.2, and the influence degree of the third correlation influence coefficient is 0.2. The comprehensive influence result after the linkage of the three is recorded as L.
[0111] A deviation degree quantization standard is constructed, and a quantization index is obtained. The standard combines the semantic fault tolerance range of the restaurant ordering scenario, such as that "Han" can be fault-tolerant recognized as "hamburger", the multi-modal data adaptive fluctuation interval, such as that the reasonable fluctuation of noise level is 30-50 decibels, and the dimension conversion connection requirement, such as the normal connection threshold of pronunciation duration and semantic expression. The coefficient linkage influence result L is compared with the standard to obtain the quantization index M corresponding to the deviation degree level, for example, the deviation level corresponding to L is moderate, and the quantization index M is 0.5.
[0112] The coefficient linkage influence result L is numerically converted to obtain a numerical conversion result N, for example, L=0.7 is converted to N=70; the coefficient correlation weight information is fused, for example, the semantic deviation weight W1=0.4, the data adaptation state weight W2=0.3, and the dimension conversion correlation strength weight W3=0.3, and the second error correction value=N×(W1×semantic deviation coefficient+W2×data adaptation coefficient+W3×dimension conversion coefficient). Taking the restaurant ordering scenario as an example, the semantic deviation coefficient is 0.8, the data adaptation coefficient is 0.6, and the dimension conversion coefficient is 0.7. Substituting the data into the equation, the second error correction value=70×(0.4×0.8+0.3×0.6+0.3×0.7)=70×(0.32+0.18+0.21)=70×0.71=49.7.
[0113] The target comparison verification benchmark, the first error correction value, and the second error correction value are processed to obtain actual error correction values one and two, which specifically include the following steps:
[0114] The benchmark verification specification is extracted from the target comparison verification benchmark; wherein the benchmark verification specification includes a judgment standard for semantic accuracy, a compliance range for data adaptation, and a connection requirement for dimension conversion;
[0115] After comparing the second error correction value with the benchmark verification specification, a deviation component that exceeds the benchmark compliance range is obtained, and a deviation correction factor is formed after judging the influence amplitude of the deviation component;
[0116] calibrating the second error correction value according to the deviation correction factor to obtain a second calibrated error correction value; and judging a correlation fit degree between a scene adaptation deviation of the first error correction value and a data dimension deviation of the second calibrated error correction value to obtain a complementary feature;
[0117] The first error correction value is taken as a leading component, and the second calibrated error correction value is taken as a deviation supplement component, and an actual error correction value one is obtained through component superposition, in a manner of taking the scene adaptation priority as a guide.
[0118] The second calibrated error correction value is taken as a check component, and the first error correction value is taken as a scene adaptation constraint component, and an actual error correction value two is obtained through component balance and complementary cooperation, in a manner of taking the data dimension deviation as a core.
[0119] A reference check specification is extracted from a target comparison verification benchmark. The target comparison verification benchmark is a check rule set matching the scene, for example, a judgment standard of semantic accuracy is that a point order sentence needs to completely contain an item and a quantity, and is judged as a semantic deviation if any part is missing; for example, a compliance range of data adaptation is that a pronunciation duration needs to be between 0.5-1 seconds, and a noise level needs to be lower than 50 decibels; for example, a requirement of connection of dimension conversion from a pronunciation duration dimension to a semantic integrity dimension needs to guarantee a matching degree of more than 90%.
[0120] A deviation correction factor and the second calibrated error correction value are obtained by processing the second error correction value, assuming that a noise level corresponding to the second error correction value is 55 decibels, exceeding a compliance range by 5 decibels, and a pronunciation duration is 0.4 seconds, exceeding a compliance range by 0.1 second, and the values of these exceeding parts are deviation components. An influence amplitude of the deviation components is judged: an influence amplitude of the noise level exceeding a standard is (55-50) / 50=0.1, an influence amplitude of the pronunciation duration exceeding a standard is (0.5-0.4) / 0.5=0.2, and a mean value of the two is taken to obtain the deviation correction factor F=0.15. The second error correction value is calibrated by using the deviation correction factor, and the second calibrated error correction value=the second error correction value x (1-F), and the second calibrated error correction value is obtained as 49.7 x (1-0.15)=42.245 by substituting the numerical value.
[0121] A correlation fit degree of the first error correction value and the second calibrated error correction value is judged. The first error correction value of the restaurant point order scene is a deviation value obtained by combining a standard scene (a general point order specification) and an adaptation portrait scene (user habit simplified expression), for example, the deviation value is 35. A scene adaptation deviation corresponding to the first error correction value is that the user habit says 'Han' instead of 'Hamburger', and a data dimension deviation corresponding to the second calibrated error correction value is, for example, noise and pronunciation duration exceeding a standard, and the correlation fit degree of the two is that the simplified expression and the data exceeding a standard jointly amplify semantic deviation, and thus a complementary feature is obtained, that is, the scene adaptation deviation and the data dimension deviation have a superposition influence.
[0122] In the restaurant ordering scenario, if the user is a regular customer and the scene adaptation priority is high, the first error correction value is used as the dominant component, and the second calibration error correction value is used as the deviation supplement component. The actual error correction value one is calculated by component superposition: actual error correction value one = first error correction value x 0.7 + second calibration error correction value x 0.3. Substituting the numerical value, the actual error correction value one is 35 x 0.7 + 42.245 x 0.3 = 24.5 + 12.6735 = 37.1735.
[0123] If the restaurant is in peak period, the data dimension deviation is more significant, and the second calibration error correction value is used as the verification component, and the first error correction value is used as the scene adaptation constraint component. The actual error correction value two is calculated as: actual error correction value two = second calibration error correction value x 0.8 + first error correction value x 0.2. Substituting the numerical value, the actual error correction value two is 42.245 x 0.8 + 35 x 0.2 = 33.796 + 7 = 40.796.
[0124] According to the actual error correction value one or the actual error correction value two, the target speech transcription text is corrected to obtain the correct recognition result, which includes the following steps:
[0125] According to the actual error correction value one or the actual error correction value two, the corresponding scheme of the multi-modal error correction strategy library is matched to obtain the target error correction scheme;
[0126] According to the target error correction scheme, the deviation content of the target speech transcription text is corrected to obtain the correct recognition result.
[0127] The multi-modal error correction strategy library is a collection of error correction schemes covering different scenes and different deviation types. In the restaurant ordering scenario, the actual error correction value one or the actual error correction value two is matched to the corresponding scheme according to the numerical value and characteristics. For example, the actual error correction value one corresponds to the scene adaptation priority deviation, and the scheme that combines user expression habit correction and data dimension deviation fine-tuning is matched. The actual error correction value two corresponds to the data dimension priority deviation, and the scheme that first calibrates the data dimension feature and constrains the scene adaptation deviation is matched.
[0128] Taking the scheme corresponding to the actual error correction value one as an example, the semantic correction is the scene adaptation deviation of replacing "Hamburger" with "Han" according to user habits, and the commonly used vocabulary library in the ordering scenario is used to supplement "Han" to "Hamburger" in the transcription text. The data dimension fine-tuning is aimed at the deviation of short pronunciation duration and excessive noise level, so as to adjust the syllable integrity of the transcription text, such as supplementing "Kele" to "Kele" corresponding to the pronunciation duration of 0.4 seconds.
[0129] The target speech transcription text is corrected. Assuming that the target speech transcription text of the restaurant ordering scene is "I want a hamburger", the deviation content includes: the semantic level "H" and "K" belong to incomplete ideographic, and the data dimension level is truncated due to short pronunciation duration and high noise. According to the target correction scheme, the semantic deviation is processed, and after calling the vocabulary library of the ordering scene, "H" corresponds to "hamburger" and "K" corresponds to "cola"; the data dimension deviation is processed, and the truncated syllable is completed according to the average pronunciation duration in the same scene.
[0130] The final correct recognition result is "I want a hamburger and coke". The actual correction value determines the priority of the correction strategy in the whole process, and the target correction scheme specifically covers the semantic and data dimension deviations, which takes into account the user's expression habits and ensures the semantic accuracy and scene adaptability of the text.
[0131] If the scheme corresponding to the actual correction value two is processed, the data dimension feature is first calibrated, the noise interference on speech recognition is weakened, and then the syllable length of the transcription text is corrected according to the compliance range of pronunciation duration (0.5-1 second). "H" and "K" are first completed to "hamburger" and "cola" corresponding to complete syllables, and then the scene adaptation deviation is constrained, so that the text conforms to the general semantic specification of the ordering scene, and finally the accurate recognition result is obtained.
[0132] A multi-modal speech recognition correction system, comprising:
[0133] The processing module obtains a first correction value of the target speech transcription text in the standard scene and the adaptive image scene according to the target speech transcription text to be corrected and the corresponding multi-modal associated data.
[0134] The first extraction module extracts a first associated influence coefficient of the recognition deviation affected by the non-adaptive variation state if different data dimension features in the historical multi-modal associated data are in a non-adaptive variation state.
[0135] The second extraction module extracts a second associated influence coefficient of the recognition deviation affected by the cross-dimension information adaptive conversion if different data dimension features in the historical multi-modal associated data are in an adaptive stable state, and extracts a third associated influence coefficient of the recognition deviation affected by the same-dimension internal feature adaptive conversion.
[0136] The judgment module obtains a second correction value according to the deviation degree of the target speech transcription text in the data adaptive state and the dimension conversion influence according to the first associated influence coefficient, the second associated influence coefficient and the third associated influence coefficient.
[0137] The output module: the target comparison verification datum is screened from the comprehensive comparison verification datum, and the target speech transcription text error correction verification is performed. The target comparison verification datum, the first error correction value and the second error correction value are processed to obtain actual error correction value one and actual error correction value two. The target speech transcription text is corrected according to the actual error correction value one or the actual error correction value two to obtain the correct recognition result.
[0138] The device embodiments described above are only schematic, wherein the units described as separate components can or can not be physically separated, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed on multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the embodiment scheme. Those skilled in the art can understand and implement without creative labor.
[0139] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be realized by means of software and the necessary general hardware platform, and of course, it can also be realized by hardware. Based on such understanding, the above technical solutions can be embodied in the form of software products, which can be stored in a computer readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in each embodiment or some parts of the embodiment.
[0140] Finally, it should be noted that: the above examples are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing examples, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing examples, or make equivalent replacement for some technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A multi-modal speech recognition error correction method, characterized by, The method comprises the following steps: According to the target speech transcription text to be corrected and the corresponding multi-modal association data, the first correction value of the target speech transcription text in the standard scene and the adaptive image scene is obtained, specifically comprising the following steps: The semantic structure between the text core ideographic element and the semantic association is obtained by performing semantic splitting processing on the target speech transcription text; the multi-dimensional feature system is constructed by separating the sound features, environment features related to the use environment, and individual features of the user in the multi-modal association data; A standard scene verification benchmark is built according to the semantic structure and the multi-dimensional feature system, the standard scene verification benchmark integrates general language expression specifications, conventional environment voice transmission rules and semantic logic criteria, the semantic structure is compared with the standard scene verification benchmark to obtain a semantic deviation item, and after determining the number and influence degree of the semantic deviation item, a standard scene deviation information is formed; An adaptive image scene verification benchmark is constructed according to the individual features, language expression habits of the user, expression preferences in specific scenes, and commonly used semantic combination methods, and after matching the semantic structure with the adaptive image scene verification benchmark, adaptive deviation items that conform to general specifications but are contrary to the expression habits of the user are identified, and after judging the necessity of correction and the associated influence range of the adaptive deviation items, adaptive image scene deviation information is obtained; Based on the standard scene deviation information and the adaptive image scene deviation information, a deviation weight distribution rule is established, the influence weights of the standard scene deviation information and the adaptive image scene deviation information are determined according to the scene use priority and the specification requirements, and the standard scene deviation information and the adaptive image scene deviation information are fused according to the weight distribution rule to obtain a first correction value reflecting the scene adaptability and the semantic accuracy; If the different data dimension features in the historical multi-modal association data show a non-adaptive variation state, a first associated influence coefficient for identifying deviation is extracted; If the different data dimension features in the historical multi-modal association data show an adaptive stable state, a second associated influence coefficient for identifying deviation of cross-dimension information adaptive conversion is extracted, and a third associated influence coefficient for identifying deviation of internal features of the same dimension adaptive conversion is extracted; According to the first associated influence coefficient, the second associated influence coefficient and the third associated influence coefficient, the deviation degree of the target speech transcription text in the data adaptive state and the dimension conversion influence is judged to obtain a second correction value; The target comparison verification benchmark for correcting and verifying the target speech transcription text is selected from the comprehensive comparison verification benchmark, the target comparison verification benchmark, the first correction value and the second correction value are processed to obtain actual correction value one and actual correction value two, and the target speech transcription text is corrected according to the actual correction value one or the actual correction value two to obtain a correct recognition result.
2. The multi-modal speech recognition error correction method of claim 1, wherein, The multi-modal association data includes pronunciation duration, syllable interval, intonation change, accent type, language habit, noise level, scene type and dialogue scene.
3. The multi-modal speech recognition error correction method of claim 1, wherein, The first associated influence coefficient for identifying deviation of non-adaptive variation state is extracted, specifically comprising the following steps: A historical transcription text one is selected from all historical speech transcription texts; extracting recognition bias rate data one belonging to a dependent data dimension from the historical transcription text one; obtaining a first correlation influence coefficient of a non-adaptation variation state affecting recognition bias according to the non-adaptation variation degree value of the recognition bias rate data one.
4. The multi-modal speech recognition error correction method of claim 3, wherein, extracting a second correlation influence coefficient of cross-dimension information adaptation conversion affecting recognition bias, specifically including the following steps: selecting a historical transcription text two from all historical voice transcription texts; extracting recognition bias rate data two in the core data dimension and the auxiliary data dimension from the historical transcription text two; obtaining the second correlation influence coefficient of cross-dimension information adaptation conversion affecting recognition bias according to the data adaptation stability value of the recognition bias rate data two.
5. The multi-modal speech recognition error correction method of claim 1, wherein, extracting a third correlation influence coefficient of same-dimension internal feature adaptation conversion affecting recognition bias, specifically including the following steps: extracting recognition bias rate data three in the dependent data dimension, the auxiliary data dimension and the core data dimension from the historical transcription text two; obtaining the third correlation influence coefficient of same-dimension internal feature adaptation conversion affecting recognition bias according to the data adaptation stability value of the recognition bias rate data three.
6. The multi-modal speech recognition error correction method of claim 5, wherein, obtaining a second correction value of the bias degree of the data adaptation state and the dimension conversion influence of the target voice transcription text according to the first correlation influence coefficient, the second correlation influence coefficient and the third correlation influence coefficient, specifically including the following steps: performing characteristic judgment on the first correlation influence coefficient, the second correlation influence coefficient and the third correlation influence coefficient to obtain coefficient characteristic condition information; wherein the coefficient characteristic condition information includes bias influence type, action range and conduction path; obtaining coefficient correlation weight information according to the semantic structure of the target voice transcription text and the feature system of the multi-modal correlation data; wherein the coefficient correlation weight information includes semantic bias, data adaptation state and dimension conversion correlation strength; establishing a coefficient influence conduction link according to the coefficient characteristic condition information and the coefficient correlation weight information; obtaining a coefficient linkage influence result according to the bias influence situation of the first correlation influence coefficient, the second correlation influence coefficient and the third correlation influence coefficient according to the coefficient influence conduction link; constructing a bias degree quantization standard based on the coefficient linkage influence result; wherein the bias degree quantization standard combines the semantic tolerance range of the general language expression, the reasonable fluctuation interval of the multi-modal data adaptation and the normal connection requirement of the dimension conversion, compares the coefficient linkage influence result with the bias degree quantization standard to obtain a quantization index corresponding to a bias degree level; performing numerical conversion on the coefficient linkage influence result according to the quantization index to obtain a numerical conversion result, and adjusting the numerical conversion result by fusing the coefficient correlation weight information to obtain a second correction value of the data adaptation state bias and the dimension conversion influence degree.
7. The multi-modal speech recognition error correction method of claim 6, wherein, processing the target comparison verification benchmark, the first correction value and the second correction value to obtain actual correction value one and actual correction value two, specifically including the following steps: extracting a benchmark verification specification from the target comparison verification benchmark; wherein the benchmark verification specification includes a judgment standard of semantic accuracy, a compliance range of data adaptation and a connection requirement of dimension conversion; The second error correction value is compared with the benchmark verification specification to obtain a deviation component that exceeds the benchmark compliance range in the second error correction value, and a deviation correction factor is formed after judging the influence amplitude of the deviation component; The second error correction value is calibrated according to the deviation correction factor to obtain a second calibrated error correction value; and the complementary feature is obtained by judging the associated fit degree of the scene adaptation deviation to which the first error correction value belongs and the data dimension deviation to which the second calibrated error correction value belongs; The first error correction value is taken as a leading component, and the second calibrated error correction value is taken as a deviation supplementary component, and an actual error correction value one is obtained through component superposition, with the scene adaptation priority as a guide; The second calibrated error correction value is taken as a verification component, and the first error correction value is taken as a scene adaptation constraint component, and an actual error correction value two is obtained through component balance and complementary cooperation, with the data dimension deviation as a core.
8. The multi-modal speech recognition error correction method of claim 1, wherein, The correct recognition result is obtained by performing error correction processing on the target speech transcription text according to the actual error correction value one or the actual error correction value two, specifically including the following steps: The target error correction scheme is obtained by matching the corresponding scheme of the multi-modal error correction strategy library according to the actual error correction value one or the actual error correction value two; The correct recognition result is obtained by performing error correction processing on the deviation content of the target speech transcription text according to the target error correction scheme.
9. A multi-modal speech recognition error correction system, applied to the multi-modal speech recognition error correction method of any one of claims 1-8, characterized in that, It includes: The processing module obtains a first error correction value of the target speech transcription text in a standard scene and an adaptation image scene according to the target speech transcription text to be corrected and corresponding multi-modal associated data, specifically including the following steps: The semantic structure between the text core ideographic element and the semantic association is obtained by performing semantic splitting processing on the target speech transcription text; a multi-dimensional feature system is constructed by separating the sound production features of the speech expression, the environment features related to the use environment, and the individual features of the user in the multi-modal associated data; The standard scene verification benchmark is built according to the semantic structure and the multi-dimensional feature system; the standard scene verification benchmark integrates the general language expression specification, the speech transmission rule of the conventional environment, and the semantic logic criterion; the semantic deviation items are obtained by comparing the semantic structure with the standard scene verification benchmark; the number and influence degree of the semantic deviation items are determined to form the standard scene deviation information; The adaptation image scene verification benchmark is constructed according to the individual features of the multi-dimensional feature system, the language expression habits of the user, the expression preferences of the specific scene, and the commonly used semantic combination methods; the adaptation deviation items that conform to the general specification and are contrary to the expression habits of the user are identified by matching the semantic structure with the adaptation image scene verification benchmark after the matching; the necessity of correction and the associated influence range of the adaptation deviation items are judged to obtain the adaptation image scene deviation information; The deviation weight distribution rule is established based on the standard scene deviation information and the adaptation image scene deviation information; the influence weights of the standard scene deviation information and the adaptation image scene deviation information are determined according to the scene use priority and the specification requirements; the standard scene deviation information and the adaptation image scene deviation information are fused to obtain the first error correction value reflecting the scene adaptability and the semantic accuracy according to the weight distribution rule; The first extraction module extracts the first associated influence coefficient of the recognition deviation of the non-adaptation variation state if the different data dimension features in the historical multi-modal associated data are in a non-adaptation variation state. The second extraction module: if the different data dimension features in the historical multi-modal correlation data are in an adaptive stable condition, a second correlation influence coefficient of cross-dimension information adaptive conversion affecting recognition bias is extracted, and a third correlation influence coefficient of same-dimension internal feature adaptive conversion affecting recognition bias is extracted; The judgment module: according to the first correlation influence coefficient, the second correlation influence coefficient and the third correlation influence coefficient, the judgment module judges the deviation degree of the target speech transcription text in the data adaptive state and the dimension conversion influence to obtain a second correction value; The output module: the target comparison and verification benchmark for the target speech transcription text correction verification is selected from the comprehensive comparison and verification benchmark, the target comparison and verification benchmark, the first correction value and the second correction value are processed to obtain actual correction value one and actual correction value two, and the target speech transcription text is corrected according to the actual correction value one or the actual correction value two to obtain a correct recognition result.
Citation Information
Patent Citations
Exhibition hall intelligent guide voice question and answer optimization system based on multi-modal data
CN121122252A
Conversation-based multi-modal feature analysis method and electronic equipment
CN121144772A