An AI-Assisted Cross-Language Automatic Note-Taking and Term Annotation System

Through artificial intelligence assisted cross-language automatic notes and term labeling systems, the problem of cross-language term recognition and labeling in complex audio environments is solved, and efficient and accurate information recording is achieved.

CN120183408BActive Publication Date: 2025-07-25SHAANXI RAILWAY INST
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510670653.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-07-25
Estimated Expiration
2045-05-23

AI Technical Summary

Technical Problem

The prior art is difficult to accurately identify and label professional terms in complex audio environments, especially in multilingual or mixed language scenarios, resulting in the impact of the accuracy and practical value of automatic notes.

Method used

The artificial intelligence-assisted cross-language automatic notes and term labeling system is adopted to accurately identify and annotate cross-language terms through audio processing, cross-language speech recognition, text proofreading, term candidate recognition, cross-language term matching analysis and labeling decision-making modules.

Benefits of technology

It significantly improves the efficiency and accuracy of information recording in cross-language communication scenarios, and generates high-quality calibration note texts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120183408B_ABST
    Figure CN120183408B_ABST
Patent Text Reader

Abstract

The present invention discloses an artificial intelligence-assisted cross-language automatic note-taking and term annotation system, belonging to the technical field of language annotation. The system first acquires and processes audio data, generates preliminary text through cross-language speech recognition, identifies term candidate words in the text, and performs pronunciation association matching with a term knowledge base. It performs language recognition and context analysis on the candidate words and their contexts, calculates a comprehensive score by integrating pronunciation matching, language recognition, and context scores, so as to accurately select the cross-language term that best matches the current context as the final candidate word, and obtains its annotation information. Integrate the final candidate word and its annotation information into the calibrated note text, and output an automatic note with cross-language term annotation. The present invention can automatically process complex audio including language mixing, generate high-quality calibrated note text, and significantly improve the information recording efficiency and accuracy in cross-language communication scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of language annotation, and particularly to an artificial intelligence-assisted cross-language automatic note-taking and term annotation system. Background Art

[0002] In the field of automated note generation, generating text from audio through speech recognition technology has received extensive attention and application. However, in the face of complex audio environments, such as the presence of background noise, multiple speakers, especially when dealing with speech inputs containing different languages or even mixed languages, accurately performing speech-to-text itself is challenging. On this basis, automatically identifying and annotating specific information in the transcribed text, such as technical terms, proper nouns, or cross-language concepts, is even more difficult for the prior art. Current automated term annotation methods often rely on simple text matching, statistical features, or rules based on limited context, and it is difficult to effectively handle the pronunciation variations of terms, the meaning differences in different contexts, and the complexity in cross-language scenarios. Especially when terms appear in non-target languages or mixed languages in the original speech, existing methods are prone to misidentification, missed annotation, or providing inaccurate annotation information, seriously affecting the accuracy and practical value of automatic notes. Summary of the Invention

[0003] To address the above technical problems, an artificial intelligence-assisted cross-language automatic note-taking and term annotation system of the present invention includes an audio processing and cross-language speech recognition module, which is used to receive audio data, perform preprocessing, and perform cross-language speech recognition on the preprocessed audio data to output preliminary text; a text proofreading and term candidate recognition module, which is used to receive the preliminary text, perform reciprocating matching and proofreading to output calibrated note text, and identify term candidate words from the preliminary text or the calibrated note text; a cross-language term matching analysis and annotation decision-making module, which is used to receive the term candidate words and the calibrated note text, obtain the pronunciation representation of the term candidate words, compare with the cross-language pronunciation representations of known terms in the term knowledge base and pronunciation model library to generate a preliminary cross-language term matching list, perform language recognition on the term candidate words or their contexts, perform context analysis on each known term in the preliminary cross-language term matching list and calculate the context matching score, calculate the comprehensive score by integrating the pronunciation association matching result, language recognition result and context matching score, select the final candidate word that best matches the current context, and obtain the annotation information of the final candidate word from the term knowledge base and pronunciation model library; an automatic note generation and annotation output module, which is used to receive the calibrated note text and the final candidate word and its annotation information, integrate the final candidate word and its annotation information into the calibrated note text, and generate an automatic note with cross-language term annotation. The term knowledge base and pronunciation model library are used to store terms, meanings and corresponding cross-language pronunciation representations in different languages, and provide query support for the cross-language term matching analysis and annotation decision-making module.

[0004] Further, the automatic note-taking and term annotation method includes the following steps:

[0005] Step S1, obtain and process audio data, and perform cross-language speech recognition to obtain preliminary text.

[0006] Step S2, perform reciprocating matching and proofreading on the preliminary text to generate calibrated note text, and identify term candidate words.

[0007] Step S3, when term candidate words are recognized, perform cross-language feature judgment, pronunciation association matching and context analysis, and comprehensively score to determine the cross-language terms to be annotated and their information.

[0008] Step S4, integrate the final candidate word and its annotation information into the note text, and output the note text with annotations.

[0009] Further, the step S1 includes,

[0010] Step S101: Obtain a multi-channel or single-channel audio data stream to be processed through an audio input device. The audio data stream may include voices of different speakers, mixed languages, and environmental noises.

[0011] Step S102: Perform enhanced preprocessing and acoustic feature extraction on the audio data. The preprocessing at least includes multi-channel signal processing, voice activity detection, noise reduction, acoustic environment adaptation, gain adjustment, and frame segmentation. The acoustic feature extraction is used to obtain a sequence of feature vectors suitable for speech recognition.

[0012] Step S103: Use a cross-language speech recognition model based on a deep learning model to perform speech-to-text operations on the sequence of acoustic feature vectors. The cross-language speech recognition model can process speech inputs containing mixed languages and output confidence information of the speech recognition result and / or language recognition confidence information of the text segment during the recognition process; output a preliminary text stream or text segment.

[0013] Furthermore, step S2 includes:

[0014] Step S201: Chunk the preliminary text output in step S1 according to a preset time length or semantic boundary to form text segments to be proofread and term-identified.

[0015] Step S202: Adopt a context-dependent reciprocating matching mode, and use the speech recognition result corresponding to the subsequent audio received and / or subsequent text context information to perform operations such as re-scoring, re-searching, or post-processing based on a language model on the currently and previously transcribed text segments, and dynamically correct the text content in combination with the confidence information of the speech recognition result to generate a more accurate calibrated note text.

[0016] Step S203: Scan the preliminary text or the calibrated note text to identify words or multi-word phrases that meet the preset features as the term candidate words.

[0017] Furthermore, step S3 includes:

[0018] Step S301: Extract the current text context where the term candidate words are located.

[0019] Step S302: Obtain the pronunciation representation of the term candidate words, compare it with the pronunciation representations of known cross-language terms in the term knowledge base, perform pronunciation-related word center matching, and generate a preliminary cross-language term matching list.

[0020] Step S303: Based on the language recognition result of the term candidate or its context, and the center matching result of the pronunciation correlation word, preliminarily judge the cross-language feature of the term candidate, and analyze the compliance degree of each known term in the preliminary cross-language term matching list with the context, and calculate the context matching score.

[0021] Step S304: Based on the center matching result of the pronunciation correlation word, the language recognition result and the context matching score, calculate the comprehensive score of each known term in the preliminary cross-language term matching list, and select the known term with the highest score and higher than the preset threshold as the final candidate word according to the comprehensive score. At the same time, obtain the annotation information of the final candidate word from the term knowledge base.

[0022] Further, the step S4 includes:

[0023] Step S401: According to the characteristics or the comprehensive score of the final candidate word obtained in step S3, determine the annotation level or display priority of the final candidate word; and integrate the final candidate word and its annotation information into the position where the corresponding word appears in the calibration note text in the preset integration format corresponding to the determined annotation level or display priority. The preset integration format at least includes inserting a specific mark in the calibration note text, changing the word style, adding a hover prompt, generating a sidebar note, creating an expandable detail.

[0024] Step S402: Output the finally generated automatic note text with cross-language term annotation. The automatic note text includes the calibration note text and the structured annotation data associated with the final candidate word to support subsequent viewing, editing or interactive display. And the output format of the automatic note text at least includes rich text format, HTML format, Markdown format.

[0025] Further, in step S304, when calculating the comprehensive score of each known term in the preliminary cross-language term matching list, the following formula is adopted:

[0026] , where represents the term candidate, represents the current text context, represents the th known term in the preliminary cross-language term matching list, and the known term is from the term knowledge base; represents the pronunciation representation of the term candidate and the known term The pronunciation similarity score between cross - language pronunciation representations, where the pronunciation representations are obtained based on acoustic feature extraction of the speech of the term candidate, phoneme sequence prediction, or generation of pronunciation vectors, and the cross - language pronunciation representations are the standard pronunciation representations of the known term in different languages, and the pronunciation similarity score is obtained by calculating the distance or similarity metric between the pronunciation representations, and this score is greater than 0; representing the known term and the current text context between the context matching scores, where the context matching scores are calculated by analyzing the semantic relevance between the known term and the context and the semantic relevance analysis can be implemented based on word vectors, language models, topic models, and text matching algorithms, and this score value is greater than 0; representing the language recognition confidence score that the system determines the term candidate belongs to the main language of the known term and the confidence score is obtained by performing language recognition analysis on the term candidate itself or its context and this score value is greater than 0 and less than or equal to 1; representing the known term presetting the main language label in the term knowledge base; representing the weight coefficient, which is used to adjust the contribution ratio of each score item in the comprehensive score; representing a small constant greater than zero, which is used to avoid calculation anomalies when taking the logarithm of the score item.

[0027] Furthermore, the weight coefficient is determined according to a preset strategy, and the strategy is used to balance the relative importance of pronunciation similarity, context matching, and language recognition in the comprehensive evaluation to optimize the recognition accuracy of cross - language terms.

[0028] The beneficial effects of the present invention compared with the prior art are: (1) The present invention provides an artificial intelligence - assisted method that can automatically process complex audio containing language mixing and generate high - quality calibrated note text. (2) By integrating pronunciation, context, and language recognition information, the present invention can accurately identify and label cross - language terms, significantly improving the information recording efficiency and accuracy in cross - language communication scenarios. Brief Description of the Drawings

[0029] Figure 1 is an exemplary step - by - step flowchart of the automatic note and term annotation method of the present invention.

[0030] Figure 2This is an exemplary flowchart of the steps for processing audio data in the present invention.

[0031] Figure 3 This is an exemplary flowchart of the steps for performing reciprocating proofreading in the present invention.

[0032] Figure 4 This is an exemplary flowchart of the steps for determining candidate terms in the present invention.

[0033] Figure 5 This is an exemplary flowchart of the steps for integrating text results in the present invention. Detailed implementation manners

[0034] This application provides an artificial intelligence-assisted cross-language automatic note-taking and term annotation system, which aims to automatically generate notes with accurate cross-language term annotations by processing audio data from different sources, significantly improving the information recording efficiency and quality in scenarios such as cross-language communication, learning, and meeting minutes.

[0035] In one embodiment, the system may include an audio processing and cross-language speech recognition module, which is used to receive audio data, perform preprocessing, and perform cross-language speech recognition on the preprocessed audio data to output preliminary text; a text proofreading and term candidate recognition module, which is used to receive the preliminary text, perform reciprocating matching and proofreading to output calibrated note text, and identify candidate terms from the preliminary text or the calibrated note text; a cross-language term matching analysis and annotation decision-making module, which is used to receive the candidate terms and the calibrated note text, obtain the pronunciation representation of the candidate terms, compare it with the cross-language pronunciation representations of known terms in the term knowledge base and the pronunciation model library to generate a preliminary cross-language term matching list, perform language recognition on the candidate terms or their contexts, perform context analysis on each known term in the preliminary cross-language term matching list and calculate the context matching score, calculate the comprehensive score by integrating the pronunciation association matching result, the language recognition result, and the context matching score, select the final candidate term that best fits the current context, and obtain the annotation information of the final candidate term from the term knowledge base and the pronunciation model library; an automatic note generation and annotation output module, which is used to receive the calibrated note text and the final candidate term and its annotation information, integrate the final candidate term and its annotation information into the calibrated note text, and generate automatic notes with cross-language term annotations. The term knowledge base and the pronunciation model library are used to store terms, meanings, and corresponding cross-language pronunciation representations in different languages, and provide query support for the cross-language term matching analysis and annotation decision-making module.

[0036] As Figure 1 shown, this is an exemplary flowchart of the steps of the automatic note-taking and term annotation method in this embodiment, including the following steps:

[0037] Step S1: Obtain and process audio data, and perform cross - language speech recognition to obtain preliminary text.

[0038] Step S2: Perform reciprocating matching and proofreading on the preliminary text to generate a calibrated note text, and identify term candidate words.

[0039] Step S3: When term candidate words are identified, perform cross - language feature judgment, pronunciation - related matching, and context analysis, and comprehensively score to determine the cross - language terms to be annotated and their information.

[0040] Step S4: Integrate the final candidate words and their annotation information into the note text, and output the note text with annotations.

[0041] As Figure 2 shown is an exemplary step - flow diagram for processing audio data in step S1 of this embodiment, including,

[0042] Step S101: Through an audio input device, obtain a multi - channel or single - channel audio data stream to be processed. The audio data stream may contain voices of different speakers, mixed languages, and environmental noise.

[0043] In one embodiment, the technical solution of step S101 is elaborated: The system receives the original audio data through various types of audio input devices. Exemplarily, these devices may include built - in microphones, external microphones, microphone arrays, voice recorders, conference systems, or audio streams received through a network. The obtained audio data stream can be single - channel or multi - channel, and may contain voices from different speakers, a mixture of voices in multiple languages, and environmental background noise, such as the hum of a meeting room, keyboard typing sounds, or other interference sounds.

[0044] Step S102: Perform enhanced pre - processing and acoustic feature extraction on the audio data. The pre - processing at least includes multi - channel signal processing, voice activity detection, noise reduction, acoustic environment adaptation, gain adjustment, and frame segmentation. The acoustic feature extraction is used to obtain a sequence of feature vectors suitable for speech recognition.

[0045] In one embodiment, the technical solution of step S102 is elaborated: The received original audio data first undergoes enhanced pre - processing. The pre - processing can at least include: If the audio is multi - channel, perform multi - channel signal processing, exemplarily, such as beamforming to enhance the speech in a specific direction; perform voice activity detection to identify the segments with speech in the audio; apply a noise reduction algorithm to reduce background noise; perform acoustic environment adaptation adjustment to adapt to different recording environments; perform gain adjustment to standardize the loudness of the audio signal. The pre - processed audio is segmented into short - time frames, and acoustic features suitable for speech recognition are extracted from each frame.

[0046] Step S103: Use a cross - language speech recognition model based on a deep - learning model to perform speech - to - text operation on the acoustic feature vector sequence. The cross - language speech recognition model can process speech inputs containing mixed languages and output confidence information of the speech recognition result and / or language recognition confidence information of text segments during the recognition process; output a preliminary text stream or text segments.

[0047] In one embodiment, the technical solution of step S103 is elaborated: The extracted acoustic feature vector sequence is input into a cross - language speech recognition model based on a deep - learning architecture. This model is trained to process speech inputs containing different languages or even mixed languages. Exemplarily, the model can adopt connectionist temporal classification, attention mechanism, sequence - to - sequence model, or a Transformer - based architecture. When performing speech - to - text operation, the model outputs a preliminary text stream or outputs text results in segments. In addition, to improve the accuracy of subsequent processing, the model can output the recognition confidence score of each word or text segment during the recognition process, and perform language recognition on each text segment and output its corresponding language confidence information. For example, determine whether the segment is Chinese, English, or other languages, and the probability of recognition.

[0048] As Figure 3 shown, an exemplary step - flow diagram for reciprocating proofreading in step S2 of this embodiment includes:

[0049] Step S201: Block - process the preliminary text output in step S1 according to a preset time length or semantic boundary to form text segments to be proofread and term - recognized.

[0050] In one embodiment, the preliminary text stream output from step S1 is segmented into more easily processed text segments. Exemplarily, the basis for blocking can be a fixed time length, detected end - of - sentence punctuation marks, or other semantic boundaries identified through natural language processing techniques.

[0051] Step S202: Adopt a context - dependent reciprocating matching pattern, and use the speech recognition result corresponding to the received subsequent audio and / or subsequent text context information to perform operations such as re - scoring, re - searching, or post - processing based on a language model on the currently and previously transcribed text segments, and dynamically correct the text content in combination with the confidence information of the speech recognition result to generate a more accurate calibrated note text.

[0052] In one embodiment, the system employs an iterative or reciprocal proofreading mechanism to improve the accuracy of the preliminary text. Exemplarily, using the recognition results of subsequently received audio segments and longer text context information, the system can re-evaluate the processed text segments. This may include re-scoring the acoustic or language model based on a broader context, re-searching in the recognition network to find a better word sequence, or applying post-processing steps based on a large language model for grammar, spelling, and semantic corrections. Combining with the speech recognition confidence information output in step S103, the system can dynamically identify and correct potential errors, thereby generating a more accurate calibrated note text.

[0053] Step S203: Scan the preliminary text or the calibrated note text, and identify words or multi-word phrases that conform to the preset features as candidate terms.

[0054] Such as Figure 4 Shown is an exemplary step flow chart for determining candidate terms in step S3 of this embodiment, including:

[0055] Step S301: Extract the current text context where the candidate term is located.

[0056] In one embodiment, for each candidate term identified in step S203, the system extracts the surrounding text information as its current context. Exemplarily, this context can be a complete sentence, the current paragraph containing the candidate term, or a fixed number of words or sentences before and after the candidate term. Extracting the context is for context relevance analysis in subsequent steps.

[0057] Step S302: Obtain the pronunciation representation of the candidate term, compare it with the pronunciation representations of known cross-language terms in the term knowledge base, perform pronunciation-related term center matching, and generate a preliminary cross-language term matching list.

[0058] In one embodiment, the system obtains the pronunciation representation of the candidate term. This can be achieved by performing acoustic analysis on the segment in the original audio corresponding to the candidate term, extracting acoustic features or phoneme sequences, or generating its representation in a certain standard phonetic notation through a text-to-phoneme model. Then, compare this pronunciation representation with the standard pronunciation representations of known cross-language terms stored in the term knowledge base and the pronunciation model base. By calculating the pronunciation similarity, perform pronunciation-related matching, and generate a preliminary cross-language term matching list, which contains the known terms with pronunciations similar to the candidate term and their similarity scores.

[0059] Step S303: Based on the language recognition results of the term candidate or its context, and the center matching results of pronunciation-related words, preliminarily judge the cross-language features of the term candidate, and analyze the compliance of each known term in the preliminary cross-language term matching list with the context, and calculate the context matching score.

[0060] Step S304: Based on the center matching results of pronunciation-related words, language recognition results, and context matching scores, calculate the comprehensive scores of each known term in the preliminary cross-language term matching list, and select the known term with the highest score and higher than the preset threshold as the final candidate word, and obtain the annotation information of the final candidate word from the term knowledge base.

[0061] In Step S304, to calculate the comprehensive scores of each known term in the preliminary cross-language term matching list, the following formula is used:

[0062] , where represents the term candidate, represents the current text context, represents the th known term in the preliminary cross-language term matching list, and the known term is from the term knowledge base; represents the pronunciation similarity score between the pronunciation representation of the term candidate and the cross-language pronunciation representation of the known term . The pronunciation representation is obtained by extracting acoustic features, predicting phoneme sequences, or generating pronunciation vectors from the speech of the term candidate. The cross-language pronunciation representation is the standard pronunciation representation of the known term in different languages. The pronunciation similarity score is obtained by calculating the distance or similarity measure between the pronunciation representations, and this score is greater than 0; represents the context matching score between the known term and the current text context . The context matching score is calculated by analyzing the semantic relevance between the known term and the context . The semantic relevance analysis can be implemented based on word vectors, language models, topic models, and text matching algorithms, and this score value is greater than 0; represents the language recognition confidence score that the system judges the term candidate belongs to the main language of the known term . The confidence score is obtained by performing language recognition analysis on the term candidate itself or its context . This score value is greater than 0 and less than or equal to 1; represents the preset main language label of the known term in the term knowledge base; represents a weight coefficient, which is used to adjust the contribution ratio of each scoring item in the comprehensive score; represents a small constant greater than zero, which is used to avoid calculation anomalies when taking the logarithm of the scoring item.

[0063] In one embodiment, the technical solution of step S304 is elaborated: the system comprehensively calculates the pronunciation association matching result of step S302, the language recognition result of step S303, and the context matching score of step S303 to obtain the comprehensive scores of each known term in the preliminary matching list. Exemplarily, the comprehensive score can be calculated using the above formula, where the weight coefficient can be adjusted according to the actual application scenario and expected preferences. For example, in a scenario with high requirements for pronunciation accuracy, increase the weight, and in a scenario with strong context association, increase the weight. The system selects the known term with the highest score as the most likely final candidate word according to the calculated comprehensive score. However, only when the highest score exceeds a preset threshold is it considered that an effective cross-language term match has been found. Once the final candidate word is determined, the system retrieves the detailed annotation information associated with it from the term knowledge base and the pronunciation model library. Exemplarily, it includes its translation in the target language, complete definition, usage examples, relevant domain knowledge, or any other auxiliary information that helps understanding and usage.

[0064] The weight coefficient is determined according to a preset strategy. The strategy is used to balance the relative importance of pronunciation similarity, context matching, and language recognition in the comprehensive evaluation to optimize cross-language

[0065] Such as Figure 5 shown is an exemplary step flowchart for integrating text results in step S4 of this embodiment, including:

[0066] Step S401: According to the characteristics or comprehensive scores of the final candidate words obtained in step S3, determine the annotation level or display priority of the final candidate words; and integrate the final candidate words and their annotation information into the positions where the corresponding words appear in the calibration note text in a preset integration format corresponding to the determined annotation level or display priority. The preset integration format includes at least inserting specific marks in the calibration note text, changing the word style, adding hover tips, generating sidebar notes, creating expandable details.

[0067] In one embodiment, the technical solution of step S401 is elaborated: Based on certain characteristics of the last candidate word determined in step S304 or its comprehensive score, the system determines the annotation level or display priority of the term in the final note. Exemplarily, terms with high scores and strong importance may be given a "first-level annotation", while terms with lower scores or secondary importance are given a "second-level annotation". Then, the system integrates the last candidate word and its annotation information obtained from the knowledge base into the position where the word actually appears in the calibrated note text according to a preset integration format corresponding to the determined annotation level / priority. The preset integration format may at least include: inserting specific markers in the text; changing the visual style of the word to attract the user's attention; adding a hover tip to quickly display its brief information when the user hovers the mouse over the annotated word; generating an independent sidebar annotation area to centrally display the detailed annotation information of the term; or creating a clickable link or icon to allow the user to expand and view the detailed information of the term.

[0068] Step S402, output the finally generated automatic note text with cross-language term annotations. The automatic note text includes the calibrated note text and the structured annotation data associated with the last candidate word to support subsequent reference, editing, or interactive display, and the output format of the automatic note text at least includes rich text format, HTML format, and Markdown format.

[0069] In one embodiment, the technical solution of step S402 is elaborated: The system finally generates and outputs an automatic note text with cross-language term annotations. This text not only includes the content of the calibrated note text generated in step S202, but also includes the structured annotation data associated with the last candidate word determined in step S304. Exemplarily, this structured annotation data can be stored inline with the text content or associated with the text file as an independent metadata file. This structured annotation is designed to support the subsequent reference, editing, and interactive display functions of the note. The output format of the automatic note text may at least include: rich text format, which supports retaining the text style and basic structure; HTML format, which is convenient for display in a web browser and supports complex interactive functions; Markdown format, which is convenient for lightweight editing and supports embedding annotations or links through specific syntax; or other formats that support embedding structured data.

[0070] The above content is only an example and illustration of the present invention. Those skilled in the art of this technology can make various modifications or supplements to the described specific embodiments or use similar methods to replace them. As long as they do not deviate from the scope defined by the invention, they should fall within the protection scope of the present invention.

Claims

1. An artificial intelligence-assisted cross-language automatic note-taking and term annotation system, characterized in that: Including an audio processing and cross - language speech recognition module, which is used to receive audio data, perform pre - processing, and perform cross - language speech recognition on the pre - processed audio data to output preliminary text; a text proofreading and term candidate recognition module, which is used to receive the preliminary text, perform reciprocal matching and proofreading to output calibrated note text, and recognize term candidate words from the preliminary text or the calibrated note text; A cross - language term matching analysis and annotation decision module, which is used to receive the term candidate words and the calibrated note text, obtain the pronunciation representation of the term candidate words, compare them with the cross - language pronunciation representations of known terms in the term knowledge base and the pronunciation model base, perform pronunciation related - word center matching to generate a preliminary cross - language term matching list, perform language recognition on the term candidate words or their contexts, perform context analysis on the compliance of each known term in the preliminary cross - language term matching list with the current text context where the term candidate words are located, calculate the context matching score, calculate the comprehensive score by integrating the pronunciation association matching result, language recognition result and context matching score, select the final candidate word that best fits the current context, and obtain the annotation information of the final candidate word from the term knowledge base and the pronunciation model base; An automatic note generation and annotation output module, which is used to receive the calibrated note text and the final candidate word and its annotation information, integrate the final candidate word and its annotation information into the calibrated note text to generate an automatic note with cross - language term annotation; A term knowledge base and a pronunciation model base, which are used to store terms, meanings and corresponding cross - language pronunciation representations in different languages, and provide query support for the cross - language term matching analysis and annotation decision module.

2. An artificial intelligence-assisted cross-language automatic note-taking and term annotation system according to claim 1, characterized in that: The automatic note and term annotation method includes the following steps: Step S1, obtain and process audio data, and perform cross - language speech recognition to obtain preliminary text; Step S2, perform reciprocal matching and proofreading on the preliminary text to generate calibrated note text, and recognize term candidate words; Step S3, when term candidate words are recognized, perform cross - language feature judgment, pronunciation association matching and context analysis, and comprehensively score to determine the cross - language terms to be annotated and their information; Step S4, integrate the final candidate word and its annotation information into the note text, and output the note text with annotation.

3. An artificial intelligence-assisted cross-language automatic note-taking and term annotation system according to claim 2, characterized in that: The said Step S1 includes, Step S101, through an audio input device, obtain a multi - channel or single - channel audio data stream to be processed, where the audio data stream contains the voices of different speakers, mixed languages and environmental noises; Step S102, perform enhanced pre - processing and acoustic feature extraction on the audio data, and the pre - processing at least includes multi - channel signal processing, voice activity detection, noise reduction, acoustic environment adaptation, gain adjustment, frame segmentation, and the acoustic feature extraction is used to obtain a feature vector sequence suitable for speech recognition; Step S103: Use a cross - language speech recognition model based on a deep - learning model to perform speech - to - text operation on the feature vector sequence. The cross - language speech recognition model can process speech inputs containing mixed languages and output confidence information of the speech recognition result and / or language recognition confidence information of text segments during the recognition process; Output a preliminary text stream or text segment.

4. An artificial intelligence-assisted cross-language automatic note-taking and terminology annotation system according to claim 2, characterized in that: Step S2 includes: Step S201: Block - process the preliminary text output in Step S1 according to a preset time length or semantic boundary to form text segments to be proofread and term - recognized; Step S202: Adopt a context - dependent reciprocating matching mode, use the speech recognition result corresponding to the received subsequent audio and / or subsequent text context information to perform re - scoring, re - searching or language - model - based post - processing operations on the currently and previously transcribed text segments, and dynamically correct the text segments in combination with the confidence information of the speech recognition result to generate a more accurate calibrated note text; Step S203: Scan the preliminary text or the calibrated note text to identify words or multi - word phrases that meet the preset features as the term candidate words.

5. An artificial intelligence-assisted cross-language automatic note-taking and term annotation system according to claim 2, characterized in that: Step S3 includes: Step S301: Extract the current text context where the term candidate words are located; Step S302: Obtain the pronunciation representation of the term candidate words, compare it with the pronunciation representations of known cross - language terms in the term knowledge base and pronunciation model library, perform pronunciation - related word - center matching, and generate a preliminary cross - language term matching list; Step S303: Based on the language recognition result of the term candidate words or their context, and the pronunciation - related word - center matching result, preliminarily judge the cross - language features of the term candidate words, and analyze the degree of compliance of each known term in the preliminary cross - language term matching list with the context, and calculate the context matching score; Step S304: Synthesize the pronunciation - related word - center matching result, the language recognition result and the context matching score, calculate the comprehensive score of each known term in the preliminary cross - language term matching list, and select the known term with the highest score and higher than the preset threshold as the final candidate word according to the comprehensive score. At the same time, obtain the annotation information of the final candidate word from the term knowledge base.

6. An artificial intelligence-assisted cross-language automatic note-taking and term annotation system according to claim 2, characterized in that: Step S4 includes: Step S401: Determine the annotation level or display priority of the final candidate word according to the characteristics or the comprehensive score of the final candidate word obtained in Step S3; And integrate the final candidate word and its annotation information into the position where the corresponding word appears in the calibrated note text in a preset integration format corresponding to the determined annotation level or display priority. The preset integration format at least includes inserting a specific mark in the calibrated note text, changing the word style, adding a hover prompt, generating a sidebar note, creating an expandable detail. Step S402: Output the finally generated automatic note text with cross - language term annotations. The automatic note text includes the calibrated note text and the structured annotation data associated with the last candidate word to support subsequent reference, editing, or interactive display. And the output format of the automatic note text includes at least Rich Text Format, HTML Format, and Markdown Format.

7. An artificial intelligence-assisted cross-language automatic note-taking and term annotation system according to claim 5, characterized in that: In step S304, calculate the comprehensive score of each known term in the preliminary cross - language term matching list using the following formula: , where represents the candidate term, represents the current text context, represents the th known term in the preliminary cross - language term matching list, and the known term is from the term knowledge base; represents the pronunciation similarity score between the pronunciation representation of the candidate term and the cross - language pronunciation representation of the known term . The pronunciation representation is obtained by performing acoustic feature extraction, phoneme sequence prediction, or generating a pronunciation vector on the speech of the candidate term. The cross - language pronunciation representation is the standard pronunciation representation of the known term in different languages. The pronunciation similarity score is obtained by calculating the distance or similarity measure between the pronunciation representations, and this score is greater than 0; represents the context matching score between the known term and the current text context . The context matching score is calculated by analyzing the semantic relevance between the known term and the context . The semantic relevance analysis is implemented based on word vectors, language models, topic models, and text matching algorithms, and this score value is greater than 0; represents the language recognition confidence score of the system's judgment that the candidate term belongs to the main language of the known term . The confidence score is obtained by performing language recognition analysis on the candidate term itself or its context . This score value is greater than 0 and less than or equal to 1; represents the preset main language label of the known term in the term knowledge base; represents a weight coefficient used to adjust the contribution ratio of each scoring item in the comprehensive score; represents a small constant greater than zero, used to avoid calculation anomalies when taking the logarithm of the scoring items.

8. An artificial intelligence-assisted cross-language automatic note-taking and term annotation system according to claim 7, characterized in that: The weight coefficient is determined according to a preset strategy, which is used to balance the relative importance of the pronunciation similarity, context matching, and the language recognition in the comprehensive evaluation to optimize the recognition accuracy of cross-language terms.

Citation Information

Patent Citations

  • Voice recognition method and device, electronic equipment and computer readable storage medium

    CN110517693A

  • Intelligent voice automatic translation system based on AI recognition

    CN119964573A