Multi-language character transcription method and system

Through collaborative processing between the AR device and the edge device, multimodal data is collected and processed, the problem of insufficient computing resources on the AR device is solved, efficient and real-time multilingual text transcription is achieved, and user experience is improved.

CN120029453APending Publication Date: 2025-05-23GUDONG TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510036060.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

Due to limited computing resources and storage space on the AR device, it is difficult to achieve high-quality and real-time multilingual text transcription, especially in scenarios such as complex speech environments, multilingual mixing and long-term continuous transcription.

Method used

Through the interaction between the AR device and the edge device, multimodal data to be transcribed including voice, visual and action data is collected, pre-processed and feature extraction is performed, and preliminary language recognition is used using the computing resources of the AR device, and then the data is sent to the edge device for more complex text transcription model processing.

Benefits of technology

It improves the accuracy and response speed of text transcription on the AR device side, achieves efficient and real-time multilingual text transcription, and improves user interaction experience and task efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029453A_ABST
    Figure CN120029453A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a multi-language text transcription method and system, and is applied to an AR device end, the method comprises the following steps: collecting first to-be-transcribed data of a user through the AR device end, and preprocessing the first to-be-transcribed data to obtain preprocessed to-be-transcribed data, the first to-be-transcribed data comprising voice data, visual data and action data; performing feature extraction on the first to-be-transcribed data to obtain first to-be-transcribed feature data, the first to-be-transcribed feature data including first voice feature data, first visual feature data and first action feature data; inputting the first voice feature data into a preset language recognition model, outputting preliminary target language data, and sending the first to-be-transcribed feature data and the preliminary target language data to an edge device end; and receiving the text transcription result, and displaying the text transcription result on the AR equipment end, so that the problem of insufficient computing resources of the AR equipment end can be effectively solved, and the transcription accuracy and reaction speed of the AR equipment end are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of Internet of Things, and in particular to a multi-language text transcription method and system. Background Art

[0002] With the rapid development of augmented reality (AR) technology, AR devices have been widely used in various fields. However, due to the limited computing resources and storage space of AR devices themselves, there are many challenges in achieving high-quality, real-time multilingual text transcription.

[0003] Currently, existing text transcription technologies mainly rely on deploying complex deep learning models on AR devices, such as end-to-end speech recognition models or sequence-to-sequence text transcription models. These models usually have a large number of parameters and complex network structures, and consume a lot of computing resources and storage space. However, due to the hardware conditions of AR devices, such as limited processor performance, memory size, and battery capacity, directly deploying and running these highly complex models on AR devices faces the problem of insufficient resources.

[0004] Due to the limitation of computing resources, it is difficult for existing AR devices to achieve efficient and real-time multilingual text transcription. When faced with complex voice environments, multilingual mixing, long-term continuous transcription and other scenarios, the transcription model on the AR device often cannot achieve the ideal transcription accuracy and speed. This limits the text transcription capabilities of the AR device in actual applications, affecting the user's interactive experience and task efficiency. Summary of the invention

[0005] The present application provides a multilingual text transcription method and system, which can effectively solve the problem of insufficient computing resources on the AR device side and improve the accuracy and response speed of transcription on the AR device side through the interaction between the AR device side and the edge device side.

[0006] In a first aspect, the present application provides a multilingual text transcription method, which is applied to an AR device, and the method includes: When the AR device is in an online state, first data to be transcribed of the user is collected through the AR device, and the first data to be transcribed is preprocessed to obtain preprocessed data to be transcribed, where the first data to be transcribed includes voice data, visual data, and action data; Performing feature extraction on the first data to be transcribed to obtain first feature data to be transcribed, wherein the first feature data to be transcribed includes first voice feature data, first visual feature data, and first action feature data; Inputting the first speech feature data into a preset language recognition model, outputting preliminary target language data, and sending the first feature data to be transcribed and the preliminary target language data to an edge device; Receive the text transcription result and display it on the AR device.

[0007] By adopting the above technical solution, multimodal data to be transcribed, including voice, visual and motion data, is collected through the AR device, and preprocessed and feature extracted to obtain first voice feature data, first visual feature data and first motion feature data. This multimodal data collection and feature extraction method fully utilizes the perception capabilities of the AR device and obtains rich and comprehensive user interaction information. By preprocessing and feature extraction of data in different modalities, data redundancy is effectively reduced, data quality is improved, and a good foundation is laid for subsequent language recognition and text transcription.

[0008] By inputting the first voice feature data into the preset language recognition model, the language of the user's input voice can be quickly and accurately determined to obtain preliminary target language data. This step utilizes the computing resources of the AR device to complete the preliminary judgment of language recognition locally, reducing the computing burden on the edge device. At the same time, by sending the first feature data to be transcribed and the preliminary target language data to the edge device together, collaborative processing between the AR device and the edge device is achieved, giving full play to the computing advantages of both.

[0009] After the edge device receives the first feature data to be transcribed and the preliminary target language data sent by the AR device, it can use its powerful computing resources to run a more accurate and complex text transcription model. By comprehensively utilizing speech, vision, and motion feature data, combined with preliminary target language information, the edge device can accurately transcribe the input multimodal data into corresponding text results. This collaborative processing method not only improves the accuracy of text transcription, but also makes full use of the computing resources of the AR device and the edge device to achieve efficient and real-time transcription tasks.

[0010] Finally, the text transcription results generated by the edge device are sent back to the AR device and displayed, so that the user can intuitively see the transcribed text content. This real-time feedback mechanism greatly improves the user's interactive experience and task efficiency. The user can view the transcription results instantly on the AR device and perform subsequent editing, processing or sharing operations as needed.

[0011] Optionally, the method further includes: When the AR device is in an offline state, the AR device collects second data to be transcribed of the user, and pre-processes the first data to be transcribed to obtain pre-processed second data to be transcribed, where the second data to be transcribed includes second voice data; Preprocessing the second voice data to obtain preprocessed second voice data, and extracting features from the preprocessed second voice data to obtain second voice feature data; The second speech feature data belongs to a preset language recognition model to obtain preliminary target language data, and according to the preliminary target language data, the corresponding first text transcription model is loaded; The second voice feature data is input into the corresponding first text transcription model, a text transcription result is output, and the text transcription result is displayed on the AR device.

[0012] By adopting the above technical solution, when the AR device is offline, it is unable to interact with the edge device for data and offload computing tasks, so it is necessary to rely on the local computing resources of the device to complete the text transcription task.

[0013] In an offline state, the AR device collects the user's second data to be transcribed, mainly including the second voice data, through its own sensors. The second voice data is preprocessed to remove noise interference and improve voice quality, thereby obtaining the preprocessed second voice data. Then, feature extraction is performed on the preprocessed second voice data to obtain second voice feature data. This process makes full use of the computing resources of the AR device, completes the preprocessing and feature extraction of voice data locally, and reduces the delay in data transmission and processing.

[0014] In order to realize language recognition in offline state, the AR device inputs the second voice feature data into the preset language recognition model to obtain preliminary target language data. The language recognition model is pre-trained and can quickly and accurately determine the language of the user's input voice locally. According to the preliminary target language data, the AR device loads the corresponding first text transcription model from local storage. This targeted model loading method avoids loading transcription models of multiple languages, reduces the waste of computing resources, and improves transcription efficiency.

[0015] By inputting the second voice feature data into the loaded first text transcription model, the text transcription task can be completed locally on the AR device and the corresponding text transcription results can be output. The first text transcription model is optimized for a specific language and can provide high transcription accuracy in an offline state. Finally, the text transcription results are displayed on the AR device, and users can view the transcribed text content in real time for a smooth interactive experience.

[0016] In a second aspect, the present application also provides a multilingual text transcription method, which is applied to an edge device, and the method includes: Receiving preliminary target language data and multiple first feature data to be transcribed sent by the AR device; Extracting acoustic features from the first speech feature data through an acoustic model to obtain acoustic features, and inputting the acoustic features and preliminary target language data into a preset target language classification model to output final target language data; Performing feature fusion on the first voice feature data, the first visual feature data, and the first action feature data to obtain multimodal feature data; Loading the corresponding second transcription model based on the final target language data, and inputting the multimodal feature data into the second transcription model to obtain a text transcription result; Send the text transcription results to the AR device.

[0017] By adopting the above technical solution, in the process of edge device segment processing, the computing resources and data processing capabilities of the edge device are fully utilized to achieve efficient and accurate text transcription. The edge device first receives the preliminary target language data and the first feature data to be transcribed sent by the AR device, including the first voice feature data, the first visual feature data and the first action feature data. This data transmission method realizes collaborative processing between devices, and transmits the multimodal data and preliminary language recognition results collected by the AR device to the edge device, making full use of the computing power of the edge device.

[0018] The acoustic features of the first speech feature data are extracted through the acoustic model to obtain acoustic features, which further enhances the representation ability of speech data and extracts more refined and robust speech features. The extracted acoustic features and preliminary target language data are input into the preset target language classification model to obtain the final target language data. The introduction of a professional and accurate language classification model verifies and corrects the preliminary language recognition results, improves the accuracy of language recognition, and lays a good foundation for subsequent text transcription.

[0019] In order to make full use of the complementarity and relevance of multimodal data, the edge device fuses the first voice feature data, the first visual feature data, and the first action feature data to obtain multimodal feature data. Feature fusion takes into account the association and complementary information between different modal data, and realizes cross-modal information interaction and enhancement by mapping them to a common feature space. Multimodal feature fusion can more comprehensively and accurately represent the user's input intention and improve the accuracy of text transcription.

[0020] According to the final target language data, the corresponding second transcription model is loaded on the edge device. The second transcription model is optimized for a specific language and can better capture the language characteristics and transcription rules of the language. The multimodal feature data is input into the loaded second transcription model to obtain the final text transcription result. The targeted model loading and transcription process fully considers the particularities of different languages ​​and improves the accuracy and efficiency of transcription.

[0021] Finally, the edge device sends the text transcription results to the AR device, realizing real-time transmission and display of the transcription results. Users can view the transcribed text content in real time through the AR device, and get a smooth and natural interactive experience.

[0022] In summary, this request-response mechanism improves the flexibility and modularity of the management system and can meet the needs of complex edge computing environments.

[0023] In a third aspect of the present application, a multilingual text transcription system is provided, which is applied to an AR device, and the system includes: A first data acquisition module 1 is used for collecting first data to be transcribed of the user through the AR device end when the AR device end is in an online state, and preprocessing the first data to be transcribed to obtain the preprocessed data to be transcribed, where the first data to be transcribed includes voice data, visual data and action data; A first feature extraction module 2, used for extracting features from the first data to be transcribed to obtain first feature data to be transcribed, where the first feature data to be transcribed includes first speech feature data, first visual feature data and first action feature data; A first preliminary target language determination module 3, used to input the first speech feature data into a preset language recognition model, output preliminary target language data, and send the first feature data to be transcribed and the preliminary target language data to the edge device end; The text transcription result receiving module 4 is used to receive the text transcription result and display the text transcription result on the AR device.

[0024] In a fourth aspect of the present application, a multilingual text transcription system is provided, which is applied to an edge device, and the system includes: The data receiving module 5 is used to receive the preliminary target language data and the first feature data to be transcribed sent by the AR device end; The final target language determination module 6 is used to extract acoustic features from the first speech feature data through an acoustic model to obtain acoustic features, and input the acoustic features and preliminary target language data into a preset target language classification model to output final target language data; A feature fusion module 7, used for fusing the first speech feature data, the first visual feature data and the first action feature data to obtain multimodal feature data; A first transcription module 8, used to load a corresponding second transcription model based on the final target language data, and input the multimodal feature data into the second transcription model to obtain a text transcription result; The text transcription result sending module 9 is used to send the text transcription result to the AR device end.

[0025] In a fifth aspect of the present application, a computer-readable storage medium is provided, wherein the computer storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor and executing the above-mentioned method steps.

[0026] In summary, one or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages: 1. This application collects multimodal data to be transcribed, including voice, visual and motion data, through the AR device, and performs preprocessing and feature extraction to obtain first voice feature data, first visual feature data and first motion feature data. This multimodal data collection and feature extraction method fully utilizes the perception capabilities of the AR device and obtains rich and comprehensive user interaction information. By preprocessing and feature extraction of data in different modalities, data redundancy is effectively reduced, data quality is improved, and a good foundation is laid for subsequent language recognition and text transcription.

[0027] By inputting the first voice feature data into the preset language recognition model, the language of the user's input voice can be quickly and accurately determined to obtain preliminary target language data. This step utilizes the computing resources of the AR device to complete the preliminary judgment of language recognition locally, reducing the computing burden on the edge device. At the same time, by sending the first feature data to be transcribed and the preliminary target language data to the edge device together, collaborative processing between the AR device and the edge device is achieved, giving full play to the computing advantages of both.

[0028] 2. In the offline state, the AR device collects the user's second data to be transcribed, mainly including the second voice data, through its own sensor. The second voice data is preprocessed to remove noise interference and improve voice quality to obtain the preprocessed second voice data. Then, feature extraction is performed on the preprocessed second voice data to obtain second voice feature data. This process makes full use of the computing resources of the AR device, completes the preprocessing and feature extraction of voice data locally, and reduces the delay of data transmission and processing.

[0029] 3. In the process of edge device segment processing, the present application fully utilizes the computing resources and data processing capabilities of the edge device side to achieve efficient and accurate text transcription. The edge device side first receives the preliminary target language data and the first feature data to be transcribed, including the first voice feature data, the first visual feature data, and the first action feature data, sent by the AR device side. This data transmission method realizes collaborative processing between devices, and transmits the multimodal data and preliminary language recognition results collected by the AR device side to the edge device side, making full use of the computing power of the edge device side.

[0030] The acoustic features of the first speech feature data are extracted through the acoustic model to obtain acoustic features, which further enhances the representation ability of speech data and extracts more refined and robust speech features. The extracted acoustic features and preliminary target language data are input into the preset target language classification model to obtain the final target language data. The introduction of a professional and accurate language classification model verifies and corrects the preliminary language recognition results, improves the accuracy of language recognition, and lays a good foundation for subsequent text transcription. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 This is a flowchart of a multilingual text transcription method applied to an AR device end provided by an embodiment of the present application; Figure 2 This is a flowchart of a multilingual text transcription method applied to an edge device provided by an embodiment of the present application; Figure 3 A system architecture diagram for an AR device provided in an embodiment of the present application; Figure 4 A diagram of the system architecture provided for an edge device in an embodiment of the present application. DETAILED DESCRIPTION

[0032] In order to enable technicians in this field to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this application, not all of the embodiments.

[0033] In the description of the embodiments of the present application, words such as "for example" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "for example" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "for example" or "for example" is intended to present related concepts in a specific way.

[0034] In the description of the embodiments of the present application, the meaning of the term "multiple" refers to two or more. For example, multiple systems refer to two or more systems, and multiple screen terminals refer to two or more screen terminals. In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance or implicitly indicating the indicated technical features. Thus, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. The terms "include", "comprise", "have" and their variations all mean "including but not limited to", unless otherwise specifically emphasized.

[0035] The following will provide a clear and complete description of the technical solutions in the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all of the embodiments.

[0036] On this basis, the present application embodiment provides a multilingual text transcription method and system. Figure 1 This is a flowchart of a multilingual text transcription method applied to an AR device end provided by an embodiment of the present application. The method can be implemented by a computer program or run as an independent tool application. Specifically, the method is applied to a main probe on an edge device end. The method includes: S101, when the AR device is in an online state, collecting first data to be transcribed of the user through the AR device, and preprocessing the first data to be transcribed to obtain preprocessed data to be transcribed, where the first data to be transcribed includes voice data, visual data, and action data; Specifically, when the AR device is online, the user's first data to be transcribed is collected through the AR device. The AR device uses its built-in microphone, camera, and motion sensor to simultaneously collect the user's voice data, visual data, and motion data to form the first data to be transcribed. The purpose of collecting multimodal data is to fully capture the user's interactive information and provide rich input information for subsequent text transcription. Voice data includes the user's voice commands and spoken content, visual data includes the user's facial expressions, lip shape changes, and gestures, and motion data includes the user's head movements and body postures. By collecting data from these different modalities, the AR device can more accurately understand the user's intentions and semantics, and improve the accuracy of text transcription.

[0037] In order to improve data quality and processing efficiency, the AR device preprocesses the collected first data to be transcribed. The preprocessing process includes steps such as data cleaning, noise removal, and feature enhancement. For voice data, the voice activity detection (VAD) algorithm is used to remove silent segments, and the speech enhancement algorithm is used to suppress background noise to improve the signal-to-noise ratio and clarity of the voice. For visual data, the image denoising algorithm is used to eliminate noise and blur in the image, and the image enhancement algorithm is used to improve the contrast and sharpness of the image. For motion data, the smoothing filtering algorithm is used to eliminate jitter and mutation in the motion data, and the data interpolation algorithm is used to supplement the missing data points. After preprocessing, the quality of the first data to be transcribed is significantly improved, laying a good foundation for subsequent feature extraction and text transcription.

[0038] S102, extracting features from the first data to be transcribed to obtain first feature data to be transcribed, where the first feature data to be transcribed includes first voice feature data, first visual feature data, and first action feature data; Specifically, for speech data, commonly used speech feature extraction methods such as Mel frequency cepstral coefficients (MFCC) and linear prediction cepstral coefficients (LPCC) are used to extract the first speech feature data reflecting the spectral characteristics and formant characteristics of the speech signal. First, the pre-processed speech data is frame segmented and windowed, and then the time domain signal is converted into a frequency domain signal by fast Fourier transform (FFT). In the frequency domain, according to the auditory characteristics of the human ear, a Mel filter group is applied to the spectrum to obtain a Mel spectrum. Finally, the Mel spectrum is subjected to discrete cosine transform (DCT) to obtain an MFCC feature vector. These first speech feature data can effectively characterize key features such as the timbre, pitch, and formant of the speech, and provide important discriminant information for language recognition and text transcription.

[0039] For visual data, deep learning models such as convolutional neural networks (CNN) are used to extract the first visual feature data reflecting facial expressions, changes in mouth shape, and gestures. First, face detection and alignment are performed on the preprocessed visual data to extract the region of interest (ROI). Then, the ROI is input into a pre-trained CNN model, such as VGGNet, ResNet, etc., and features are extracted layer by layer through convolutional layers and pooling layers. In the last layer or the last few layers of CNN, high-level semantic features are obtained as the first visual feature data. These features can capture subtle changes in facial expressions, dynamic changes in mouth shape, and the spatial trajectory of gestures, providing rich visual clues for text transcription.

[0040] For motion data, sequence models such as recursive neural networks (RNN) and long short-term memory networks (LSTM) are used to extract the first motion feature data reflecting the user's head movement and body posture. First, the preprocessed motion data is sequenced and normalized. Then, the motion sequence is input into the RNN or LSTM model, and the time dependency and long-term and short-term context information of the motion data are captured through the recurrent connection and gating mechanism. In the output layer of the model, the hidden state vector representing the motion feature is obtained as the first motion feature data. These features can characterize the user's head movement, body posture and other motion information, and provide additional auxiliary information for text transcription.

[0041] Based on the above embodiment, as an optional embodiment, feature extraction is performed on the first data to be transcribed to obtain first feature data to be transcribed, where the first feature data to be transcribed includes first voice feature data, first visual feature data, and first action feature data, including: S201, extracting features from speech data using a GMM-HMM model to obtain first speech feature data; Specifically, the edge device first preprocesses the received voice data, such as voice activity detection, endpoint detection, noise reduction, etc., to obtain a pure voice signal. Then, the preprocessed voice signal is frame segmented and windowed to divide the continuous voice signal into several short time frames. The length of each frame is usually 20-30 milliseconds, and the frame shift is about 10 milliseconds. Fast Fourier transform (FFT) is performed on each frame of voice signal to obtain its frequency domain representation. Next, the GMM-HMM model is used to model and extract the frequency domain features of each frame of voice signal. The GMM-HMM model is a speech recognition model that combines the Gaussian mixture model (GMM) and the hidden Markov model (HMM). Among them, GMM is used to probabilistically model the frequency domain features of each frame of voice signal and characterize its statistical distribution characteristics. HMM is used to model the temporal relationship of voice signals and capture the transfer rules between voice units.

[0042] In the GMM part, the frequency domain features of each frame of speech signal are extracted using common speech features such as Mel-frequency cepstral coefficients (MFCC) or linear prediction cepstral coefficients (LPCC) to obtain a feature vector of fixed dimension. Then, a mixed Gaussian model of multiple Gaussian components is used to estimate the probability density of the feature vector to obtain the posterior probability that the frame of speech belongs to each Gaussian component. These posterior probabilities constitute the GMM observation probability distribution of the frame of speech.

[0043] In the HMM part, the continuous GMM observation probability distribution sequence is used as the HMM observation sequence, and the Baum-Welch algorithm is used to train and estimate the parameters of the HMM, including the state transition probability matrix and state emission probability distribution. Through HMM modeling, the state transition rules of speech signals at different times can be characterized and the contextual relationship between speech units can be captured.

[0044] After feature extraction and modeling of the GMM-HMM model, the first speech feature data is obtained, including the GMM observation probability distribution sequence and HMM state sequence of each frame of speech. These feature data not only characterize the frequency domain characteristics of the speech signal, but also consider the temporal relationship between speech units, and have strong discriminability and robustness. The first speech feature data will serve as the input for subsequent language recognition and text transcription, providing reliable feature representation for these tasks.

[0045] Using the GMM-HMM model for speech feature extraction can effectively model the acoustic characteristics and timing rules of speech signals, and obtain robust and effective first speech feature data. Compared with traditional speech feature extraction methods, the GMM-HMM model takes into account the statistical distribution characteristics and contextual relationships of speech signals, and can better characterize the essential attributes of speech and improve the discriminative ability of features. At the same time, the GMM-HMM model has a certain degree of adaptability and robustness to noise and speaker changes, and can reduce the impact of environmental factors and individual differences on feature extraction to a certain extent.

[0046] S202, extracting features from the visual data using an image feature extraction algorithm to obtain second visual feature data; Specifically, the edge device first preprocesses the received visual data, such as image scaling, normalization, and denoising, to obtain image data of uniform size and format. Then, different image feature extraction algorithms are used to extract features from the preprocessed image data to obtain multi-level and multi-faceted visual features.

[0047] Commonly used image feature extraction algorithms include local binary pattern (LBP), scale-invariant feature transform (SIFT), histogram of oriented gradients (HOG), convolutional neural network (CNN), etc. Among them, the LBP algorithm extracts local texture features of the image by comparing the size relationship between pixels and their neighboring pixels. The SIFT algorithm extracts local structural features of the image by detecting key points in the image and calculating the gradient direction and amplitude around the key points. The HOG algorithm extracts the shape and contour features of the image by calculating the gradient direction histogram of the local area of ​​the image. The CNN algorithm automatically learns and extracts the hierarchical feature representation of the image through multi-layer convolution and pooling operations.

[0048] In practical applications, edge devices can select one or more image feature extraction algorithms based on specific task requirements and computing resources to extract visual features of different levels and types. For example, the LBP algorithm can be used to extract the texture features of the image, the SIFT algorithm can be used to extract the key point features of the image, and the HOG algorithm can be used to extract the shape features of the image. These features can be combined into a high-dimensional feature vector as the second visual feature data.

[0049] Different feature extraction strategies can be used for different types of visual data, such as text images, scene images, and face images. For example, for text images, the focus can be on extracting the shape and arrangement features of characters; for scene images, the focus can be on extracting the category and spatial relationship features of objects; for face images, the focus can be on extracting the geometry and expression features of facial organs. By selecting and combining image feature extraction algorithms in a targeted manner, more effective and comprehensive second visual feature data can be obtained.

[0050] The extracted second visual feature data will be used together with the first voice feature data as input for subsequent language recognition and text transcription. By integrating the information of the two modalities of voice and vision, they can complement and verify each other to improve the robustness and accuracy of the system. For example, visual features can provide contextual information in speech recognition and help eliminate ambiguity and uncertainty in speech signals; speech features can also provide semantic information for visual understanding and help determine key content and attributes in images.

[0051] S203, preprocessing the motion data to obtain preprocessed motion data, and extracting features from the preprocessed motion data to obtain first motion feature data, wherein the first motion feature data includes time domain statistical features, time-frequency domain features, and time series features.

[0052] Specifically, the edge device first preprocesses the received action data. The preprocessing steps include data smoothing, normalization, segmentation and other operations. Through data smoothing, such as moving average filtering, median filtering, etc., high-frequency noise and mutation points in the action signal can be removed to obtain more stable and continuous data. Through data normalization, action data of different dimensions and dimensions can be mapped to a unified scale range, which is convenient for subsequent feature extraction and comparison. Through data segmentation, continuous action data can be divided into several action segments, each segment corresponds to a complete action cycle or action unit, providing a more fine-grained processing unit for feature extraction.

[0053] After obtaining the preprocessed action data, the edge device extracts its features to obtain the first action feature data, including time-domain statistical features, time-frequency domain features, and time series features. Among them, the time-domain statistical features characterize the statistical properties of the action signal in the time dimension, such as mean, variance, kurtosis, skewness, etc., reflecting the overall intensity and change trend of the action. The time-frequency domain features characterize the joint distribution of the action signal in the two dimensions of time and frequency, such as short-time Fourier transform (STFT), wavelet transform, etc., reflecting the frequency component and energy distribution of the action. The time series features characterize the evolution law of the action signal in the time dimension, such as hidden Markov model (HMM), recurrent neural network (RNN), etc., reflecting the temporal relationship and contextual dependence of the action.

[0054] When extracting time-domain statistical features, the edge device calculates the mean, variance, kurtosis, skewness and other statistics for each action segment to obtain a fixed-dimensional feature vector. These statistics can reflect the overall intensity, range of variation, distribution shape and other properties of the action segment, providing discriminant information for the recognition and comparison of action patterns.

[0055] When extracting time-frequency domain features, the edge device performs STFT or wavelet transform on each action clip to obtain its time-frequency representation. Then, feature extraction is performed on the time-frequency representation, such as calculating frequency band energy, frequency band entropy, frequency band center, etc., to obtain a feature vector of fixed dimension. These time-frequency domain features can reflect the energy distribution and change rules of the action clips in different frequency components, and provide frequency domain information for the description of action rhythm and intensity.

[0056] When extracting time series features, the edge device inputs the continuous action fragment sequence into a time series model such as HMM or RNN, and obtains the hidden state sequence or output representation of the action sequence through model learning and inference. These time series features can characterize the evolution law and context dependency of actions in the time dimension, capture the association and transfer relationship between actions, and provide time series information for understanding the semantics of actions.

[0057] By preprocessing and extracting features from the motion data, first motion feature data including time-domain statistical features, time-frequency domain features, and time series features are obtained. These features describe the properties and laws of motion signals from different angles and granularities, and provide rich and complementary motion representations. The first motion feature data, together with the first speech feature data and the second visual feature data, are used as inputs for subsequent language recognition and text transcription, which can make full use of multimodal information and improve the accuracy and adaptability of the system.

[0058] Motion features can provide important auxiliary information for speech recognition, such as the speaker's identity, emotion, tone, etc., to help eliminate the ambiguity and uncertainty of speech signals. At the same time, motion features can also provide contextual clues for visual understanding, such as character behavior, interaction, scene, etc., to help determine the key content and events in the image. By integrating the features of the three modalities of speech, vision and motion, the target language and text content can be inferred and verified from multiple angles, improving the robustness and reliability of the entire multimodal text transcription system.

[0059] S103, inputting the first speech feature data into a preset language recognition model, outputting preliminary target language data, and sending the first feature data to be transcribed and the preliminary target language data to an edge device; Specifically, the extracted first speech feature data is input into a preset language recognition model, and the model extracts language-related high-level features layer by layer through forward propagation and feature transformation. In the output layer of the model, the probability distribution of each language category is calculated through the Softmax activation function, and the language with the highest probability is selected as the preliminary target language data. This process can be carried out in real time, and the language is dynamically recognized according to the user's voice input. The preliminary target language data represents the language category to which the model preliminarily determines the user's input voice belongs, such as Chinese, English, Japanese, etc.

[0060] After outputting the preliminary target language data, the AR device sends the first feature data to be transcribed and the preliminary target language data to the edge device. The reason for sending the data to the edge device is to utilize the more powerful computing power and storage resources of the edge device to perform more complex and sophisticated text transcription tasks. The edge device is usually deployed at the edge of the network close to the AR device, such as a local server, gateway, etc., with lower transmission delay and higher data throughput. By offloading computing tasks to the edge device, the processing burden on the AR device can be reduced, and the efficiency and performance of the entire system can be improved.

[0061] During the data transmission process, the AR device establishes a connection with the edge device through wireless communication technologies such as Wi-Fi, 5G, etc. Then, the AR device packages the first feature data to be transcribed (including the first voice feature data, the first visual feature data, and the first action feature data) and the preliminary target language data, and sends them to the edge device through the network transmission protocol. After receiving the data, the edge device unpacks and parses the data, extracts the feature data of each modality and the preliminary target language information, and prepares for subsequent processing.

[0062] S104, receiving the text transcription result, and displaying the text transcription result on the AR device.

[0063] Specifically, when the edge device completes the processing and text transcription of the first feature data to be transcribed, the generated text transcription result will be sent back to the AR device through the network transmission protocol. The AR device receives the data packet sent by the edge device through a wireless communication module, such as Wi-Fi, 5G, etc. After receiving the data packet, the AR device parses and extracts the data packet to obtain the text transcription result. The text transcription result is usually expressed in the form of a string, which contains the text content of the user's input voice.

[0064] After obtaining the text transcription results, the AR device displays them on the device's screen so that the user can view the transcribed text content in real time. The AR device can use a variety of display methods, such as overlay display, pop-up display, scrolling display, etc., according to the specific application scenario and user needs. For example, in the scenario of real-time voice interaction, the text transcription results can be superimposed in the user's field of view in the form of subtitles, and displayed in harmony with the surrounding environment. In the scenario of voice recording and note-taking, the text transcription results can be displayed in a fixed position on the screen in the form of a text box, which is convenient for users to view and edit.

[0065] When displaying the text transcription results, the AR device can also perform necessary formatting and optimization on the text content to improve the readability and aesthetics of the text. For example, punctuation marks, line breaks, etc. are automatically added according to the semantic and syntactic structure of the text content. For special texts such as recognized proper nouns, numbers, dates, etc., different colors, fonts or styles can be used to highlight them, making it easier for users to quickly locate and understand them. In addition, the AR device can also provide text editing functions, allowing users to modify, delete, copy, and other operations on the transcription results to meet the user's personalized needs.

[0066] Based on the above embodiment, as an optional embodiment, the method further includes: S105, when the AR device is in an offline state, collecting second data to be transcribed of the user through the AR device, and preprocessing the first data to be transcribed to obtain preprocessed second data to be transcribed, where the second data to be transcribed includes second voice data; Specifically, the AR device collects the user's voice input through the built-in microphone to obtain the second voice data. Unlike the online state, the second data to be transcribed in the offline state only includes the second voice data, and does not include visual data and action data. This is because in the offline state, the computing resources and battery capacity of the AR device are limited, and the overhead of processing multimodal data is large. Therefore, in order to ensure the real-time and smoothness of the transcription task, only voice data is collected and processed.

[0067] After collecting the second voice data, the AR device preprocesses it to obtain the preprocessed second data to be transcribed. The preprocessing steps are similar to those in the online state, including voice activity detection, voice enhancement, noise suppression, etc. Through voice activity detection, the silent segments in the second voice data are removed to reduce the data volume and processing overhead. Through voice enhancement and noise suppression, the signal-to-noise ratio and clarity of the voice are improved, providing higher quality input data for subsequent feature extraction and text transcription. The preprocessed second data to be transcribed will be used as input for subsequent steps, and feature extraction and text transcription will be performed locally on the AR device.

[0068] Preprocessing the second data to be transcribed offline can effectively reduce the amount of data, reduce computational complexity, and improve data quality. This is especially important for AR devices with limited resources, as it can extend battery life and device usage time while ensuring transcription performance. The preprocessed second data to be transcribed will be further processed and transcribed locally on the AR device without relying on network connection and edge device. This offline processing capability can expand the application scenarios of AR devices, and ensure that the user's speech-to-text interaction is not affected even when the network is unavailable or the edge device fails.

[0069] S106, preprocessing the second voice data to obtain preprocessed second voice data, and extracting features from the preprocessed second voice data to obtain second voice feature data; Specifically, the AR device first preprocesses the second voice data. The preprocessing steps include voice activity detection (VAD), voice enhancement, and noise suppression. Through VAD, the valid voice segments and silent segments in the voice data can be automatically detected, the silent segments can be removed, and the amount of data and processing overhead can be reduced. Through speech enhancement algorithms, such as spectral subtraction and Wiener filtering, the background noise in the voice data can be suppressed, and the signal-to-noise ratio and clarity of the voice can be improved. Through noise suppression algorithms, such as adaptive noise cancellation (ANC) and spectrogram masking, the residual noise in the voice data can be further removed, and the quality and intelligibility of the voice can be improved. After preprocessing, the preprocessed second voice data is obtained, and the data quality is significantly improved, providing a purer and more reliable input for subsequent feature extraction.

[0070] After obtaining the preprocessed second voice data, the AR device extracts features from it to obtain the second voice feature data. The feature extraction step is similar to that in the online state, and commonly used voice feature extraction methods such as Mel frequency cepstral coefficients (MFCC) and linear prediction cepstral coefficients (LPCC) are used. First, the preprocessed second voice data is frame segmented and windowed to divide the voice signal into several short time frames. Then, a fast Fourier transform (FFT) is performed on each frame to convert the time domain signal into a frequency domain signal. In the frequency domain, according to the auditory characteristics of the human ear, a Mel filter group is applied to the spectrum to obtain a Mel spectrum. Finally, a discrete cosine transform (DCT) is performed on the Mel spectrum to obtain the MFCC feature vector as the second voice feature data. These feature data can effectively characterize key features such as the timbre, pitch, and resonance peaks of the voice, and provide distinguishing and discriminative information for subsequent language recognition and text transcription.

[0071] By preprocessing and extracting features from the second speech data, more refined and effective second speech feature data can be obtained. The preprocessing step improves the quality of the speech data, removes noise and redundant information, and provides a purer and more reliable input for feature extraction. The feature extraction step extracts the most distinctive and representative features from the preprocessed speech data, such as MFCC feature vectors, which greatly reduces the data dimension and highlights the key features of the speech. These refined second speech feature data will serve as input for subsequent language recognition and text transcription, which can significantly improve the performance and efficiency of these tasks.

[0072] Based on the above embodiment, as an optional embodiment, preprocessing the second voice data to obtain the preprocessed second voice data includes: S301, performing noise reduction processing on the second voice data by spectral subtraction to obtain the second voice data after noise reduction processing, and performing voice endpoint detection on the second voice data after noise reduction processing to obtain voice endpoint data; Specifically, the AR device first performs spectral subtraction noise reduction processing on the received second voice data. Spectral subtraction is a frequency domain-based voice noise reduction algorithm that estimates the spectrum of the noise and subtracts it from the spectrum of the voice signal to obtain a pure voice spectrum. The specific steps are as follows: 1. Perform short-time Fourier transform (STFT) on the second speech data to obtain its spectrum representation.

[0073] 2. Estimate the noise spectrum in the non-speech segments of the speech signal, such as the silent segments before and after the speech. You can use energy-based or statistical model-based methods, such as minimum statistic estimation, Gaussian mixture model, etc., to get the estimated value of the noise spectrum.

[0074] 3. Subtract the estimated noise spectrum from the spectrum of the speech signal to obtain the speech spectrum after noise reduction. In order to avoid speech distortion caused by excessive noise reduction, a spectrum compensation factor can be introduced to correct the spectrum after noise reduction.

[0075] 4. Perform inverse Fourier transform on the denoised speech spectrum to obtain the second speech data after denoising in the time domain.

[0076] Through spectral subtraction noise reduction processing, the background noise in the second voice data, such as environmental noise, equipment noise, etc., can be effectively removed to obtain a purer and clearer voice signal. Compared with the traditional time domain noise reduction method, spectral subtraction directly estimates and eliminates noise in the frequency domain, has better noise reduction effect and spectral resolution, and can retain the details and characteristics of the voice signal.

[0077] After obtaining the second voice data after noise reduction processing, the AR device performs voice endpoint detection on it to obtain voice endpoint data. Voice endpoint detection refers to accurately detecting the starting and ending points of voice from continuous voice signals and dividing the time interval of voice activity. Commonly used voice endpoint detection algorithms include judgment methods based on features such as energy, zero-crossing rate, and spectral entropy, as well as classification methods based on machine learning models such as hidden Markov models (HMM) and support vector machines (SVM).

[0078] In practical applications, AR devices can use a combination of multiple features and algorithms to achieve robust and efficient speech endpoint detection. For example, time domain features such as short-time energy and zero-crossing rate can be used to roughly determine potential speech segments in speech signals. Then, these potential speech segments are further extracted and modeled, such as calculating frequency domain features such as spectral entropy and Mel-frequency cepstral coefficients (MFCC), and using models such as HMM or SVM to classify speech / non-speech to obtain accurate speech endpoint locations.

[0079] Through speech endpoint detection, the speech activity interval in the speech signal can be accurately located, and non-speech segments before and after, such as silence and noise, can be removed. This can reduce the amount of data and computing overhead for subsequent processing, and improve the real-time performance and efficiency of the system. At the same time, endpoint detection can also provide more accurate and complete speech input for speech recognition and text transcription, avoiding the interference and influence of non-speech segments on the recognition results.

[0080] S302, performing pre-emphasis processing on the second voice data after the noise reduction processing to obtain voice data after the emphasis processing, and performing voice framing processing on the voice data after the emphasis processing to obtain voice data after the voice framing; Specifically, the AR device first performs pre-emphasis on the second voice data after noise reduction. Pre-emphasis is a high-pass filtering operation on a voice signal. It enhances the energy of high-frequency components, compensates for the high-frequency attenuation of the voice signal during transmission and acquisition, and improves the clarity and intelligibility of the voice. The specific steps of pre-emphasis are as follows: 1. Perform Z transform on the second speech data after noise reduction processing to obtain its frequency domain representation S(z).

[0081] 2. Design a first-order high-pass filter H(z)=1-a*z^(-1), where a is the pre-emphasis coefficient, usually 0.9~0.95. The filter has gain in the high-frequency part and attenuation in the low-frequency part, which can compensate for the high-frequency attenuation of the speech signal.

[0082] 3. Multiply the frequency domain representation S(z) of the speech signal by the pre-emphasis filter H(z) to obtain the emphasized frequency domain representation S'(z)=S(z)*H(z).

[0083] 4. Perform an inverse Z transform on the emphasized speech frequency domain representation S'(z) to obtain the emphasized speech data in the time domain.

[0084] Pre-emphasis processing can effectively enhance the high-frequency components of the speech signal, making its energy distribution more uniform and improving the clarity and intelligibility of the speech. This is very beneficial for subsequent feature extraction and text transcription, and can provide richer and more effective speech information. At the same time, pre-emphasis can also suppress low-frequency noise to a certain extent, further improving the signal-to-noise ratio of the speech.

[0085] After obtaining the amplified voice data, the AR device performs voice framing processing on it to obtain voice data after voice framing. Voice framing refers to dividing the continuous voice signal into several short-time frames, each of which contains a voice segment of a fixed length. The specific steps of framing are as follows: 1. Determine the frame length and frame shift. The frame length indicates the time length of each speech frame.

[0086] 2. Divide the speech data after the emphasis processing into frames. Starting from the starting point of the speech data, take a speech segment of one frame length each time as a frame, and use the frame shift as the interval between two adjacent frames, and slide and intercept them in sequence until the end of the speech data.

[0087] 3. Perform windowing on each speech segment. In order to reduce the discontinuity at the edge of the frame, a Hamming window or a Hanning window is usually used to weight each speech segment so that it decays smoothly to zero at the edge of the frame. The windowed speech segment is more consistent with the stationarity assumption of the speech signal.

[0088] S303, performing windowing processing on the voice data after voice framing to obtain windowed voice data, and fusing the windowed voice data with the voice endpoint data to obtain pre-processed second voice data.

[0089] Specifically, after obtaining the voice data after voice framing, the AR device performs windowing processing on it. Windowing means multiplying each frame of voice fragment by a window function so that it decays smoothly to zero at the edge of the frame, reducing the discontinuity between frames.

[0090] The specific steps of windowing each frame of speech segment are as follows: 1. Based on the selected window function type and frame length, generate a window function vector with the same size as the frame length.

[0091] 2. Multiply the window function vector by each frame of speech fragment element by element to obtain the windowed speech fragment.

[0092] 3. The windowed speech segments are concatenated in sequence to obtain the windowed speech data.

[0093] Through windowing, the discontinuity caused by speech framing can be effectively reduced, so that each frame of speech segments smoothly transitions at the edge of the frame, which is more in line with the assumption of the stability of speech signals. This is very beneficial for subsequent feature extraction and text transcription, and can provide more accurate and reliable speech information and reduce the interference and distortion caused by frame edge effects.

[0094] After obtaining the windowed voice data, the AR device fuses it with the previously obtained voice endpoint data to obtain the preprocessed second voice data. The voice endpoint data contains the start and end point information of the voice activity interval, indicating the time range of the effective voice in the voice signal. Fusion of the windowed voice data with the voice endpoint data is to extract the voice frames within the voice activity interval to obtain a complete and accurate voice segment.

[0095] The specific steps of fusion are as follows: 1. According to the voice endpoint data, determine the frame index range of the voice activity interval in the windowed voice data.

[0096] 2. Extract the speech frames within the speech activity interval to obtain a new speech data sequence.

[0097] 3. The extracted speech data sequence is used as the pre-processed second speech data and output to the subsequent feature extraction and text transcription module.

[0098] By fusing the windowed speech data with the speech endpoint data, the effective speech segments in the speech signal can be accurately extracted, and invalid data such as silence and noise outside the speech activity interval can be removed. This can significantly reduce the amount of data and computing overhead for subsequent processing, and improve the real-time performance and efficiency of the system. At the same time, the fused speech data contains complete and accurate speech information, providing high-quality input for feature extraction and text transcription, which helps to improve the performance and accuracy of the entire offline text transcription system.

[0099] S107, applying the second speech feature data to a preset language recognition model to obtain preliminary target language data, and loading a corresponding first text transcription model according to the preliminary target language data; Specifically, the AR device inputs the extracted second voice feature data into a preset language recognition model. This language recognition model is the same as the model used in the online state. It uses machine learning models such as deep neural network (DNN) and convolutional neural network (CNN). It is trained through a large amount of multilingual voice data to learn the acoustic features and language characteristics of different languages. The input of the model is the second voice feature data, and the output is the preliminary target language data, which indicates the language category to which the model determines the user input voice belongs, such as Chinese, English, Japanese, etc.

[0100] After obtaining the preliminary target language data, the AR device loads the corresponding first text transcription model based on the data. The first text transcription model is a speech recognition and text transcription model for a specific language. It usually uses sequence models such as hidden Markov model (HMM) and recurrent neural network (RNN), combined with language model and pronunciation dictionary, to convert the speech signal into a text sequence. The AR device pre-stores the first text transcription models of multiple languages, and each model is specifically optimized for a language. According to the preliminary target language data, the AR device selects the first text transcription model corresponding to the language from the stored model library, loads it into the memory, and prepares for offline speech recognition and text transcription.

[0101] S108: Input the second voice feature data into the corresponding first text transcription model, output a text transcription result, and display the text transcription result on the AR device.

[0102] Specifically, the AR device inputs the extracted second voice feature data into the corresponding first text transcription model. The first text transcription model is a speech recognition and text transcription model that is loaded based on the preliminary target language data and matches the language of the user's input voice. The model uses sequence models such as hidden Markov models (HMM) and recurrent neural networks (RNN) and is trained through a large amount of speech-text alignment data to learn the mapping relationship between speech signals and text sequences. The input of the model is the second voice feature data, and the output is the text transcription result, which means that the model converts the input voice signal into the corresponding text content.

[0103] During the inference process of the model, the second speech feature data is first fed into the acoustic model, and the probability distribution of the phonemes or characters corresponding to each frame of speech features is calculated through algorithms such as HMM or RNN. Then, the obtained probability distribution sequence is input into the language model, which scores and sorts the candidate text sequences based on contextual information and prior knowledge, and selects the text with the highest probability as the output. Finally, the characters in the text are converted into corresponding words through the pronunciation dictionary to obtain the final text transcription result. The entire process is performed locally on the AR device side without relying on network connection and edge device side.

[0104] After obtaining the text transcription results, the AR device displays them on the device's screen so that the user can view the transcribed text content in real time.

[0105] In the second aspect of this application, see Figure 2 , Figure 2 The present application also provides a multilingual text transcription method, which is applied to an edge device end, and the method includes the following steps: S401, receiving preliminary target language data and a plurality of first feature data to be transcribed sent by an AR device; Specifically, after completing the preprocessing and feature extraction of the speech, vision and motion data, the AR device obtains preliminary target language data and multiple first feature data to be transcribed. The preliminary target language data is the result of the AR device making a preliminary judgment on the target language based on the speech, vision and motion features using the local language recognition model. The multiple first feature data to be transcribed include the first speech feature data, the second visual feature data and the first motion feature data, which respectively represent the multimodal features extracted from the speech, vision and motion data on the AR device.

[0106] In order to further improve the accuracy of language recognition and text transcription, the AR device sends the preliminary target language data and the first feature data to be transcribed to the edge device. This process is usually carried out through wireless networks, such as WiFi, 4G / 5G, etc. The AR device packages the data into one or more network data packets and sends the data packets to the IP address and port of the edge device through the network protocol stack and network interface.

[0107] The edge device listens to the data transmission request from the AR device through the pre-configured network interface and protocol. After receiving the network data packet sent by the AR device, the edge device parses and extracts the data packet to obtain the preliminary target language data and the first feature data to be transcribed.

[0108] After receiving this data, the edge device can use its powerful computing and storage resources to further process and analyze the data. Compared with AR devices, edge devices usually have larger model capacity, richer training data and higher computing performance, which can achieve more accurate and comprehensive language recognition and text transcription.

[0109] S402, extracting acoustic features from the first speech feature data through an acoustic model to obtain acoustic features, and inputting the acoustic features and preliminary target language data into a preset target language classification model to output final target language data; Specifically, the edge device first uses the acoustic model to further extract acoustic features from the first voice feature data uploaded by the AR device. The first voice feature data usually includes the basic features of the voice signal, such as Mel-frequency cepstral coefficients (MFCC), linear prediction cepstral coefficients (LPCC), etc., which reflect the spectrum and resonance peak characteristics of the voice signal. However, these basic features are not enough to fully characterize the acoustic properties of the voice signal, especially in cross-language scenarios.

[0110] In order to extract richer and more effective acoustic features, edge devices use pre-trained acoustic models, such as deep neural networks (DNNs) and convolutional neural networks (CNNs), to further transform and abstract the first voice feature data. Acoustic models are usually trained on large-scale multilingual speech datasets to learn the differences and commonalities in acoustic features of different languages.

[0111] The input of the acoustic model is the first speech feature data, and the output is a higher-dimensional and more abstract acoustic feature. The specific steps for extracting acoustic features are as follows: 1. Arrange and align the first speech feature data according to the input format of the acoustic model to construct an input tensor.

[0112] 2. Send the input tensor to the acoustic model and obtain the output tensor of the model through forward propagation calculation.

[0113] 3. Post-process the output tensor, such as dimensionality reduction and normalization, to obtain the final acoustic features.

[0114] The acoustic features extracted by the acoustic model not only retain the basic spectrum and resonance peak characteristics of the speech signal, but also capture more high-level and abstract acoustic properties, such as the rhythm, intonation, and pronunciation of the speech, which is very helpful for language recognition.

[0115] After obtaining the acoustic features, the edge device inputs them together with the preliminary target language data uploaded by the AR device into the preset target language classification model to further judge and confirm the target language. The target language classification model usually uses machine learning methods, such as support vector machine (SVM), random forest (RF), neural network, etc., and is trained on large-scale multilingual speech data sets to learn the discrimination boundaries of different languages ​​in terms of acoustic features and other features.

[0116] The specific steps for inputting acoustic features and preliminary target language data into the target language classification model are as follows: 1. Arrange and align the acoustic features and preliminary target language data according to the input format of the classification model to construct the input vector.

[0117] 2. Send the input vector into the classification model and obtain the output probability distribution of the model through forward propagation calculation.

[0118] 3. Post-process the output probability distribution, such as threshold judgment, softmax normalization, etc., to obtain the final target language data.

[0119] The target language classification model comprehensively considers the acoustic features and the preliminary judgment results of the AR device segment to identify the target language more accurately and reliably. Through the acoustic features, the classification model can capture the language-related acoustic properties contained in the speech signal, such as speech rhythm, pronunciation, etc. Through the preliminary target language data, the classification model can understand the local language judgment tendency of the AR device end, and verify and correct it.

[0120] S403, performing feature fusion on the first voice feature data, the first visual feature data, and the first action feature data to obtain multimodal feature data; Specifically, the first voice feature data, the first visual feature data, and the first motion feature data respectively describe the feature information of the user input from three different modal perspectives: voice, vision, and motion. The first voice feature data usually includes acoustic features such as the spectrum, fundamental frequency, and resonance peak of the voice signal, reflecting the acoustic properties of the voice. The first visual feature data usually includes visual features such as key points, expressions, and lip movements of the user's face, reflecting the facial visual clues of the user when speaking. The first motion feature data usually includes motion features such as user gestures, posture, and motion trajectories, reflecting the body motion information of the user when speaking.

[0121] The edge device uses multimodal feature fusion technology to fuse the first voice feature data, the first visual feature data, and the first action feature data at the feature level to obtain a unified multimodal feature representation. Commonly used multimodal feature fusion methods include splicing fusion, attention fusion, bilinear pooling fusion, etc. Taking splicing fusion as an example, the specific steps are as follows: 1. Align the first speech feature data, the first visual feature data, and the first action feature data along the time axis to construct three feature sequences.

[0122] 2. Concatenate the three feature sequences in the feature dimension to obtain a high-dimensional multimodal feature sequence.

[0123] 3. Perform dimensionality reduction and normalization on the spliced ​​multimodal feature sequence to obtain the final multimodal feature data.

[0124] Through splicing and fusion, the feature data of different modalities are unified into a common feature space, realizing feature-level information fusion. The fused multimodal feature data not only retains the feature information of a single modality, but also captures the correlation and complementarity between different modalities, providing a more comprehensive and in-depth representation of user input.

[0125] Based on the above embodiment, as an optional embodiment, feature fusion is performed on the first voice feature data, the first visual feature data and the first action feature data to obtain multimodal feature data, including: S501, concatenating the first speech feature data, the first visual feature data and the first action feature data to obtain a high-dimensional feature vector, and mapping the first high-dimensional feature vector pair to a common feature space to obtain a second high-dimensional feature vector; Specifically, assume that the dimension of the first voice feature data is d1, the dimension of the first visual feature data is d2, and the dimension of the first action feature data is d3. The edge device first concatenates the three feature data in the feature dimension to obtain a first high-dimensional feature vector of d1+d2+d3 dimensions. The concatenation order can be arbitrary, but it needs to be consistent.

[0126] For example, the first speech feature data can be placed in the first d1 dimensions of the first high-dimensional feature vector, the first visual feature data can be placed in the middle d2 dimensions, and the first action feature data can be placed in the last d3 dimensions. In this way, the first high-dimensional feature vector contains complete feature information of the three modalities of speech, vision, and action.

[0127] It should be noted that since the feature data of different modalities usually have different numerical ranges and distribution characteristics, direct splicing may cause the features of certain modalities to dominate and affect the fusion effect. Therefore, before splicing, it is usually necessary to normalize or standardize the feature data of each modality so that its numerical distribution is within a similar range. Secondly, the first high-dimensional feature vector is mapped to a common feature space to obtain the second high-dimensional feature vector. The purpose of this step is to map the features of different modalities into a unified feature representation space, eliminate the differences between modalities, and extract the common features of different modalities.

[0128] S502, using the attention mechanism to adaptively perform weighted fusion on the second high-dimensional feature vector to obtain a third high-dimensional feature vector, and using a preset deep neural network to fuse the fourth high-dimensional feature vector to obtain multimodal feature data.

[0129] Specifically, the attention mechanism generates a weight vector by calculating the relevance of different modal features to the current task or context, and performs weighted summation on each feature dimension in the second high-dimensional feature vector to obtain a third high-dimensional feature vector. The calculation of the weight vector is usually based on the query-key-value attention model, where the query vector represents the current task or context, and the key vector and value vector represent each feature dimension in the second high-dimensional feature vector.

[0130] Based on the above embodiment, as an optional embodiment, the second high-dimensional feature vector is adaptively weighted and fused using an attention mechanism to obtain a third high-dimensional feature vector, including: S601, for each feature vector in the second high-dimensional feature vector, calculating the attention weight of the feature vector; Specifically, for each feature vector in the second highest dimensional feature vector, the edge device uses the attention mechanism to calculate a scalar weight, which indicates the relevance and importance of the feature vector to the current task or context. The larger the weight value, the more important the feature vector is, and a higher weight should be given when fusion is performed; the smaller the weight value, the less important the feature vector is, and a lower weight should be given when fusion is performed.

[0131] Assuming that the second highest dimensional feature vector is composed of n d-dimensional feature vectors, it can be represented as an n×d matrix. The attention mechanism calculates the relevance of these n feature vectors with the current task or context to obtain an n-dimensional attention weight vector.

[0132] Common methods for calculating attention weights include dot product attention, additive attention, scaled dot product attention, etc. Taking scaled dot product attention as an example, it first performs a dot product of each d-dimensional feature vector with a d-dimensional query vector to obtain a scalar relevance score. The query vector represents the current task or context, which can be obtained through learning or set according to the task prior. Then, these n relevance scores are divided by a scaling factor, usually √d, to improve the stability of the gradient. Finally, the scaled relevance scores are normalized by the Softmax function to obtain the attention weights of the n feature vectors. The Softmax function ensures that the sum of all weights is 1 and the weights are non-negative, which conforms to the properties of the probability distribution.

[0133] For example, suppose the second highest dimensional feature vector contains three 4-dimensional feature vectors, and the query vector is also 4-dimensional. First, do a dot product of each feature vector with the query vector to get three scalar relevance scores, 2.1, 1.5, and 0.8. Then, divide these three scores by √4=2 to get 1.05, 0.75, and 0.4. Finally, normalize these three scaled scores through the Softmax function to get the attention weights of the three feature vectors, which are 0.47, 0.35, and 0.18, respectively.

[0134] These attention weights quantitatively reflect the importance differences of different feature vectors. Feature vectors with larger weights will play a greater role in fusion, and feature vectors with smaller weights will play a smaller role in fusion. Therefore, the attention mechanism can adaptively adjust the feature fusion strategy, highlight more relevant and important features, and improve the fusion effect.

[0135] S602, weighting each feature vector in the second high-dimensional feature vector based on the attention weight to obtain a weighted feature vector; Specifically, the edge device multiplies each feature vector in the second highest dimensional feature vector by its corresponding attention weight to obtain a weighted feature vector. Since the attention weight is a scalar, the weighting operation does not change the dimension of the feature vector, but only changes the numerical value of the feature vector.

[0136] S603: Fusing the weighted feature vectors to obtain a third high-dimensional feature vector.

[0137] Specifically, edge devices can fuse weighted feature vectors in a variety of ways, such as concatenation, summation, averaging, and nonlinear transformation. Concatenation is the simplest and most commonly used fusion method, which is to connect all weighted feature vectors end to end to obtain a longer feature vector. Summation and averaging are to add or average the elements at corresponding positions of all weighted feature vectors to obtain a fused feature vector of the same length as a single feature vector. Nonlinear transformation is to perform nonlinear function transformation on the concatenated, summed, or averaged feature vectors, such as Sigmoid, Tanh, ReLU, etc., to improve the nonlinearity and abstraction of feature representation.

[0138] Assume that the second high-dimensional feature vector is composed of n d-dimensional feature vectors, and after weighting, n d-dimensional weighted feature vectors are obtained. If the splicing fusion method is adopted, the third high-dimensional feature vector obtained is nd-dimensional; if the summation or averaging fusion method is adopted, the third high-dimensional feature vector obtained is d-dimensional; if a nonlinear transformation is performed after summation or averaging, the dimension of the third high-dimensional feature vector obtained is consistent with the output dimension of the transformation function.

[0139] For example, continuing the calculation results of the previous step, assume that the three weighted 4-dimensional feature vectors are fused. If the splicing fusion method is used, a 12-dimensional third-high-dimensional feature vector is obtained. If the summation fusion method is used, the elements of the corresponding positions of the three feature vectors are added to obtain a 4-dimensional third-high-dimensional feature vector. If the ReLU transformation is performed after the summation, a 4-dimensional non-negative third-high-dimensional feature vector is obtained.

[0140] The fused third high-dimensional feature vector integrates the feature information of different modalities, and after adaptive weighting by the attention mechanism, it further highlights the features related to the current task or context. This compact and high-level feature representation can make full use of the complementary information of different modalities, reduce feature redundancy and noise, and improve the quality and discrimination ability of feature representation.

[0141] S404, loading a corresponding second transcription model based on the final target language data, and inputting the multimodal feature data into the second transcription model to obtain a text transcription result; Specifically, the edge device first determines the target language category of the user's voice input, such as Chinese, English, Japanese, etc., based on the final target language data obtained in the previous steps. Since different languages ​​have significant differences in pronunciation, vocabulary, grammar, etc., in order to obtain the best transcription effect, it is necessary to build a dedicated transcription model for each language and make full use of language-specific linguistic knowledge and training data.

[0142] Therefore, the edge device will pre-train multiple second transcription models for different languages, each of which is a language-specific end-to-end text transcription model, such as an attention-based encoder-decoder model, a Transformer-based sequence-to-sequence model, etc. These models are trained on large-scale speech transcription datasets in specific languages ​​to learn the mapping relationship between speech and text, as well as language-specific pronunciation, vocabulary, and grammar rules.

[0143] After obtaining the final target language data, the edge device selects the model corresponding to the target language from multiple pre-trained second transcription models according to the language label, and loads its parameters and configuration. In this way, the subsequent text transcription process can make full use of the linguistic knowledge of the target language to obtain more accurate and natural transcription results.

[0144] After loading the language-specific second transcription model, the edge device inputs the multimodal feature data obtained in the previous step into the model for text transcription inference. Multimodal feature data integrates speech, vision, and motion information to provide a more comprehensive and rich representation of user speech input. The second transcription model encodes the multimodal feature data through an encoder, converts it into a latent space representation, and then decodes the latent space representation through a decoder to generate the corresponding transcribed text.

[0145] The specific steps of inputting multimodal feature data into the second transcription model and performing inference are as follows: 1. Arrange and align the multimodal feature data according to the input format of the second transcription model to construct the input tensor.

[0146] 2. The input tensor is fed into the encoder of the second transcription model and the latent space representation is obtained through forward propagation calculation.

[0147] 3. The latent space representation is sent to the decoder of the second transcription model to generate a transcribed text sequence through autoregressive decoding or beam search decoding.

[0148] 4. Post-process the transcribed text sequence, such as removing duplicates, removing stop words, formatting, etc., to obtain the final text transcription result.

[0149] Through the inference process of the second transcription model, multimodal feature data that integrates speech, vision and motion information is converted into corresponding transcription text. Since the second transcription model is language-specific and has been fully trained on a large-scale corpus, the transcription text it generates is usually highly accurate, fluent and readable.

[0150] S405: Send the text transcription result to the AR device.

[0151] Specifically, after the edge device completes the text transcription task, it will get a high-quality transcribed text, such as "Meeting in the company lobby at 3 pm tomorrow." In order for the transcribed text to be effectively used by the AR device, the edge device needs to send the text transcription result to the AR device through network communication.

[0152] The specific steps for the edge device to send the text transcription results are as follows: 1. Encapsulate the text transcription results according to the agreed data format and construct a transmission data packet. Common data formats can be JSON, XML, etc. The data packet usually includes meta information such as the transcribed text content, timestamp, and confidence level.

[0153] 2. Create a network connection based on the network address and port number of the AR device. The network connection can be a long connection based on the TCP / IP protocol or a short connection based on application layer protocols such as HTTP and WebSocket.

[0154] 3. Send the encapsulated transmission data packet to the AR device through the network connection. The sending process can be synchronous or asynchronous, depending on the specific application scenario and performance requirements.

[0155] 4. Wait for the AR device to receive confirmation information to ensure that the text transcription result is sent successfully. If the confirmation information is not received for a long time, the edge device can resend the data packet or adopt other exception handling mechanisms.

[0156] After receiving the text transcription results sent by the edge device, the AR device can process and display them according to specific application requirements. For example, the AR device can display the transcribed text in real time in the user's field of view as subtitles or prompt information; it can also parse the transcribed text into specific voice commands to trigger corresponding functions or operations; it can also store the transcribed text locally or in the cloud as historical records or data analysis materials.

[0157] By sending text transcription results to the AR device through the edge device, effective coordination and integration of cloud-edge resources can be achieved.

[0158] In the third aspect of this application, please refer to Figure 3 , Figure 3 This is a system architecture diagram for an AR device provided in an embodiment of the present application. The present application provides a multilingual text transcription system, which is applied to an AR device. The system includes: A first data acquisition module 1 is used for collecting first data to be transcribed of the user through the AR device end when the AR device end is in an online state, and preprocessing the first data to be transcribed to obtain the preprocessed data to be transcribed, where the first data to be transcribed includes voice data, visual data and action data; A first feature extraction module 2, used for extracting features from the first data to be transcribed to obtain first feature data to be transcribed, where the first feature data to be transcribed includes first speech feature data, first visual feature data and first action feature data; A first preliminary target language determination module 3, used to input the first speech feature data into a preset language recognition model, output preliminary target language data, and send the first feature data to be transcribed and the preliminary target language data to the edge device end; The text transcription result receiving module 4 is used to receive the text transcription result and display the text transcription result on the AR device.

[0159] In the fourth aspect of this application, please refer to Figure 4 , Figure 4 This is a system architecture diagram for edge device end provided in an embodiment of the present application. The present application provides a multilingual text transcription system, which is applied to edge device end. The system includes: The data receiving module 5 is used to receive the preliminary target language data and the first feature data to be transcribed sent by the AR device end; The final target language determination module 6 is used to extract acoustic features from the first speech feature data through an acoustic model to obtain acoustic features, and input the acoustic features and preliminary target language data into a preset target language classification model to output final target language data; A feature fusion module 7, used for fusing the first speech feature data, the first visual feature data and the first action feature data to obtain multimodal feature data; A first transcription module 8, used to load a corresponding second transcription model based on the final target language data, and input the multimodal feature data into the second transcription model to obtain a text transcription result; The text transcription result sending module 9 is used to send the text transcription result to the AR device end.

[0160] In a fifth aspect of the present application, a computer storage medium is provided, wherein the computer storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor and executing the above-mentioned method steps.

[0161] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0162] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a memory and includes several instructions for a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned memory includes: various media that can store program codes, such as USB flash drives, mobile hard drives, magnetic disks or optical disks.

[0163] The above are only exemplary embodiments of the present disclosure and cannot be used to limit the scope of the present disclosure. That is, any equivalent changes and modifications made according to the teachings of the present disclosure are still within the scope of the present disclosure. After considering the disclosure of the specification and practice, those skilled in the art will easily think of other embodiments of the present disclosure.

[0164] This application is intended to cover any variation, use or adaptation of the present disclosure, which follows the general principles of the present disclosure and includes common knowledge or customary technical means in the art not described in the present disclosure. The description and examples are to be regarded as exemplary only, and the scope and spirit of the present disclosure are defined by the claims.

Claims

1. A multilingual text transcription method, characterized in that: Applied to an AR device, the method includes: When the AR device is in an online state, first data to be transcribed of the user is collected through the AR device, and the first data to be transcribed is preprocessed to obtain preprocessed data to be transcribed, where the first data to be transcribed includes voice data, visual data, and action data; Performing feature extraction on the first data to be transcribed to obtain first feature data to be transcribed, wherein the first feature data to be transcribed includes first voice feature data, first visual feature data, and first action feature data; Inputting the first speech feature data into a preset language recognition model, outputting preliminary target language data, and sending the first feature data to be transcribed and the preliminary target language data to an edge device; The text transcription result is received, and the text transcription result is displayed on the AR device.

2. The method according to claim 1, characterized in that The method further comprises: When the AR device is in an offline state, collecting second data to be transcribed of the user through the AR device, and preprocessing the first data to be transcribed to obtain preprocessed second data to be transcribed, where the second data to be transcribed includes second voice data; Preprocessing the second voice data to obtain preprocessed second voice data, and extracting features from the preprocessed second voice data to obtain second voice feature data; The second speech feature data belongs to the preset language recognition model to obtain the preliminary target language data, and according to the preliminary target language data, load the corresponding first text transcription model; The second voice feature data is input into the corresponding first text transcription model, a text transcription result is output, and the text transcription result is displayed on the AR device.

3. The method according to claim 1, characterized in that The feature extraction of the first data to be transcribed is performed to obtain first feature data to be transcribed, wherein the first feature data to be transcribed includes first voice feature data, first visual feature data and first action feature data, including: Performing feature extraction on the speech data through a GMM-HMM model to obtain the first speech feature data; Extracting features from the visual data using an image feature extraction algorithm to obtain the second visual feature data; The action data is preprocessed to obtain preprocessed action data, and features are extracted from the preprocessed action data to obtain the first action feature data, wherein the first action feature data includes time domain statistical features, time-frequency domain features and time series features.

4. The method according to claim 2, characterized in that: The preprocessing of the second voice data to obtain the preprocessed second voice data includes: Performing noise reduction processing on the second voice data by spectral subtraction to obtain second voice data after noise reduction processing, and performing voice endpoint detection on the second voice data after noise reduction processing to obtain voice endpoint data; Performing pre-emphasis processing on the second voice data after the noise reduction processing to obtain voice data after the emphasis processing, and performing voice framing processing on the voice data after the emphasis processing to obtain voice data after the voice framing; The voice data after the voice framing is subjected to windowing processing to obtain the windowed voice data, and the windowed voice data is fused with the voice endpoint data to obtain the pre-processed second voice data.

5. A multilingual text transcription method, characterized in that: Applied to the edge device side, the method includes: Receiving the preliminary target language data and the plurality of first feature data to be transcribed sent by the AR device; Extracting acoustic features from the first speech feature data using the acoustic model to obtain acoustic features, and inputting the acoustic features and the preliminary target language data into a preset target language classification model to output final target language data; Performing feature fusion on the first speech feature data, the first visual feature data, and the first action feature data to obtain multimodal feature data; Loading a corresponding second transcription model based on the final target language data, and inputting the multimodal feature data into the second transcription model to obtain the text transcription result; The text transcription result is sent to the AR device.

6. The method according to claim 5, characterized in that The step of fusing the first voice feature data, the first visual feature data, and the first action feature data to obtain multimodal feature data includes: concatenating the first speech feature data, the first visual feature data, and the first action feature data to obtain a high-dimensional feature vector, and mapping the first high-dimensional feature vector pair to a common feature space to obtain a second high-dimensional feature vector; The second high-dimensional feature vector is adaptively weighted and fused using an attention mechanism to obtain a third high-dimensional feature vector, and the fourth high-dimensional feature vector is fused using a preset deep neural network to obtain the multimodal feature data.

7. The method according to claim 6, characterized in that The step of adaptively performing weighted fusion on the second high-dimensional feature vector using the attention mechanism to obtain a third high-dimensional feature vector includes: For each feature vector in the second high-dimensional feature vector, calculating an attention weight of the feature vector; Weighting each of the feature vectors in the second high-dimensional feature vector based on the attention weight to obtain a weighted feature vector; The weighted feature vectors are fused to obtain the third high-dimensional feature vector.

8. A multilingual text transcription system, characterized in that: Applied to an AR device, the system includes: a first data acquisition module, configured to collect first data to be transcribed of the user through the AR device end when the AR device end is in an online state, and preprocess the first data to be transcribed to obtain preprocessed data to be transcribed, wherein the first data to be transcribed includes voice data, visual data, and action data; A first feature extraction module, used for performing feature extraction on the first data to be transcribed to obtain first feature data to be transcribed, wherein the first feature data to be transcribed includes first voice feature data, first visual feature data and first action feature data; A first preliminary target language determination module, configured to input the first speech feature data into a preset language recognition model, output preliminary target language data, and send the first feature data to be transcribed and the preliminary target language data to an edge device end; The text transcription result receiving module is used to receive the text transcription result and display the text transcription result on the AR device.

9. A multilingual text transcription system, characterized in that: Applied to the edge device side, the system includes: A data receiving module, configured to receive the preliminary target language data and the plurality of first feature data to be transcribed sent by the AR device; A final target language determination module, configured to extract acoustic features from the first speech feature data using the acoustic model to obtain acoustic features, input the acoustic features and the preliminary target language data into a preset target language classification model, and output final target language data; A feature fusion module, used for fusing the first speech feature data, the first visual feature data and the first action feature data to obtain multimodal feature data; A first transcription module, configured to load a corresponding second transcription model based on the final target language data, and input the multimodal feature data into the second transcription model to obtain the text transcription result; The text transcription result sending module is used to send the text transcription result to the AR device end.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores instructions, and when the instructions are executed, the method steps according to any one of claims 1 to 7 are performed.

Citation Information

Cited By

  • Multi-language automatic identification method and system

    CN121687012A

  • A multilingual automatic recognition method and system

    CN121687012B