Digital twin Web application control method and device, computer equipment and storage medium
Through voice recognition technology, users can control digital twin web applications through voice, solving the problems of complex traditional operations and insufficient barrier-free design, and achieving a more intuitive and intelligent operation experience.
Patent Information
- Application Number
- CN202411785557.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-06
- Publication Date
- 2025-06-06
AI Technical Summary
Traditional digital twin web applications have complex operations, high learning costs, and lack of barrier-free design, making it difficult to meet the intelligent and humanized needs of modern users.
Obtain user voice information through the microphone to determine whether it is a wake-up speech. If so, continue to obtain user voice and convert it into text. According to the conversion result, identify user intentions and call the SDK API of the digital twin rendering engine for rendering.
It realizes a more intuitive operation of digital twin web applications without memory of cumbersome instructions, improving the practicality and user experience of the application, especially providing a more friendly operating experience for users with physiological disorders.
Smart Images

Figure CN120104092A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a digital twin system, and more specifically to a digital twin Web application control method, device, computer equipment and storage medium. Background Art
[0002] Digital twin web applications are online platforms that simulate real objects or systems through digital means. Such applications allow users to monitor, analyze and manage physical objects or processes in real time, usually involving data collection, visualization and interactive functions. Through digital twin technology, users can better understand and optimize complex systems, improve efficiency and decision-making capabilities.
[0003] Traditional digital twin Web applications usually rely on manual operation of the mouse, keyboard or touch screen for control. This approach has many limitations. First, the complex operating instructions are not friendly to novices, resulting in high learning costs; second, the lack of barrier-free design makes it difficult for users with visual or motor impairments to use them smoothly. In addition, current applications are relatively weak in intelligence and humanization, and cannot meet the expectations of modern users.
[0004] Therefore, it is necessary to design a new method to realize the operation of digital twin Web applications more intuitively, without the need to remember cumbersome instructions, and to improve the practicality of the application. Summary of the invention
[0005] The purpose of the present invention is to overcome the defects of the prior art and provide a digital twin Web application control method, device, computer equipment and storage medium.
[0006] To achieve the above object, the present invention adopts the following technical solution: a digital twin Web application control method, comprising:
[0007] Acquire the voice input by the microphone to obtain voice information;
[0008] Determining whether the voice information is a wake-up word;
[0009] If the voice information is a wake-up word, continue to obtain the user's voice;
[0010] Convert the user's speech into text to obtain a conversion result;
[0011] Identifying user intent and corresponding parameters for the conversion result;
[0012] According to the user intention and corresponding parameters, the relevant digital twin rendering engine SDK API is called to render the digital twin.
[0013] A further technical solution is: the step of determining whether the voice information is a wake-up word includes:
[0014] Performing frame processing on the voice information to obtain a frame result;
[0015] The frequency feature of each frame in the framing result is converted into a Mel frequency scale using the Mel spectrum to obtain a frequency conversion result;
[0016] Inputting the frequency conversion result into a classification model to obtain a recognition result;
[0017] Determine whether the voice information is a wake-up word according to the recognition result.
[0018] A further technical solution is: inputting the frequency conversion result into a classification model to obtain a recognition result, including:
[0019] A sliding window is applied to the audio stream, and the audio blocks corresponding to the frequency conversion results are analyzed frame by frame. The audio blocks are input into the classification model for hot word detection to obtain recognition results.
[0020] A further technical solution is: converting the user's voice into text to obtain a conversion result includes:
[0021] The user speech is converted into text through the SpeechRecognition interface in the Web Speech API to obtain a conversion result.
[0022] A further technical solution is: the identifying of the user intention and the corresponding parameters of the conversion result includes:
[0023] The conversion result is input into the recognition model to recognize the user's intention and determine the corresponding parameters.
[0024] Its further technical solution is: the intent recognition model includes an intent classification model and a parameter extraction model, wherein the intent classification model is obtained by training a pre-trained model of natural language processing using a number of natural language commands covering common scenarios of all digital twin applications and the corresponding SDK API as a sample set, and the parameter extraction model is obtained by training a named entity recognition model or a custom rule using a mapping relationship constructed by a number of natural language commands covering common scenarios of all digital twin applications and the corresponding SDK API as a sample set.
[0025] Its further technical solution is: the parameters include device ID or status, and related digital twin API parameters.
[0026] The present invention also provides a digital twin Web application control device, comprising:
[0027] A voice acquisition unit, used to acquire the voice input by the microphone to obtain voice information;
[0028] A judging unit, used to judge whether the voice information is a wake-up word;
[0029] A continuous acquisition unit, configured to continuously acquire user voice if the voice information is a wake-up word;
[0030] A conversion unit, used for converting the user speech into text to obtain a conversion result;
[0031] An identification unit, used to identify the user intention and corresponding parameters of the conversion result;
[0032] The calling unit is used to call the relevant digital twin rendering engine SDK API according to the user intention and corresponding parameters to perform digital twin rendering.
[0033] The present invention further provides a computer device, which includes a memory and a processor. The memory stores a computer program, and the processor implements the above method when executing the computer program.
[0034] The present invention also provides a storage medium, wherein the storage medium stores a computer program, and the computer program implements the above method when executed by a processor.
[0035] The beneficial effects of the present invention compared with the prior art are: the present invention obtains the user's voice information through a microphone and determines whether it is a wake-up word. If it is a wake-up word, the present invention continues to obtain the user's voice and converts it into text, identifies the user's intention based on the conversion result, and calls the SDK API of the digital twin rendering engine for rendering; the digital twin Web application can be operated more intuitively without memorizing cumbersome instructions, thereby improving the practicality of the application.
[0036] The present invention is further described below in conjunction with the accompanying drawings and specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying any creative work.
[0038] Figure 1 A schematic diagram of an application scenario of a digital twin Web application control method provided by an embodiment of the present invention;
[0039] Figure 2A schematic diagram of a flow chart of a digital twin Web application control method provided by an embodiment of the present invention;
[0040] Figure 3 A schematic diagram of a sub-process of a digital twin Web application control method provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0041] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0042] It should be understood that when used in this specification and the appended claims, the terms "include" and "comprises" indicate the presence of described features, integers, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or combinations thereof.
[0043] It should also be understood that the terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.
[0044] It should be further understood that the term "and / or" used in the present description and the appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0045] See also Figure 1 and Figure 2 , Figure 1 A schematic diagram of an application scenario of the digital twin Web application control method provided in an embodiment of the present invention. Figure 2 A schematic flow chart of a digital twin Web application control method provided in an embodiment of the present invention. The digital twin Web application control method is applied to a server. The server interacts with a microphone and a digital twin Web application for data, and the introduction of natural language interaction can effectively solve these problems. Through natural language, users can operate more intuitively without having to remember cumbersome instructions, while providing a more friendly operating experience for users with physiological disorders. In addition, the interactive method not only improves the intelligence level of the application, but also enhances its market competitiveness, allowing users to experience the convenience of cutting-edge technology.
[0046] Figure 2is a flow chart of a digital twin Web application control method provided by an embodiment of the present invention. Figure 2 As shown, the method includes the following steps S110 to S160.
[0047] S110: Acquire voice input by a microphone to obtain voice information.
[0048] In this embodiment, the user's voice input is captured by a microphone of the device to form voice information.
[0049] S120: Determine whether the voice information is a wake-up word.
[0050] In this embodiment, the wake-up word is a user-defined word, and the user can customize the wake-up word through the Web interface to meet personal needs.
[0051] In one embodiment, see Figure 3 , the above-mentioned step S120 may include steps S121 to S124.
[0052] S121. Perform frame processing on the voice information to obtain a frame processing result.
[0053] In this embodiment, the frame division result refers to the speech segment formed after the speech information is framed.
[0054] Specifically, the continuous speech signal is divided into multiple short time frames for detailed analysis. Each frame is usually short to ensure that the rapid changes in speech can be captured.
[0055] S122. Use the Mel spectrum to convert the frequency feature of each frame in the frame division result into a Mel frequency scale to obtain a frequency conversion result.
[0056] In this embodiment, the frequency conversion result refers to the speech segment after the frequency feature of each frame is converted into the Mel frequency scale.
[0057] The Mel spectrum is used to convert the frequency characteristics of each frame into the Mel frequency scale. This scale is more in line with the auditory characteristics of the human ear and can better represent the audio signal.
[0058] S123: Input the frequency conversion result into a classification model to obtain a recognition result.
[0059] In this embodiment, a sliding window is applied to the audio stream, the audio block corresponding to the frequency conversion result is analyzed frame by frame, and the audio block is input into the classification model for hot word detection to obtain the recognition result.
[0060] A sliding window is applied to the entire audio stream to continuously analyze small audio blocks. These audio blocks are input into the classifier model for hot word detection. If a hot word is detected, a response will be made to wake up the subsequent speech recognition tool.
[0061] Specifically, the frequency conversion results are input into the classification model, and the frame-by-frame sliding analysis is performed to detect hot words and identify whether there is a wake-up word.
[0062] The above classification model is trained as follows:
[0063] Data Collection:
[0064] Diversity: Collect audio data from a variety of contexts, including speakers of different genders, ages, and accents.
[0065] Sample size: Make sure the sample size is large enough so that the model can learn rich features.
[0066] Data preprocessing:
[0067] Frame processing: Divide the continuous audio signal into multiple short-time frames, usually 20-40 milliseconds per frame, for easy processing.
[0068] Feature extraction: Mel spectrum: Mel spectrum is calculated using Fast Fourier Transform (FFT) to capture frequency features. MFCC (Mel Frequency Cepstral Coefficient): further extracts features to reduce data dimensionality and improve model recognition capabilities.
[0069] Standardization: Standardize the extracted features so that their mean is 0 and their variance is 1 to improve the model training effect.
[0070] Label annotation: Classification label: Assign a label to each audio sample to indicate whether it contains the target wake-up word.
[0071] Balanced dataset: Make sure the number of wake-up word and non-wake-up word samples is as balanced as possible to prevent the model from being biased towards one category.
[0072] Model selection
[0073] Select an algorithm: Traditional machine learning: such as support vector machine (SVM), random forest, etc. Deep learning: such as convolutional neural network (CNN) and recurrent neural network (RNN), which are more suitable for processing complex audio features. Architecture design: Choose an appropriate model architecture based on data characteristics, such as multiple convolutional layers and pooling layers, followed by a fully connected layer.
[0074] Training process:
[0075] Loss function: Select a suitable loss function (such as cross entropy loss) to measure the gap between the model output and the true label.
[0076] Optimization algorithm: Use optimization algorithms such as Adam or SGD to adjust model weights and reduce losses.
[0077] Hyperparameter adjustment: such as learning rate, batch size, regularization, etc., are optimized through cross-validation and other methods.
[0078] Iterative training: Perform multiple rounds of iterations until the model's performance on the validation set becomes stable.
[0079] Model evaluation: Evaluate model performance on an independent validation set. Common metrics include accuracy, recall, and F1 score.
[0080] Overfitting detection: monitor the performance difference between the training set and the validation set to prevent the model from overfitting.
[0081] Real-time recognition: Integrate the trained model into practical applications to detect real-time voice streams.
[0082] Update and maintenance: The model is updated regularly based on actual usage to improve recognition accuracy and adaptability.
[0083] Through these steps, the model can effectively identify the voice features that match the wake-up word, thereby improving the user experience of the system.
[0084] S124. Determine whether the voice information is a wake-up word according to the recognition result.
[0085] According to the output of the classification model, it is determined whether the input voice information is the wake-up word set by the user. If the recognition result contains the wake-up word, it indicates that the voice information is the wake-up word, otherwise, the voice information is not the wake-up word.
[0086] In this embodiment, the user is allowed to set the wake-up word according to personal needs to improve the personalized experience; the combination of Mel spectrum and classification model can improve the recognition accuracy of the wake-up word; the frame segmentation and sliding window technology realizes the real-time analysis of the audio stream and improves the response speed of the system; it can adapt to the wake-up words of different users and meet various application scenarios.
[0087] S130: If the voice information is a wake-up word, continue to obtain the user's voice.
[0088] When the voice information is a wake-up word, the microphone starts recording the user's voice and performs subsequent processing.
[0089] S140: Convert the user's voice into text to obtain a conversion result.
[0090] In this embodiment, the conversion result refers to the text information corresponding to the user's voice.
[0091] Specifically, the user speech is converted into text through the SpeechRecognition interface in the Web Speech API to obtain a conversion result.
[0092] The implementation steps of the SpeechRecognition interface are as follows:
[0093] First, perform browser compatibility processing. Specifically, use prefixed attributes to ensure that the speech recognition function of different browsers is supported. For example, use webkitSpeechRecognition to support WebKit-based browsers (such as Chrome) to improve the cross-browser compatibility of the application, ensure that users can enjoy the speech recognition function in different browsers, and enhance the user experience. Next, initialize the SpeechRecognition object. Specifically, create a SpeechRecognition instance and set the properties: lang: set the recognition language to Chinese (zh-CN); continuous: set to false so that the recognition ends after the user stops speaking; interimResults: set to true to obtain real-time intermediate results; flexibly configure recognition parameters to meet different application requirements, improve recognition accuracy and real-time performance, and enable users to interact with the system more naturally. Next, process the recognition results. Specifically, call the start method to start speech recognition. The system will request access to the microphone and start recording speech. The recognition results are obtained through the result event callback; users can quickly see the results of speech-to-text conversion, which enhances interactivity and is particularly suitable for application scenarios that require real-time feedback, such as online customer service or voice assistants. Then, to improve recognition accuracy and speed, use the grammars attribute to define specific grammar rules and vocabulary to limit the possible options for speech recognition. For example:
[0094] javascript
[0095] grammar='#JSGF V1.0;grammar commands;public <command> =(Highlight|Jump|Switch|Display);';
[0096] By explicitly specifying possible phrases, the speech recognition system is forced to make fewer guesses, significantly improving recognition accuracy and speed, especially in noisy environments.
[0097] Finally, close the speech recognition tool. Specifically, after the interactive dialogue ends, call the stop method to end the speech recognition so as to prepare for the next wake-up. Effectively manage the life cycle of speech recognition, avoid unnecessary resource occupation and delay, and provide users with a smoother experience.
[0098] Through the implementation of the above steps, the SpeechRecognition interface can provide users with an accurate and efficient speech recognition experience. Whether used in an application or interacting on a web page, these settings greatly improve the user's interactive experience, especially for scenarios that require natural language input, such as smart assistants, educational tools, and barrier-free applications.
[0099] S150: Identify user intention and corresponding parameters for the conversion result.
[0100] In this embodiment, the parameters include device ID or status, as well as other relevant digital twin API parameters such as location information, flight time, etc.
[0101] Specifically, the conversion result is input into a recognition model to identify the user's intention and determine the corresponding parameters.
[0102] The intent recognition model includes an intent classification model and a parameter extraction model, wherein the intent classification model is obtained by training a pre-trained model for natural language processing using a number of natural language commands covering common scenarios of all digital twin applications and corresponding SDK APIs as a sample set, and the parameter extraction model is obtained by training a named entity recognition model or a custom rule using a mapping relationship constructed by a number of natural language commands covering common scenarios of all digital twin applications and corresponding SDK APIs as a sample set.
[0103] Specifically, first, collect and annotate data: Build a comprehensive command library by collecting natural language commands that users may enter. For example, collect instructions such as "check device 123 status" and annotate these commands as corresponding SDK API calls, such as sdk.get_device_status(device_id=123). Each sample must clearly indicate the relationship between the input and the API so that the training model can understand it.
[0104] The system can cover a variety of user input situations, ensuring the diversity and comprehensiveness of model training and improving accuracy. Data annotation provides a clear learning goal for the model and enhances the effectiveness of subsequent processing.
[0105] Next, construct intent and parameter mapping to map natural language commands to SDK APIs, such as mapping "check device status" to the get_device_status intent. Through natural language processing, extract the parameters required for API calls, such as device ID, status, etc.
[0106] This clear mapping enables the system to quickly understand user intent and necessary parameters, improving response speed and accuracy.
[0107] Next, select a suitable model for intent classification and parameter extraction. For example, use a pre-trained model such as BERT or GPT for fine-tuning to better suit specific scenarios. For parameter extraction, you can use a named entity recognition (NER) model or custom rules.
[0108] By leveraging the powerful capabilities of pre-trained models, high-accuracy intent classification and parameter extraction can be achieved more quickly, reducing the time and cost of training from scratch.
[0109] Finally, in the model training phase, the intent classification model takes the user's natural language text as input and outputs the corresponding intent label. At the same time, the parameter extraction model will identify and extract the parameters in the command (such as device ID, status) based on the training data.
[0110] Through training, the model can efficiently identify user intentions and extract necessary information, thereby improving the overall interactive experience and enabling users to interact with the digital twin system in a natural language manner.
[0111] Through the above steps, the natural language processing module can efficiently and accurately identify user intentions and call the corresponding digital twin SDK API. The optimization of this process not only improves the response speed and accuracy of the system, but also enhances the user experience, making the digital twin application more interactive and practical.
[0112] S160. Call the relevant digital twin rendering engine SDK API according to the user intention and corresponding parameters to render the digital twin.
[0113] Specifically, fully trained intent classification models and parameter extraction models are packaged and ready to be deployed to the server. These models can understand and parse the user's natural language input, identify their intent, and extract the required parameter information.
[0114] Configure the necessary environment on the server, including deep learning frameworks (such as TensorFlow or PyTorch), dependent libraries, and API interfaces. Ensure that the server has sufficient computing resources to support real-time request processing.
[0115] Create an API interface that accepts the user's speech recognition results as input. The interface's request parameter format can be JSON, containing the text entered by the user. After the interface receives the user's speech recognition results, it first performs preprocessing, such as removing extra spaces or punctuation marks, to ensure the standardization of the input.
[0116] The preprocessed text is input into the intent classification model, which analyzes and returns the recognized user intent. For example, the user inputs "check the status of device 123", and the intent classification model recognizes that this is a request to "check the status of the device", and the SDK API corresponding to this request is the get_device_status method. At the same time, the input is passed to the parameter extraction model to identify the parameters related to the intent. In this example, the model extracts the device ID "123".
[0117] Based on the classification results and extracted parameters, a response object is constructed, which contains the identified intent and related parameters. This information will be sent back to the front end. After receiving the response, the front-end application parses the required digital twin interface and parameters. For example, the interface may be get_device_status, and the parameter is device_id=123.
[0118] The front end calls the corresponding digital twin SDK interface and passes the extracted parameters. This process can be asynchronous to ensure a smooth user experience. After the digital twin system processes the request, it returns the result to the front end, which then displays the result to the user. For example, the current status information of the device "123" is displayed.
[0119] After the entire process is completed, the user can see the result of his request, and this round of voice interaction ends. The system is ready to receive the next user's request and continue to provide services.
[0120] Through the above process, the user's voice commands are effectively converted into digital twin system operations, improving interaction efficiency and user experience. The deployment of the model, the design of the interface, and the coordination of the front-end and back-end make voice interaction more natural and efficient.
[0121] If the voice information is not a wake-up word, the step S110 is executed.
[0122] In order to avoid performance loss caused by the system continuously monitoring the user's voice, a wake-up word function has been added. Only when the user says the wake-up word (the default is "hi Xiaoyuan"), the voice recognition tool will be activated for subsequent voice recognition and operations. After the system wakes up, the voice recognition converts the user's voice command into text and supports multi-language output. Natural language processing (NLP) needs to be pre-trained first, and models such as BERT or GPT are used for fine-tuning to recognize the commands and intentions of digital twin services. According to the results of voice recognition, the user's intention is analyzed and the corresponding front-end operation is triggered to complete the interaction. This method effectively solves the shortcomings of manual operation: speeds up user learning, simplifies complex instructions, provides barrier-free operation, meets the needs of users with physical disabilities, improves the intelligent experience and user satisfaction, and makes the application more competitive.
[0123] The above-mentioned digital twin Web application control method obtains the user's voice information through the microphone and determines whether it is a wake-up word. If it is a wake-up word, it continues to obtain the user's voice and converts it into text. According to the conversion result, the user's intention is identified, and the SDK API of the digital twin rendering engine is called for rendering; the digital twin Web application can be operated more intuitively without memorizing cumbersome instructions, thereby improving the practicality of the application.
[0124] Corresponding to the above digital twin Web application control method, the present invention also provides a digital twin Web application control device. The digital twin Web application control device includes a unit for executing the above digital twin Web application control method, and the device can be configured in a server. Specifically, the digital twin Web application control device includes a voice acquisition unit, a judgment unit, a continuous acquisition unit, a conversion unit, a recognition unit, and a calling unit.
[0125] A voice acquisition unit is used to acquire voice input by a microphone to obtain voice information; a judgment unit is used to judge whether the voice information is a wake-up word; a continuous acquisition unit is used to continuously acquire user voice if the voice information is a wake-up word; a conversion unit is used to convert the user voice into text to obtain a conversion result; an identification unit is used to identify user intent and corresponding parameters for the conversion result; a calling unit is used to call the relevant digital twin rendering engine SDK API according to the user intent and the corresponding parameters to perform digital twin rendering.
[0126] In one embodiment, the determination unit includes a framing subunit, a conversion subunit, an input subunit, and a determination subunit.
[0127] The framing subunit is used to perform framing processing on the voice information to obtain a framing result; the conversion subunit is used to use the Mel spectrum to convert the frequency characteristics of each frame in the framing result into a Mel frequency scale to obtain a frequency conversion result; the input subunit is used to input the frequency conversion result into a classification model to obtain a recognition result; the determination subunit is used to determine whether the voice information is a wake-up word based on the recognition result.
[0128] In one embodiment, the input subunit is used to apply a sliding window on the audio stream, analyze the audio block corresponding to the frequency conversion result frame by frame, and input the audio block into the classification model for hot word detection to obtain a recognition result.
[0129] In one embodiment, the conversion unit is used to convert the user speech into text through the SpeechRecognition interface in the Web Speech API to obtain a conversion result.
[0130] In one embodiment, the recognition unit is used to input the conversion result into a recognition model to recognize the user's intention and determine the corresponding parameters.
[0131] It should be noted that technicians in the relevant field can clearly understand that the specific implementation process of the above-mentioned digital twin Web application control device and each unit can refer to the corresponding description in the aforementioned method embodiment. For the convenience and conciseness of the description, it will not be repeated here.
[0132] The above-mentioned digital twin Web application control device can be implemented in the form of a computer program, which can be run on a computer device.
[0133] The computer device may be a server, wherein the server may be an independent server or a server cluster composed of multiple servers.
[0134] The computer device includes a processor, a memory and a network interface connected via a system bus, wherein the memory may include a non-volatile storage medium and an internal memory.
[0135] The non-volatile storage medium can store an operating system and a computer program. The computer program includes program instructions, and when the program instructions are executed, the processor can execute a digital twin Web application control method.
[0136] The processor is used to provide computing and control capabilities to support the operation of the entire computer equipment.
[0137] The internal memory provides an environment for the operation of a computer program in a non-volatile storage medium. When the computer program is executed by a processor, the processor can execute a digital twin Web application control method.
[0138] The network interface is used to communicate with other devices over the network. It can be understood by those skilled in the art that the structure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0139] The processor is used to run a computer program stored in the memory to implement the following steps:
[0140] Acquire voice input by a microphone to obtain voice information; determine whether the voice information is a wake-up word; if the voice information is a wake-up word, continue to acquire user voice; convert the user voice into text to obtain a conversion result; identify user intent and corresponding parameters for the conversion result; call the relevant digital twin rendering engine SDK API according to the user intent and the corresponding parameters to perform digital twin rendering.
[0141] In one embodiment, when the processor implements the step of determining whether the voice information is a wake-up word, the processor specifically implements the following steps:
[0142] The voice information is framed to obtain a frame result; the frequency feature of each frame in the frame result is converted into a Mel frequency scale using a Mel spectrum to obtain a frequency conversion result; the frequency conversion result is input into a classification model to obtain a recognition result; and whether the voice information is a wake-up word is determined according to the recognition result.
[0143] In one embodiment, when the processor implements the step of inputting the frequency conversion result into the classification model to obtain the recognition result, the processor specifically implements the following steps:
[0144] A sliding window is applied to the audio stream, and the audio blocks corresponding to the frequency conversion results are analyzed frame by frame. The audio blocks are input into the classification model for hot word detection to obtain recognition results.
[0145] In one embodiment, when the processor implements the step of converting the user voice into text to obtain a conversion result, the processor specifically implements the following steps:
[0146] The user speech is converted into text through the SpeechRecognition interface in the Web Speech API to obtain a conversion result.
[0147] In one embodiment, when the processor implements the step of identifying the user intention and the corresponding parameter of the conversion result, the processor specifically implements the following steps:
[0148] The conversion result is input into the recognition model to recognize the user's intention and determine the corresponding parameters.
[0149] Among them, the intent recognition model includes an intent classification model and a parameter extraction model, wherein the intent classification model is obtained by training a pre-trained model of natural language processing using a number of natural language commands covering common scenarios of all digital twin applications and the corresponding SDK API as a sample set, and the parameter extraction model is obtained by training a named entity recognition model or a custom rule using a mapping relationship constructed by a number of natural language commands covering common scenarios of all digital twin applications and the corresponding SDK API as a sample set.
[0150] The parameters include device ID or status and related digital twin API parameters.
[0151] It should be understood that in the embodiments of the present application, the processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0152] It can be understood by those skilled in the art that all or part of the processes in the method for implementing the above embodiment can be completed by instructing the relevant hardware through a computer program. The computer program includes program instructions, and the computer program can be stored in a storage medium, which is a computer-readable storage medium. The program instructions are executed by at least one processor in the computer system to implement the process steps of the embodiment of the above method.
[0153] Therefore, the present invention also provides a storage medium. The storage medium may be a computer-readable storage medium. The storage medium stores a computer program, wherein when the computer program is executed by a processor, the processor executes the following steps:
[0154] Acquire voice input by a microphone to obtain voice information; determine whether the voice information is a wake-up word; if the voice information is a wake-up word, continue to acquire user voice; convert the user voice into text to obtain a conversion result; identify user intent and corresponding parameters for the conversion result; call the relevant digital twin rendering engine SDK API according to the user intent and the corresponding parameters to perform digital twin rendering.
[0155] In one embodiment, when the processor executes the computer program to implement the step of determining whether the voice information is a wake-up word, the processor specifically implements the following steps:
[0156] The voice information is framed to obtain a frame result; the frequency feature of each frame in the frame result is converted into a Mel frequency scale using a Mel spectrum to obtain a frequency conversion result; the frequency conversion result is input into a classification model to obtain a recognition result; and whether the voice information is a wake-up word is determined according to the recognition result.
[0157] In one embodiment, when the processor executes the computer program to implement the step of inputting the frequency conversion result into a classification model to obtain a recognition result, the processor specifically implements the following steps:
[0158] A sliding window is applied to the audio stream, and the audio blocks corresponding to the frequency conversion results are analyzed frame by frame. The audio blocks are input into the classification model for hot word detection to obtain recognition results.
[0159] In one embodiment, when the processor executes the computer program to implement the step of converting the user speech into text to obtain a conversion result, the processor specifically implements the following steps:
[0160] The user speech is converted into text through the SpeechRecognition interface in the Web Speech API to obtain a conversion result.
[0161] In one embodiment, when the processor executes the computer program to implement the step of identifying the user intention and the corresponding parameter of the conversion result, the processor specifically implements the following steps:
[0162] The conversion result is input into the recognition model to recognize the user's intention and determine the corresponding parameters.
[0163] Among them, the intent recognition model includes an intent classification model and a parameter extraction model, wherein the intent classification model is obtained by training a pre-trained model of natural language processing using a number of natural language commands covering common scenarios of all digital twin applications and the corresponding SDK API as a sample set, and the parameter extraction model is obtained by training a named entity recognition model or a custom rule using a mapping relationship constructed by a number of natural language commands covering common scenarios of all digital twin applications and the corresponding SDK API as a sample set.
[0164] The parameters include device ID or status and related digital twin API parameters.
[0165] The storage medium may be a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disk, etc., which are computer-readable storage media that can store program codes.
[0166] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0167] In the several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of each unit is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.
[0168] The steps in the method of the embodiment of the present invention can be adjusted in order, combined and deleted according to actual needs. The units in the device of the embodiment of the present invention can be combined, divided and deleted according to actual needs. In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0169] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, terminal, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention.
[0170] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed by the present invention, and these modifications or replacements should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be based on the protection scope of the claims.
Claims
1. A digital twin Web application control method, characterized in that: include: Acquire the voice input by the microphone to obtain voice information; Determining whether the voice information is a wake-up word; If the voice information is a wake-up word, continue to obtain the user's voice; Convert the user's speech into text to obtain a conversion result; Identifying user intent and corresponding parameters for the conversion result; The relevant digital twin rendering engine SDK API is called according to the user intention and corresponding parameters to render the digital twin.
2. The digital twin Web application control method according to claim 1, characterized in that: The determining whether the voice information is a wake-up word includes: Performing frame processing on the voice information to obtain a frame result; The frequency feature of each frame in the framing result is converted into a Mel frequency scale using the Mel spectrum to obtain a frequency conversion result; Inputting the frequency conversion result into a classification model to obtain a recognition result; Determine whether the voice information is a wake-up word according to the recognition result.
3. The digital twin Web application control method according to claim 2 is characterized in that: The step of inputting the frequency conversion result into a classification model to obtain a recognition result includes: A sliding window is applied to the audio stream, and the audio blocks corresponding to the frequency conversion results are analyzed frame by frame. The audio blocks are input into the classification model for hot word detection to obtain recognition results.
4. The digital twin Web application control method according to claim 1, characterized in that: The converting the user speech into text to obtain a conversion result includes: The user speech is converted into text through the SpeechRecognition interface in the Web Speech API to obtain a conversion result.
5. The digital twin Web application control method according to claim 1, characterized in that: The identifying the user intention and the corresponding parameters of the conversion result includes: The conversion result is input into the recognition model to recognize the user's intention and determine the corresponding parameters.
6. The digital twin Web application control method according to claim 5, characterized in that: The intent recognition model includes an intent classification model and a parameter extraction model, wherein the intent classification model is obtained by training a pre-trained model for natural language processing using a number of natural language commands covering common scenarios of all digital twin applications and corresponding SDK APIs as a sample set, and the parameter extraction model is obtained by training a named entity recognition model or a custom rule using a mapping relationship constructed by a number of natural language commands covering common scenarios of all digital twin applications and corresponding SDK APIs as a sample set.
7. The digital twin Web application control method according to claim 6, characterized in that: The parameters include device ID or status and related digital twin API parameters.
8. A digital twin web application control device, characterized in that: include: A voice acquisition unit, used to acquire the voice input by the microphone to obtain voice information; A judging unit, used to judge whether the voice information is a wake-up word; A continuous acquisition unit, configured to continuously acquire user voice if the voice information is a wake-up word; A conversion unit, used for converting the user speech into text to obtain a conversion result; An identification unit, used to identify the user intention and corresponding parameters of the conversion result; The calling unit is used to call the relevant digital twin rendering engine SDK API according to the user intention and corresponding parameters to render the digital twin.
9. A computer device, characterized in that: The computer device comprises a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.
10. A storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.