Page control method and device based on voice recognition, equipment and medium

By receiving and processing voice data transmitted by remote control terminals, combining preset command sets and cloud servers, large-screen page control without fixed wake-up instructions is realized, solving the problem of limited user experience in the existing technology, and improving the intelligence and accuracy of voice control.

CN120071925APending Publication Date: 2025-05-30BEIJING NORTH STAR DIGITAL REMOTE SENSING TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510187881.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing voice control system requires users to use fixed wake-up instructions, which limits the user experience and makes it difficult to achieve intelligent voice control without fixed wake-up instructions.

Method used

By receiving the target voice data transmitted by the remote control terminal, obtaining the target execution command based on the preset command set, and forwarding it to the large-screen terminal through the cloud server, realizing large-screen page control without a fixed wake-up command.

Benefits of technology

It realizes intelligent voice control of large-screen pages without fixed wake-up instructions, improves user experience, and enhances processing and confidence verification through voice features, improving the accuracy and computing efficiency of voice recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071925A_ABST
    Figure CN120071925A_ABST
Patent Text Reader

Abstract

The invention relates to a page control method and device based on voice recognition, equipment and a medium. The method comprises the following steps: receiving target voice data transmitted by a remote control terminal; based on a preset instruction set, obtaining a target execution command corresponding to the target voice data; sending the target execution command to a remote control terminal, so that the remote control terminal converts the target execution command into a real-time message and sends the real-time message to a real-time communication server; and the real-time communication server forwards the target execution command to the large-screen terminal with the same communication identifier as the unique identifier of the user, and the large-screen terminal executes the target execution command to complete page control. According to the invention, after voice input is carried out through the remote control terminal, the large-screen terminal can realize corresponding page control operation through processing of the cloud server, and a fixed wake-up instruction is not needed, so that intelligent voice control of the large-screen page without the fixed wake-up instruction can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of page control, and particularly to a page control method, device, equipment and medium based on speech recognition. Background Art

[0002] With the popularization of smart devices and the rapid development of Internet technology, speech control technology has been increasingly widely applied in the fields of smart home, smart office, smart entertainment, etc.

[0003] Existing speech control systems usually require users to use fixed wake-up commands to activate the speech recognition function, such as "Hey, xxx" or "OK, xxx". Although this method can effectively avoid mis-triggering, it also limits the user experience because users have to remember specific wake-up words.

[0004] Therefore, there is an urgent need for a page control solution based on speech recognition that can achieve intelligent speech control without fixed wake-up commands. Summary of the Invention

[0005] In order to achieve intelligent speech control without fixed wake-up commands, the present application provides a page control method, device, equipment and medium based on speech recognition.

[0006] In a first aspect, the present application provides a page control method based on speech recognition, including:

[0007] Receiving target speech data transmitted by a remote control terminal;

[0008] Based on a preset instruction set, obtaining a target execution command corresponding to the target speech data;

[0009] Sending the target execution command to the remote control terminal, so that the remote control terminal converts the target execution command into a real-time message and sends it to a real-time communication server, the real-time message including a user unique identifier and the target execution command; the real-time communication server forwards the target execution command to a large-screen terminal having the same communication identifier as the user unique identifier, and the large-screen terminal executes the target execution command to complete page control.

[0010] The beneficial effect of the present application is that no fixed wake-up command is required. After voice input through the remote control terminal and processing by the cloud server, the large-screen terminal can implement corresponding page control operations. At the same time, the voice input content has no fixed format, and different voice texts can implement the same operation instruction, so as to achieve intelligent speech control of the large-screen page without fixed wake-up commands.

[0011] Further, before obtaining the target execution command corresponding to the target voice data based on the preset instruction set, the following steps are also included:

[0012] Based on the short-time energy corresponding to the target voice data, determine whether the target voice data is valid speech, where the valid speech is speech data with short-time energy within a preset energy range;

[0013] If the target voice data is the valid speech, perform speech feature enhancement processing on the target voice data based on the preset Mel cepstral coefficient algorithm.

[0014] The beneficial effect of adopting the above further solution is that by using the Mel cepstral coefficient algorithm to perform spectral weighting and reconstruction on the voice data, the effect of reducing noise interference is achieved. It can effectively filter out invalid speech and extract robust speech features, providing high-quality input for subsequent text conversion and instruction parsing of the ASR model, improving the speech recognition accuracy in a noisy environment, and the computational efficiency can meet the real-time requirements.

[0015] Further, the step of obtaining the target execution command corresponding to the target voice data based on the preset instruction set includes:

[0016] Based on a preset ASR model, recognize the target voice data after the speech feature enhancement processing and convert it into a target text;

[0017] Based on a preset LLM model, match the target text and the current execution command corresponding to the instruction set, and obtain the target execution command based on the current execution command, where the target execution command includes an instruction control type and instruction control parameters.

[0018] The beneficial effect of adopting the above further solution is that by converting speech to text through the ASR model and then matching the text to specific operation instructions through the LLM model, efficient and accurate voice control can be achieved.

[0019] Further, after recognizing the target voice data after the speech feature enhancement processing and converting it into a target text, the following steps are also included:

[0020] Based on the user unique identifier, obtain a confidence threshold;

[0021] If the current confidence corresponding to the target text is lower than the confidence threshold, filter the target text.

[0022] The beneficial effect of adopting the above further solution is that after the text conversion is completed, adding a confidence check to filter out target texts with low confidence reduces the incorrect response to ambiguous speech.

[0023] Further, after matching the target text and the current execution command corresponding to the instruction set based on the preset LLM model, the following steps are also included:

[0024] Based on the historical operation data corresponding to the user unique identifier and the current operation scenario, determine whether the current execution command is a reasonable command;

[0025] If the current execution command is a reasonable command, perform the operation of obtaining the target execution command based on the current execution command.

[0026] The beneficial effect of adopting the above further solution is that by combining the historical operation data corresponding to the user unique identifier and the current operation scenario, the rationality of the current execution command can be dynamically judged, ensuring that the user's operation instruction is reasonable and executable in the current operation scenario, thereby avoiding incorrect operations and improving the user experience.

[0027] Further, obtaining the target execution command based on the current execution command includes:

[0028] Based on the historical operation data corresponding to the user unique identifier and the current operation scenario, predict the user's execution requirement information;

[0029] Based on the user's execution requirement information, adjust the current execution command to obtain the target execution command.

[0030] The beneficial effect of adopting the above further solution is that by combining the user's historical operation data and the current operation scenario, the user's execution requirement can be predicted, and the current execution command can be dynamically adjusted to generate a target execution command that better conforms to the user's intention, providing stronger intelligent capabilities and user experience for voice control.

[0031] Further, obtaining the target execution command based on the current execution command includes:

[0032] Based on the emotional characteristics corresponding to the target voice data, identify the user's emotional state information;

[0033] Based on the emotional state information, adjust the current execution command to obtain the target execution command.

[0034] The beneficial effect of adopting the above further solution is that a full-link closed loop of emotional state - response strategy - long-term optimization is established, which can provide more personalized services.

[0035] In a second aspect, the present application provides a page control device based on voice recognition, including:

[0036] A data receiving module, configured to receive the target voice data transmitted by the remote control terminal;

[0037] An identification command module, configured to obtain a target execution command corresponding to the target voice data based on a preset instruction set;

[0038] A data sending module, configured to send the target execution command to the remote control terminal, so that the remote control terminal converts the target execution command into a real-time message and sends it to a real-time communication server, where the real-time message includes a user unique identifier and the target execution command; the real-time communication server forwards the target execution command to a large-screen terminal having a communication identifier the same as the user unique identifier, and the large-screen terminal executes the target execution command to complete page control.

[0039] In a third aspect, the present application provides an electronic device, including a processor and a memory, where the processor is coupled to the memory;

[0040] The processor is configured to execute a computer program stored in the memory, so that the electronic device executes the method according to any one of the first aspect.

[0041] In a fourth aspect, the present application provides a computer-readable storage medium, including a computer program or instruction, when the computer program or instruction runs on a computer, the computer is caused to execute the method according to any one of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 It is a flowchart of a page control method based on voice recognition according to an embodiment of the present application;

[0043] Figure 2 It is a structural block diagram of a page control method based on voice recognition according to an embodiment of the present application;

[0044] Figure 3 It is a structural block diagram of a page control device based on voice recognition according to an embodiment of the present application;

[0045] Figure 4 It is a structural block diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0046] The following further describes the present application in detail with reference to the accompanying drawings.

[0047] An embodiment of the present application provides a page control method based on voice recognition. This method can be executed by a device, and the device can be a server or a terminal device. Among them, the server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The terminal device can be a desktop computer, etc., but is not limited thereto.

[0048] AsFigure 1 and Figure 2 As shown in Figure 2 , a page control method based on speech recognition takes an electronic device as the execution entity. The main process of the method is described as follows (Steps S101 to S103):

[0049] Step S101: Receive the target voice data transmitted by the remote control terminal.

[0050] In this embodiment, the remote control terminal can be a self-built remote control APP, and the electronic device can be a cloud server. The remote control APP is communicatively connected to the cloud server. The remote control APP provides a voice input function for the user. The user operates through the remote control APP to record a voice file to generate voice data, and the remote control APP uploads the voice data to the cloud server for intelligent processing.

[0051] To ensure the security of uploading the voice data to the cloud server, the uploaded voice data can be encrypted against tampering. The specific encryption process can include: sorting the file information (file name, file hash value) corresponding to the voice data in order to generate a unique string; performing MD5 encryption on the generated unique string to generate a key as a transmission parameter and passing it into the cloud server, and the cloud server generates the corresponding key according to the same rule for front-end and back-end anti-tampering comparison.

[0052] Step S102: Based on a preset instruction set, obtain the target execution command corresponding to the target voice data.

[0053] In this embodiment, the target execution command can include an instruction control type and an instruction control parameter. The instruction control type can include page jump, map positioning, opening a camera, etc. The instruction control parameter is the specific control parameter corresponding to the instruction control type. For example, the target execution command can be "turn on camera 001", the instruction control type is "open camera", and the instruction control parameter is the camera with the control number "001".

[0054] According to the instruction set, the target instruction can be matched from the target voice data input by the user. For example, if the target voice data input by the user is "locate to Wuhan" or "go to Wuhan", the instruction control type of "map positioning" and the instruction control parameter of locating to Wuhan can be matched.

[0055] Step S103: Send the target execution command to the remote control terminal, so that the remote control terminal converts the target execution command into a real-time message and sends it to the real-time communication server. The real-time message includes the user unique identifier and the target execution command; the real-time communication server forwards the target execution command to a large screen terminal with the same communication identifier as the user unique identifier, and the large screen terminal executes the target execution command to complete page control.

[0056] In this embodiment, the real-time message is the WebSocket message, and the real-time communication server is the WebSocket server. As a general service, the cloud server can serve multiple remote control terminals externally, providing voice and command conversion functions, which can be called by each system to implement voice operation commands.

[0057] By using the remote control APP as an information transfer point, when calling the cloud server service upward, the user's unique identifier is used to distinguish and push voice data and receive the recognized target execution command. When pushing the WebSocket message downward, the user's unique identifier is used as the communication identifier of the message, so as to realize the page control operation of the large screen terminal with the same user unique identifier.

[0058] After receiving the target execution command, the WebSocket server parses and converts the target execution command, and sends the user unique identifier and the target execution command pushed by the remote control APP to the large screen terminal with the same user unique identifier.

[0059] On the large screen terminal, the large screen terminal parses the WebSocket message, matches the target execution command with the pre-defined operation steps, and executes relevant operations, such as turning on the camera, page jumping or playing a video, etc.

[0060] In this embodiment, no fixed wake-up command is required. After voice input through the remote control APP, through the processing of the cloud server, the large screen terminal can realize the corresponding page control operation. At the same time, the voice input content has no fixed format, and different voice texts can realize the same operation command. For example, jumping to page 1, opening page 1, page 1, etc. can all realize page 1 for switching the large screen page, so as to realize the intelligent voice control of the large screen page without a fixed wake-up command.

[0061] In this embodiment, the same remote control APP can control multiple large screen terminals at the same time. The remote control APP and the large screen terminal can be bound according to the user's unique identifier. When multiple large screen terminals log in to the same account, the same execution command can be sent to each large screen terminal with the same user identifier through the WebSocket server to realize the simultaneous operation of multiple large screen terminals.

[0062] The remote control APP can support two page control modes: button control mode and voice control mode. Among them, the button control mode includes: when clicking the page control button on the remote control APP, directly sending real-time messages to the backend websocket server. The websocket server processes the real-time messages and forwards them to the large-screen terminal corresponding to the user's unique identifier. After receiving the instruction, the large-screen terminal performs associated operations according to the preset instructions, and automatically performs operations such as page jumping, video playing, and camera opening;

[0063] The voice control mode includes: after the user describes the operation to be performed through voice, the remote control APP calls the cloud server to convert the operation described by the user into a corresponding target execution command. The cloud server then returns the target execution command to the remote control APP, and the remote control APP sends a websocket message to the websocket server according to the target execution command.

[0064] In this embodiment, before step S102, it further includes: based on the short-time energy corresponding to the target voice data, determining whether the target voice data is valid voice, where the valid voice is voice data whose short-time energy is within a preset energy range; if the target voice data is the valid voice, then based on the preset Mel cepstral coefficient algorithm, performing voice feature enhancement processing on the target voice data.

[0065] Obtaining the short-time energy corresponding to the target voice data specifically includes: frame segmentation of the voice data, dividing the target voice data into short-time frames, and the duration of each frame can be 20 - 40 ms (such as 25 ms), and the frame shift is 10 ms; based on the preset short-time energy calculation formula, calculating the short-time energy (Short-Time Energy, STE) for each frame of signal. Exemplarily, the sampling rate is 16 kHz, the frame length is 400 samples (25 ms × 16,000 Hz), and the frame shift is 160 samples (10 ms × 16,000 Hz).

[0066] The preset energy threshold range can be [E min ,E max , the low energy threshold E min is used to filter out silence or weak noise, and the high energy threshold E max is used to filter out sudden high-energy noise (such as hitting the microphone).

[0067] If the energy E n of a certain frame is within [E min ,E maxIf it is within the range, the frame can be marked as a valid frame; if the number of consecutive valid frames exceeds a preset frame number threshold (such as 5 frames), the target voice data can be determined to be valid voice, and the next step of voice feature enhancement processing can be entered; if the target voice data is determined to be invalid voice, it can be directly discarded and the remote control terminal can be notified to re-enter. By calculating the short-time energy of the voice data and filtering out invalid voices (such as silence, low-energy noise or sudden high-energy interference), the validity of the input voice data is ensured and the recognition accuracy is improved.

[0068] The voice feature enhancement processing based on Mel Frequency Cepstral Coefficients (MFCC) specifically includes: performing high-frequency enhancement on the valid voice segment to compensate for the high-frequency attenuation of the voice signal; re-framing the pre-emphasized voice again (with the same parameters as the short-time energy calculation); applying a Hamming Window to each frame of the signal to reduce spectral leakage; performing a Fast Fourier Transform (FFT) on each frame of the signal to obtain the amplitude spectrum; converting the linear frequency scale to the Mel Scale to simulate the human ear's auditory characteristics; designing a set of triangular filters (which can be 20 - 40), covering from 0 Hz to the Nyquist frequency (such as 8 kHz); taking the logarithm of the energy output by each filter to enhance the distinguishability of low-energy components; performing Discrete Cosine Transform (DCT) on the logarithmic filter bank energy to extract MFCC coefficients (which can take the first 12 - 13 dimensions); calculating the first-order difference (Delta) and the second-order difference (Delta-Delta) to characterize the temporal variation of the features.

[0069] By using the Mel Frequency Cepstral Coefficient algorithm to weight and reconstruct the spectrum of the voice data, the effect of reducing noise interference is achieved. Through the combination of short-time energy judgment and MFCC feature enhancement, invalid voices can be effectively filtered out and robust voice features can be extracted, providing high-quality input for the subsequent text conversion and instruction parsing of the ASR model, improving the voice recognition accuracy in a noisy environment, and the calculation efficiency can meet the real-time requirements.

[0070] In this embodiment, at the beginning stage of voice input (such as the first 1 second), the default threshold can be used as the initial energy threshold. Exemplarily, E min can be 0.01, E max can be 0.5. During the voice processing, continuously detect low-energy frames (such as frames with energy lower than the current noise energy estimate), and use the exponential smoothing method to update the noise energy estimate. The noise energy estimate can be expressed as:

[0071] E noise = α 1 E noise +(1 - α 1 )E current

[0072] where, E noiseRepresents the current noise energy estimate, E current Represents the energy of the current frame, α 1 Represents the smoothing coefficient (which can take values from 0.9 - 0.99).

[0073] Based on the noise energy estimate, the energy threshold range of the effective speech can be dynamically adjusted:

[0074] E min = k1E noise

[0075] E max = k2E noise

[0076] Among them, both k1 and k2 can be empirical coefficients. Exemplarily, k1 can be 2 and k2 can be 10.

[0077] By updating the noise energy estimate in real time through the exponential smoothing method, it can adapt to the changes in environmental noise. Dynamically adjusting the energy threshold according to the noise level can avoid the failure problem of the fixed threshold in a noisy environment. Through the adaptive energy threshold technology, it can dynamically adjust the judgment standard of speech validity according to the environmental noise level, significantly improving the robustness in a noisy environment. By combining real-time noise estimation and dynamic threshold adjustment, it provides a stronger environmental adaptability for the page control scheme based on speech recognition.

[0078] In this embodiment, step S102 includes: based on a preset ASR model, identifying the target speech data after the speech feature enhancement process and converting it into a target text; based on a preset LLM model, matching the target text and the current execution command corresponding to the instruction set, and obtaining the target execution command based on the current execution command, where the target execution command includes an instruction control type and instruction control parameters.

[0079] The ASR model, that is, the Automatic Speech Recognition model, can use open-source asr resources to identify and convert the speech data input by the user into text. The ASR model can be a framework and model with strong Chinese recognition ability, high recognition accuracy, fast recognition speed, offline deployability, and low resource occupancy.

[0080] Input the target speech data after the speech feature enhancement process into the ASR model, and the ASR model can output the recognized text. Exemplarily: Input: Speech "Open camera 001"; Output: Text "Open camera 001".

[0081] The cloud server includes an AI module, which can be an LLM model, i.e., a Large Language Model. The LLM model can match the recognized text to specific operation instructions and their parameters.

[0082] The LLM model can also provide a general registration interface externally, and can dynamically create operation instructions. The dynamic creation process can specifically include: the front end sends the system identifier, voice recognition features, and operation instructions to the cloud server through the registration interface, and the cloud server performs voice and instruction binding to establish a mapping relationship between the text and the instructions.

[0083] For example, the front end can register instructions such as "jump page", "map positioning", "turn on the camera", etc. The front end sends the system identifier, voice recognition features (such as keywords), and operation instructions to the cloud server through the registration interface. The cloud server binds this information to form a mapping relationship of "voice text → operation instruction", providing a basis for subsequent instruction execution.

[0084] Converting the voice to text through the ASR model and then matching the text to specific operation instructions through the LLM model can achieve efficient and accurate voice control. Combining dynamic instruction registration and real-time communication technology makes the page control method based on voice recognition highly flexible and scalable.

[0085] In this embodiment, after recognizing the target voice data after the voice feature enhancement process and converting it into a target text, it further includes: obtaining a confidence threshold based on the user unique identifier; if the current confidence corresponding to the target text is lower than the confidence threshold, then filtering the target text.

[0086] After the ASR model recognizes, the confidence (Confidence Score) of the recognition result can be filtered by a threshold. If the confidence is lower than a preset value (such as 0.5), it can be determined as low-quality voice or mis-triggered, and the subsequent instruction matching process is terminated.

[0087] In this embodiment, the confidence threshold can be dynamically adjusted in combination with the user unique identifier. Set the initial threshold according to the user unique identifier. When the user corresponding to the user unique identifier is a new user, the confidence threshold can be defaulted to a relatively low value, such as 0.3; when the user corresponding to the user unique identifier is an old user, the confidence threshold can be set to a relatively high value, such as 0.7, according to historical data. When the confidence threshold is low, more fuzzy voices can be tolerated to avoid frequently asking the user to re-enter, thereby enhancing the initial experience of new users; when the confidence threshold is high, incorrect responses can be reduced and the system accuracy can be improved, thereby enhancing the usage efficiency of old users.

[0088] After the initial threshold is set, the confidence threshold can be gradually adjusted according to the preset threshold adjustment formula and user feedback data. The confidence threshold can be set with upper and lower limits to prevent the threshold from being adjusted too high or too low. For example, the lowest value of the confidence threshold can be 0.2, and the highest value of the confidence threshold can be 0.8.

[0089] If the current recognition result output by the ASR model is accepted by the user (such as successful execution), the threshold can be increased. The threshold adjustment formula can be expressed as:

[0090] n th =c th +α 2 ×(1-c th )

[0091] If the current recognition result output by the ASR model is rejected by the user (such as execution failure), the threshold can be lowered. The threshold adjustment formula can be expressed as:

[0092] n th =c th -α 2 ×(1-c th )

[0093] Among them, n th represents the new confidence threshold, c th represents the current confidence threshold, α 2 Indicates the threshold adjustment factor.

[0094] In this embodiment, the threshold adjustment coefficient α can also be dynamically adjusted according to the data analysis results. 2 For example, if the success rate of the recognition result being accepted by the user continues to be higher than 90%, then increase α 2 , which can speed up the improvement of the confidence threshold.

[0095] In this embodiment, after matching the target text and the current execution command corresponding to the instruction set based on the preset LLM model, it also includes: judging whether the current execution command is a reasonable command based on the historical operation data corresponding to the user unique identifier and the current operation scenario; if the current execution command is a reasonable command, executing the operation of obtaining the target execution command based on the current execution command.

[0096] User historical operations may include the command control type, command control parameters, timestamp, and execution results of each user operation, and the current operation scenario may include the user's current operation scenario information, such as device status, page location, and environmental data.

[0097] Verify the rationality of the command execution by combining the user's historical operations and the current operation scenario, including:

[0098] Analyze the user's historical operation habits, extract the frequently executed commands and instruction control parameters. Exemplarily, for user A, the frequently executed command is "Turn on camera 001", and the instruction control parameter is "001".

[0099] Obtain information such as the current device status, page position, and environmental data. Exemplarily, the device status is that camera 001 is closed, the page position is the home page, and the environmental data is light intensity = 300 lux.

[0100] Calculate the rationality score of the current executed command according to the type matching situation, parameter validity situation, and scenario adaptability situation corresponding to the current executed command.

[0101] If the rationality score exceeds the reasonable threshold (such as 0.8), then determine that the current executed command is a reasonable command; otherwise, determine that the current executed command is an unreasonable command. If the current executed command is determined to be reasonable, continue with the subsequent operations. If the current executed command is determined to be unreasonable, return an error message and prompt the user to re-enter. Exemplarily, the error message is "This operation is not supported in the current operation scenario. Please try other instructions."

[0102] In this embodiment, calculating the rationality score of the current executed command according to the type matching situation, parameter validity situation, and scenario adaptability situation corresponding to the current executed command specifically includes:

[0103] Type matching situation: Check whether the current executed command is in the user's frequently executed command set. If it is in the user's frequently executed command set, the score for the type matching situation can be +0.5. For example, the current executed command is "Turn on camera 001", which can match the frequently executed command set.

[0104] Parameter validity situation: Check whether the instruction control parameter corresponding to the current executed command is valid in the current operation scenario. If it is valid in the current operation scenario, the score for the parameter validity situation can be +0.3. For example, the parameter "001" corresponds to a camera that exists and is operable.

[0105] Scenario adaptability situation: Check whether the current executed command is suitable for the current operation scenario. If it is valid in the current operation scenario, the score for the parameter validity situation can be +0.2. For example, the current page is the home page, which supports the operation of "Turn on the camera".

[0106] Sum up the scores corresponding to the type matching situation, parameter validity situation, and scenario adaptability situation, and calculate the rationality score corresponding to the current executed command.

[0107] In this embodiment, filtering policies are respectively set in multiple stages such as the recognition of the ASR model to achieve hierarchical filtering, gradually reducing invalid operations, so that only the execution commands that pass all verifications enter the subsequent processing flow, thereby reducing the resource consumption of invalid operations and network transmissions, and significantly reducing the impacts of mis-triggering and interference.

[0108] As an alternative implementation manner of this embodiment, obtaining the target execution command based on the current execution command includes: predicting user execution requirement information based on the historical operation data corresponding to the user unique identifier and the current operation scenario; and adjusting the current execution command based on the user execution requirement information to obtain the target execution command.

[0109] The cloud server may further include a requirement prediction model, which is a traditional machine learning model or a deep learning model for predicting user requirements. Taking the historical operation data and the current operation scenario as the input features of the requirement prediction model, the output result of the requirement prediction model is the predicted user execution requirement information.

[0110] Exemplarily, if the user's historical operation is to frequently operate "turn on camera 001" and the current operation scenario is that camera 001 is closed, when the user inputs "turn on the camera", the predicted requirement is that the user may need to "turn on camera 001", and the current execution command is adjusted according to the user execution requirement information to obtain the target execution command.

[0111] In this alternative implementation manner, user habit analysis may be performed based on the recorded historical operation data to understand the user's operation habits. For example, if the user often uses a specific function on a certain page, it can be predicted that the user may use the function again on that page. When the user inputs "open the page", the predicted requirement is that the user may need to "open the page and use the function".

[0112] In this alternative implementation manner, after predicting the instructions that the user may execute, the instruction control parameters can be automatically optimized to provide better control instructions. For example: if the instruction input by the user is "locate to Wuhan", more specific target locations in Wuhan (such as "Wuhan University" or "Wuhan Railway Station") can be automatically predicted according to the user's historical operation data and provided for the user to select; if the instruction input by the user is "turn on camera 001", the photographing state or video recording state of the camera (such as automatically turning on the night vision mode or automatically focusing) can be automatically adjusted according to the user's historical operation data and the current operation scenario.

[0113] By combining the user's historical operation data and the current operation scenario, it is possible to predict the user's execution requirements, dynamically adjust the current execution command, and generate a target execution command that better conforms to the user's intention, providing stronger intelligent capabilities and user experience for voice control.

[0114] As another alternative implementation of this embodiment, obtaining the target execution command based on the current execution command includes: identifying the emotional state information of the user based on the emotional characteristics corresponding to the target voice data; adjusting the current execution command based on the emotional state information to obtain the target execution command.

[0115] Based on speech signal processing technology, the emotional characteristics corresponding to the target voice data can be extracted. The emotional characteristics can include intonation, speech rate, energy, and spectral characteristics. Among them, the intonation can be obtained by analyzing the fundamental frequency (Pitch) change of the speech, the speech rate can be obtained by calculating the speech rate (such as the number of syllables per second), the energy can be obtained by analyzing the short-term energy of the speech, and the spectral characteristics can be obtained by extracting spectral characteristics such as Mel Frequency Cepstral Coefficients (MFCC).

[0116] The cloud server may also include an emotion classification model. The emotion classification model is a traditional machine learning model or a deep learning model for classifying emotional characteristics. The traditional machine learning model can be any one of Support Vector Machine (SVM) or Random Forest, and the deep learning model can be any one of Convolutional Neural Network (CNN) or Recurrent Neural Network (RNN).

[0117] The extracted emotional characteristics can be used as the input of the emotion classification model, and the emotion classification model outputs the emotional state information of the user. The emotional state information can include happy, angry, calm, and sad, etc.

[0118] Adjust the instruction control parameters corresponding to the current execution command according to the emotional state information. The instruction control parameters can also include response voice parameters and light parameters. The response voice parameters can include speech rate, intonation, and timbre, etc.

[0119] The cloud server stores the corresponding relationship between different emotional state information and different instruction control parameters. Exemplarily, when the user says "play music", if the emotional state information is happy, the corresponding instruction control parameter is to recommend pop music; if the emotional state information is sad, the corresponding instruction control parameter is to recommend light music or white noise. When the user says "turn on the light", if the emotional state information is calm, the corresponding instruction control parameter is the default white light; if the emotional state information is angry, the corresponding instruction control parameter is to adjust to warm yellow light.

[0120] In this alternative embodiment, the execution priority of the execution command can also be adjusted according to the emotional state information. For example, when the emotional state information is anger or sadness, complex operations (such as "turn on Camera 001 and adjust parameters") can be skipped, and simple instructions (such as "turn on the camera") can be preferentially executed. When the emotional state information is happiness, extended functions can be actively recommended (such as "Turn on Camera 001 for you. Do you need to enable a certain photo-taking mode?").

[0121] Record the current interaction process, use data analysis tools (such as Python Pandas) to count the user's emotional distribution, and the high-frequency negative emotion time periods and the correlation between specific instructions and emotions can be obtained. According to the correlation between specific instructions and emotions, the instruction matching logic can be optimized. Exemplarily, "increase the volume" is often accompanied by anger, and an appeasement response can be automatically triggered; if the user is frequently angry after 8 pm, the "night soothing mode" can be automatically enabled during this time period. The "night soothing mode" is a working mode that can provide soothing page control.

[0122] By establishing a full-link closed loop of emotional state - response strategy - long-term optimization, not only can we understand what the user "says", but also sense how the user "says", so as to provide more user-friendly services.

[0123] Based on the same technical concept, the present application also provides a page control device based on voice recognition, as Figure 3 shown. The page control device 200 based on voice recognition mainly includes:

[0124] A data receiving module 201, configured to receive target voice data transmitted by a remote control terminal;

[0125] A command recognition module 202, configured to obtain a target execution command corresponding to the target voice data based on a preset instruction set;

[0126] A data sending module 203, configured to send the target execution command to the remote control terminal, so that the remote control terminal converts the target execution command into a real-time message and sends it to a real-time communication server. The real-time message includes a user unique identifier and the target execution command; the real-time communication server forwards the target execution command to a large-screen terminal having the same communication identifier as the user unique identifier, and the large-screen terminal executes the target execution command to complete page control.

[0127] Optionally, before the command recognition module 202, it further includes:

[0128] A validity judgment module, configured to judge whether the target speech data is valid speech based on the short-time energy corresponding to the target speech data, where the valid speech is speech data with short-time energy within a preset energy interval; if the target speech data is the valid speech, perform speech feature enhancement processing on the target speech data based on a preset Mel cepstrum coefficient algorithm.

[0129] Optionally, the recognition command module 202 includes:

[0130] A recognition conversion sub-module, configured to recognize the target speech data after the speech feature enhancement processing based on a preset ASR model and convert it into a target text;

[0131] An instruction matching sub-module, configured to match the target text and the current execution command corresponding to the instruction set based on a preset LLM model, and obtain the target execution command based on the current execution command, where the target execution command includes an instruction control type and instruction control parameters.

[0132] Optionally, after the recognition conversion module, it further includes:

[0133] A confidence threshold acquisition sub-module, configured to acquire a confidence threshold based on the user unique identifier;

[0134] A filtering sub-module, configured to filter the target text when the current confidence corresponding to the target text is lower than the confidence threshold.

[0135] Optionally, after the instruction matching sub-module, it further includes:

[0136] A rationality judgment sub-module, configured to judge whether the current execution command is a reasonable command based on the historical operation data and the current operation scenario corresponding to the user unique identifier;

[0137] An execution sub-module, configured to perform the operation of obtaining the target execution command based on the current execution command when the current execution command is a reasonable command.

[0138] Optionally, the instruction matching sub-module includes:

[0139] A prediction sub-module, configured to predict user execution requirement information based on the historical operation data and the current operation scenario corresponding to the user unique identifier;

[0140] A first adjustment sub-module, configured to adjust the current execution command based on the user execution requirement information to obtain the target execution command.

[0141] Optionally, the instruction matching sub-module includes:

[0142] An emotion recognition sub-module, configured to recognize the emotional state information of the user based on the emotional features corresponding to the target voice data;

[0143] A second adjustment sub-module, configured to adjust the current execution command based on the emotional state information to obtain the target execution command.

[0144] In one example, the modules in any of the above devices may be one or more integrated circuits configured to implement the above methods. For example: one or more application specific integrated circuits (ASICs), or, one or more digital signal processors (DSPs), or, one or more field programmable gate arrays (FPGAs), or a combination of at least two of these integrated circuit forms.

[0145] Again, when the modules in the device can be implemented in the form of a processing element scheduler, the processing element may be a general-purpose processor, such as a central processing unit (CPU) or other processors that can call programs. Again, these modules may be integrated together and implemented in the form of a system-on-a-chip (SOC).

[0146] In this application, names may be assigned to various objects such as various messages / information / devices / network elements / systems / devices / actions / operations / processes / concepts, etc. It can be understood that these specific names do not constitute limitations on the relevant objects, and the assigned names may change with factors such as scenarios, contexts, or usage habits. The understanding of the technical meanings of the technical terms in this application should be mainly determined from the functions and technical effects reflected / executed in the technical solutions.

[0147] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the systems, devices, and modules described above can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0148] Those of ordinary skill in the art can realize that the modules and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0149] Based on the same inventive concept, the present application also provides an electronic device, such as Figure 4 shown, the electronic device 300 includes a processor 301 and a memory 302, and may further include one or more of an information input / output (I / O) interface 303, a communication component 304, and a communication bus 305.

[0150] Among them, the processor 301 is used to control the overall operation of the electronic device 300 to complete all or part of the steps in the above-mentioned page control method based on speech recognition; the memory 302 is used to store various types of data to support the operation of the electronic device 300. These data may include, for example, instructions for any application or method operating on the electronic device 300, and application-related data. The memory 302 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk, or one or more of them.

[0151] The I / O interface 303 provides an interface between the processor 301 and other interface modules. The above-mentioned other interface modules can be a keyboard, a mouse, buttons, etc. These buttons can be virtual buttons or physical buttons. The communication component 304 is used to test the wired or wireless communication between the electronic device 300 and other devices. Wireless communication, such as Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, or 4G, or a combination of one or more of them. Therefore, the corresponding communication component 304 may include: a Wi-Fi component, a Bluetooth component, an NFC component.

[0152] The communication bus 305 may include a path for transmitting information among the above components. The communication bus 305 may be a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, or the like. The communication bus 305 may be divided into an address bus, a data bus, a control bus, etc.

[0153] The electronic device 300 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components, and is used to execute the steps of the page control method based on speech recognition given in the above embodiments.

[0154] The electronic device 300 may include, but is not limited to, mobile terminals such as digital broadcast receivers, PDAs (Personal Digital Assistants), PMPs (Portable Multimedia Players), etc., and fixed terminals such as digital TVs, desktop computers, etc., and may also be a server, etc.

[0155] Based on the same technical concept, the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above page control method based on speech recognition are implemented.

[0156] The computer-readable storage medium may include various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.

[0157] The term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device.

[0158] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In the description of the present application, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically defined.

[0159] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0160] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.

Claims

1. A page control method based on voice recognition, characterized in that: include: Receiving target voice data transmitted by the remote control terminal; Based on a preset instruction set, obtaining a target execution command corresponding to the target voice data; The target execution command is sent to the remote control terminal so that the remote control terminal converts the target execution command into a real-time message and sends it to a real-time communication server, wherein the real-time message includes a user unique identifier and the target execution command; the real-time communication server forwards the target execution command to a large-screen terminal having a communication identifier that is the same as the user unique identifier, and the large-screen terminal executes the target execution command to complete page control.

2. A page control method based on voice recognition according to claim 1, characterized in that: Before obtaining the target execution command corresponding to the target voice data based on the preset instruction set, the method further includes: Based on the short-time energy corresponding to the target voice data, determining whether the target voice data is valid voice, wherein the valid voice is voice data whose short-time energy is within a preset energy range; If the target speech data is the valid speech, speech feature enhancement processing is performed on the target speech data based on a preset Mel-cephalometric coefficient algorithm.

3. The page control method based on voice recognition according to claim 2, characterized in that: The step of obtaining a target execution command corresponding to the target voice data based on a preset instruction set includes: Based on a preset ASR model, the target speech data after the speech feature enhancement processing is recognized and converted into a target text; Based on a preset LLM model, the target text and the current execution command corresponding to the instruction set are matched, and the target execution command is obtained based on the current execution command, where the target execution command includes an instruction control type and an instruction control parameter.

4. The page control method based on speech recognition according to claim 3 is characterized in that: After the target voice data after the voice feature enhancement processing is recognized and converted into a target text, the method further includes: Based on the user unique identifier, obtaining a confidence threshold; If the current confidence level corresponding to the target text is lower than the confidence threshold, the target text is filtered.

5. A page control method based on voice recognition according to claim 3 or 4, characterized in that: After matching the target text and the current execution command corresponding to the instruction set based on the preset LLM model, the method further includes: Based on the historical operation data corresponding to the user unique identifier and the current operation scenario, determining whether the currently executed command is a reasonable command; If the current execution command is a reasonable command, the operation of obtaining the target execution command based on the current execution command is performed.

6. The page control method based on voice recognition according to claim 3, characterized in that: The obtaining the target execution command based on the current execution command includes: Predicting user execution demand information based on historical operation data and current operation scenario corresponding to the user unique identifier; Based on the user execution requirement information, the current execution command is adjusted to obtain the target execution command.

7. The page control method based on voice recognition according to claim 3, characterized in that: The obtaining the target execution command based on the current execution command includes: Based on the emotional features corresponding to the target voice data, identifying the user's emotional state information; Based on the emotional state information, the current execution command is adjusted to obtain the target execution command.

8. A page control device based on voice recognition, characterized in that: include: A data receiving module, used for receiving target voice data transmitted by a remote control terminal; A recognition command module, used for obtaining a target execution command corresponding to the target voice data based on a preset instruction set; A data sending module is used to send the target execution command to the remote control terminal so that the remote control terminal converts the target execution command into a real-time message and sends it to a real-time communication server. The real-time message includes a user unique identifier and the target execution command. The real-time communication server forwards the target execution command to a large-screen terminal having a communication identifier that is the same as the user unique identifier. The large-screen terminal executes the target execution command to complete page control.

9. An electronic device, characterized in that: comprising a processor and a memory, wherein the processor is coupled to the memory; The processor is configured to execute a computer program stored in the memory, so that the electronic device executes the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The method comprises a computer program or an instruction, which, when executed on a computer, causes the computer to execute the method according to any one of claims 1 to 7.