Intelligent terminal voice service input processing method and device and terminal

By classifying and abstracting the key information input from the smart terminal voice service, the problem of high cost and inflexible adaptation of voice service input in the prior art is solved, and more efficient and flexible voice service adaptation and user experience improvement are achieved.

CN120017889APending Publication Date: 2025-05-16SHENZHEN COOCAA NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510165011.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The voice service input adaptation cost of existing smart TV terminals is high and not flexible enough, resulting in high labor costs for TV manufacturers and external integrators.

Method used

By classifying and abstracting the key information input from the voice service of the smart terminal, it is classified into four categories: start event, end event, audio data and text data, and an independent processing process is determined based on these categories to realize speech recognition and processing of voice applications.

Benefits of technology

It reduces the adaptation cost, improves the flexibility and efficiency of adaptation, and significantly improves the quality and user experience of the overall voice service.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120017889A_ABST
    Figure CN120017889A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent terminal voice service input processing method and device and a terminal, and the method comprises the steps: obtaining the voice service input information of an intelligent terminal, and extracting the key information of the voice service input; information classification and abstraction are carried out on the extracted key information input by the intelligent terminal voice service, and the key information classification comprises a start event, an end event, audio data and / or text data; and according to the information classification determined by the key information input by the voice service of the intelligent terminal, controlling the independent processing flow of the corresponding classification to carry out voice recognition and voice application processing. According to the invention, the versatility and compatibility of voice service input of the television terminal can be improved, convenience is provided for television manufacturers and external integration parties, the quality and efficiency of the whole voice service are obviously improved, and convenience is provided for users.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent terminals, and in particular to an intelligent terminal voice service input processing method, device, intelligent terminal and storage medium. Background Art

[0002] With the development of terminal technology and the continuous improvement of people's living standards, the use of various smart terminals such as smart TVs is becoming more and more popular.

[0003] The voice services of smart TV manufacturer terminals in the prior art face a large workload when they are opened to the outside for integration or software function transplantation, especially for the input of voice services. Due to the differences in platform sound pickup, the external integrator needs to pay a high price to complete the terminal voice service adaptation. There are roughly two adaptation methods in the prior art: 1) The voice service input of the TV manufacturer terminal is software adapted according to the requirements of the external integrator's equipment; 2) The external integrator's system end performs system customization adaptation according to the voice service input of the TV manufacturer's terminal.

[0004] Disadvantages or shortcomings of the existing technology: high adaptation cost, lack of flexibility, and high labor cost investment is required for both the TV manufacturer's voice service terminal and the external integrator.

[0005] Therefore, the existing technology still needs to be improved and developed. Summary of the invention

[0006] The technical problem to be solved by the present invention is that, in view of the problems and defects of the above-mentioned prior art, a method, device, smart terminal and storage medium for processing voice service input of an intelligent terminal are provided. The present invention can improve the versatility and compatibility of voice service input of TV terminals, provide convenience for TV manufacturers and external integrators, significantly improve the quality and efficiency of the overall voice service, and provide convenience for users.

[0007] The technical solution adopted by the present invention to solve the problem is as follows: A method for processing voice service input of an intelligent terminal, comprising: Obtain information input by the voice service of the smart terminal and extract key information of the voice service input; Classifying and abstracting the key information extracted from the intelligent terminal voice service input, wherein the key information classification includes a start event, an end event, audio data and / or text data; According to the information classification determined by the key information input into the voice service of the intelligent terminal, the voice recognition and voice application processing are controlled according to the independent processing flow of the corresponding classification.

[0008] The method for processing voice service input of a smart terminal, wherein the step of obtaining information of the voice service input of the smart terminal and extracting key information of the voice service input comprises: The key information input for the smart terminal voice service is pre-divided into four categories, namely, start event information, end event information, audio data information, and text data information.

[0009] The method for processing voice service input of a smart terminal, wherein before the step of obtaining information of the voice service input of the smart terminal and extracting key information of the voice service input, the method further includes: Corresponding independent speech recognition and speech application processing flows are respectively set for the start event information, the end event information, the audio data information and / or the text data information.

[0010] The method for processing voice service input of a smart terminal, wherein the step of obtaining information of the voice service input of the smart terminal and extracting key information of the voice service input comprises: Acquire information input by a voice service of a smart terminal, and pre-process the information input by the voice service; From the pre-processed voice service input information, key information including start events, end events, audio data and / or text data keywords are extracted.

[0011] The method for processing voice service input of a smart terminal, wherein the step of classifying and abstracting the key information extracted from the voice service input of the smart terminal comprises: The key information extracted from the intelligent terminal voice service input is classified according to basic voice interaction; wherein the key information classification includes start event, end event, audio data and / or text data; Extract common features from the key information for classification and ignore unnecessary details.

[0012] The method for processing voice service input of a smart terminal, wherein the step of controlling the processing of voice recognition and voice application according to the independent processing flow of the corresponding classification based on the information classification determined by the key information of the voice service input of the smart terminal comprises: When the key information of the abstracted voice service input is: start event, end event, and audio data; Then, according to the start instruction of the start event, control the start of the voice process and control the UI interface processing; Send the input audio data to the background for speech recognition based on the key information of the audio data. The background generates the intermediate recognition results and notifies the UI interface for processing. And according to the end instruction of the end event, control to stop voice recognition, and notify the UI interface processing of the final voice recognition result, and perform voice instruction processing on the semantic understanding instruction generated in the final voice recognition result.

[0013] The method for processing voice service input of a smart terminal, wherein the step of controlling the processing of voice recognition and voice application according to the independent processing flow of the corresponding classification based on the information classification determined by the key information of the voice service input of the smart terminal comprises: When the key information of the abstracted voice service input is: a start event, an end event, and on-screen text data, the on-screen text data includes the intermediate recognized text data; the on-screen text data refers to the data that is converted into text by voice recognition technology and presented on the screen; Then, according to the start instruction of the start event, control the start of the voice process and control the UI interface processing; Convert the input voice content into text according to the audio data, send the recognized text, identify whether the text is the final recognition result, and notify the UI interface for processing if not; if yes, perform semantic understanding and generate voice instructions based on the semantic understanding; And according to the end instruction of the end event, control to stop the voice recognition and process the generated voice instruction.

[0014] A smart terminal voice service input processing device, wherein the device comprises: The voice input classification presetting module is used to pre-classify the key information of the intelligent terminal voice service input into four categories, namely, start event information, end event information, audio data information, and text data information; A classification processing setting module, used to set corresponding independent speech recognition and speech application processing flows for the start event information, end event information, audio data information and / or text data information respectively; The acquisition and extraction module is used to acquire the information input by the intelligent terminal voice service and extract the key information of the voice service input; A classification and abstraction module, used to classify and abstract the key information extracted from the intelligent terminal voice service input, wherein the key information classification includes a start event, an end event, audio data and / or text data; The classification processing control module is used to control the processing of voice recognition and voice application according to the independent processing flow of the corresponding classification based on the information classification determined by the key information input to the voice service of the intelligent terminal.

[0015] An intelligent terminal includes a memory and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by one or more processors, including the method for executing any one of the methods described above.

[0016] A computer-readable storage medium, wherein when instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute any one of the methods described above.

[0017] Beneficial effects of the present invention: The present invention provides a method, device, intelligent terminal and storage medium for processing voice service input of an intelligent terminal. The present invention has significant advantages over the prior art. First, by classifying and abstracting the key information of the voice service input of the TV terminal, the adaptation cost can be effectively reduced. It is no longer necessary for TV manufacturers and external integrators to perform high-cost customized adaptation, reducing the investment in manpower costs. Secondly, the voice service input of the TV terminal is classified into four categories: start event, end event, audio data, and text data, providing a set of standardized terminal voice service input solutions. This enables external integrators to freely dock and import, avoiding repeated investment in manpower costs due to platform changes. Furthermore, the scheme of the present invention improves the flexibility of adaptation. It is no longer limited to a specific platform, and the voice service input of different platforms can be efficiently adapted through this set of solutions. The present invention can improve the versatility and compatibility of the voice service input of the TV terminal, provide convenience for TV manufacturers and external integrators, significantly improve the quality and efficiency of the overall voice service, and provide convenience for users. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0019] Figure 1 It is a flowchart of a method for processing voice service input of a smart terminal provided by an embodiment of the present invention.

[0020] Figure 2 It is a processing flow diagram of a method for processing voice service input of a smart terminal provided by an embodiment of the present invention.

[0021] Figure 3 It is a processing flow diagram of a second method of processing voice service input of a smart terminal provided by an embodiment of the present invention.

[0022] Figure 4It is a processing flow diagram of mode 3 of the intelligent terminal voice service input processing method provided by an embodiment of the present invention.

[0023] Figure 5 1 is a schematic diagram of a fourth processing flow of a method for processing voice service input of a smart terminal provided in an embodiment of the present invention.

[0024] Figure 6 A principle block diagram of an embodiment of a smart terminal voice service input processing device provided by the present invention.

[0025] Figure 7 It is a block diagram of the internal structure principle of the intelligent terminal provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0026] In order to make the purpose, technical solution and advantages of the present invention clearer and more specific, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0027] It should be noted that if the embodiments of the present invention involve directional indications (such as up, down, left, right, front, back, etc.), the directional indications are only used to explain the relative position relationship, movement status, etc. between the components under a certain specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indication will also change accordingly.

[0028] Regarding the shortcomings or deficiencies of existing technologies: 1). The adaptation cost is high and not flexible enough. Whether it is the TV manufacturer's voice service terminal or the external integrator, it requires a high manpower cost investment; 2). Changing the platform requires a new round of docking, resulting in repeated investment in manpower costs; 3). The TV manufacturer's voice service terminal input lacks a set of standardized terminal voice service input solutions to be opened to external integrators for free docking and import.

[0029] Based on the characteristics of the existing TV terminal voice service input content, this application innovatively implements a more abstract and general TV terminal voice service input technology architecture. Through this voice service input technology architecture, rapid access to TV terminal voice service input can be achieved. The key technology is to enable any device end of the external integrator to quickly integrate into the TV terminal voice service, and to meet a variety of input interactions, such as accepting input from the integrator through a mobile terminal.

[0030] Based on the above invention objectives, the present invention classifies the key information of the TV terminal voice service input, and abstracts it with reference to the input interactions of various voice services currently supported by the industry, and roughly classifies the TV terminal voice service input into four categories: start events, end events, audio data, and text data. Based on these four categories, a complete TV terminal voice service input architecture is defined.

[0031] like Figure 1 As shown, a method for processing voice service input of a smart terminal according to Embodiment 1 of the present invention comprises the following steps: Step S100: Categorize the key information input by the intelligent terminal voice service into four categories in advance, namely, start event information, end event information, audio data information, and text data information.

[0032] In the voice service of a smart terminal such as a smart TV, the present invention pre-sets the input key information to be divided into four categories, namely start event information, end event information, audio data information, and text data information, in order to better manage and process the user's voice data.

[0033] The start event information refers to the start signal of the voice input. This information contains the user's activation instruction, such as "Hello, assistant" or "Start voice input". This signal indicates that the voice recognition system should start monitoring and processing subsequent audio input.

[0034] For example, when the user voice is detected saying "turn on the music", this sentence is the start event information, indicating that the user wants to start voice control, and the present invention is set to start recognizing subsequent content at this time as the start event. In this way, the start signal with clear start event information can reduce the useless monitoring time of the system and improve processing efficiency. It can also save resources because it is only activated when necessary, which can save the computing and power resources of the device.

[0035] The end event information is used to mark the end of the user's voice input, which can be a clear pause or a specific voice command, such as "end" or "done". This information helps the system understand when to stop recording and start processing the input data. For example, the user says "OK" after using the instruction. At this time, "OK" is interpreted as end event information, indicating that the system should start processing the previously recorded audio. The present invention uses a clear end signal to facilitate subsequent steps to fully judge the input content and reduce the situation of misjudgment; and the user does not need to worry about whether the voice input has been completed, and can interact more naturally.

[0036] The audio data information refers to the original audio stream of the captured user's voice. Such data is usually an essential part of the transcription process. In the speech recognition system of the present invention, the audio data needs to be processed to generate corresponding text. For example, each sentence spoken by the user in the embodiment of the present invention will generate a corresponding audio stream, such as "How is the weather?" This part of the audio stream will then be analyzed and converted into text.

[0037] The present invention can apply different audio analysis techniques, such as noise reduction, feature extraction, etc., through audio data information to improve recognition accuracy. In addition, audio data can be used to improve the speech recognition model and continuously optimize the performance of the system.

[0038] The text data information refers to the text content generated after the audio data information is analyzed and recognized. This information is a written record of the user's voice and is usually used for subsequent operations or feedback. For example, after the user says "today's weather forecast", the system converts this audio into the text message "today's weather forecast", and then provides the user with relevant weather information. The recognized text data can be conveniently used to execute instructions, search for information, or interact with other applications. Users can check the accurately recognized text content to ensure that the system understands their instructions.

[0039] In the embodiment of this step, the key information input by the intelligent terminal voice service is divided into these four categories, which can effectively improve the efficiency and accuracy of voice interaction and improve the user experience. This structured data processing method can help developers and systems better understand and respond to user needs. Through clear event recognition and data classification, it not only optimizes the use of system resources, but also provides a basis for the further development and application of voice recognition technology.

[0040] Step S200: setting corresponding independent speech recognition and speech application processing flows for the start event information, the end event information, the audio data information and / or the text data information respectively.

[0041] In the embodiment of the present invention, the start event information refers to the start signal of the recognized voice event, such as the moment when the user starts to speak. In order to accurately capture this signal, the present invention will design an independent processing flow to detect when to start processing the input.

[0042] The end event information refers to the end signal of the voice input, such as the moment when the user stops speaking. The present invention also sets an independent process for the end event to ensure that the system can accurately identify the end of the voice input.

[0043] The audio data information is the audio data related to the user's speech. The present invention also provides special processing for the audio data information, such as noise reduction, feature extraction, etc., so as to improve the accuracy of speech recognition.

[0044] The text data information refers to the text data recognized from the audio. The present invention also sets up an independent processing flow for the text data information. Independently processing this part of the data can better perform subsequent operations, such as semantic analysis or instruction recognition.

[0045] It can be seen that the present invention can ensure that each link is optimized by adopting independent processing flows for different types of information, thereby improving the overall recognition accuracy. For example, in historical data, it may be found that specific background noise affects the detection of start and end events. Through independent processing, an algorithm that is more suitable for a specific scenario can be designed. In addition, the independent processing flow allows the system to be adjusted according to different application scenarios. For example, the processing required in a noisy environment may be different from the processing method in a quiet environment, and then applied to the needs of various end users. And when errors occur during the recognition process of the system, it is easier to track the performance of each independent processing flow, so as to locate and solve the problem more quickly.

[0046] Step S300: Acquire the information of the voice service input of the intelligent terminal, and extract the key information of the voice service input; In this step, the smart terminal, such as a smart TV, first captures the user's voice input, which includes recording the user's speech through a microphone, converting it into a digital signal, and passing it to the voice recognition system. This process requires the device to accurately and timely capture the user's instructions or requests.

[0047] After receiving the user's voice input, the next task is to analyze and understand the voice to extract key information. Key information can be specific instructions, requests, emotional information or keywords, etc. For example, the user may say "Xiao K Xiao K, open the voice search function and help me check tomorrow's meeting", then the key information extracted is "open", "search", "check", "tomorrow", "meeting".

[0048] In this way, the present invention can provide a more accurate response by accurately identifying and extracting key information from the user's voice. This efficient identification can enhance the user's interactive experience. After extracting key information, the present invention can implement intelligent processing according to the user's intentions and needs, such as quickly responding or executing related tasks based on keywords, thereby improving efficiency. And when the user's voice command is accurately extracted, the intelligent terminal can understand and execute tasks faster, reducing the user's waiting time, thereby more conveniently meeting user needs.

[0049] Furthermore, the step S300 specifically includes: S301, obtaining information input by a voice service of a smart terminal, and preprocessing the information input by the voice service; In this step, the information input by the voice service of the smart terminal is obtained, that is, voice capture is performed first, and the smart terminal (such as a smart TV) is controlled to collect the user's voice signal through a microphone. This process requires ensuring that the device effectively captures the user's voice in a good environment. The device continuously monitors and determines when to start recording the user's voice through a "wake-up word" or continuous recognition. Then digital conversion is performed, and the captured analog sound signal is converted into digital data for subsequent processing by the computer.

[0050] Then, the information input by the voice service is preprocessed, including: Noise reduction processing controls the smart terminal to first eliminate or reduce background noise from the acquired voice service input information, making the user's voice clearer. This usually involves signal processing technology to remove low-frequency or high-frequency interference sounds for subsequent recognition.

[0051] Then, the audio signal is divided into small frames to facilitate the subsequent key feature extraction. The time length of each frame is generally between 20 and 40 milliseconds, which can better capture the instantaneous characteristics of the audio signal.

[0052] Feature extraction: Extract key features (such as Mel-frequency cepstral coefficients, MFCC) from the framed audio signal. These features can effectively represent the essence of the speech signal for use by subsequent speech recognition algorithms.

[0053] S302: Extract key information including start events, end events, audio data and / or text data keywords from the pre-processed voice service input information.

[0054] In this embodiment, after the above steps, the speech data becomes clearer and more standardized after the pre-processing steps such as noise reduction, framing, and feature extraction. At this point, the data is ready for further analysis and recognition.

[0055] Then extract key information, including: Start event: refers to the start recognition point of speech input, usually marking the beginning of the user's voice. This can be detected by analyzing the volume and feature changes of the audio data.

[0056] End event: refers to the end recognition point of the voice input, usually detected when the user stops speaking. Accurately defining the end event is very important to ensure accurate recognition, and it helps the system know when to stop processing the recorded voice information.

[0057] Audio data: refers to the processed digital audio signal. The system will analyze this signal and extract key information and features, such as the main sound frequency, spectrum, etc.

[0058] Text data: After being processed by the speech recognition model, the audio signal is converted into text to obtain the user's specific request or instruction. Text data is a key information carrier that helps with subsequent understanding and processing.

[0059] Keyword extraction: Extract important keywords from the identified text data. These keywords are usually the core of user intent and can directly reflect user needs. For example, when a user says "check tomorrow's weather for me", the keywords may be "check", "tomorrow", and "weather".

[0060] It can be seen that, by extracting these key information, the intelligent assistant or voice service can better understand the user's intention and provide accurate responses and services based on this. This step not only optimizes the user experience, but also lays the foundation for subsequent voice interaction.

[0061] Step S400: classifying and abstracting the key information extracted from the intelligent terminal voice service input, wherein the key information classification includes a start event, an end event, audio data and / or text data; This step mainly classifies and abstracts the key information extracted from the smart terminal voice service input, wherein the classification includes start events, end events, audio data and / or text data.

[0062] Specifically, the step S400 includes: S401, classifying the key information extracted from the intelligent terminal voice service input according to basic voice interaction; wherein the key information classification includes start event, end event, audio data and / or text data; In the previous steps, the present invention has extracted a series of information related to the user input, including the start event, the end event, the audio data and the text data.

[0063] Start event: marks the time point when user input begins, helping the system know when to start processing voice signals. End event: indicates the time point when user input ends, so that the system can stop processing voice information and recognize it. Audio data: a digital signal containing the user's voice characteristics, which can reflect the user's voice quality and emotional state. Text data: the text information obtained after conversion by the speech recognition model, which specifically describes the content of the user's request.

[0064] By classifying these key information, the present invention can more clearly understand the structure of the user's voice and perform targeted processing according to different types of information.

[0065] S402: Perform abstraction to extract common features from the classified key information and ignore unnecessary details.

[0066] In the embodiment of the present invention, after completing the information classification, the focus is on analyzing and extracting the common features of the information. The common features may include: the duration of the voice input, the repeated use of keywords, the emotional and intonation features of various types of information, the fluency and clarity of the voice, etc.

[0067] Ignoring unnecessary details: In the process of extracting common features, the system will recognize that some information may be noise or irrelevant information, which directly affects the efficiency of subsequent processing and response. For example, if the audio signal is mixed with background noise, spoken words (such as "uh", "ah", etc.) or irrelevant statements, these are details that can be ignored.

[0068] In this way, processing efficiency can be optimized, and by focusing on common features and ignoring unnecessary details, the present invention can more efficiently perform analysis and recognition in subsequent processing. This helps to improve the accuracy of voice interaction, shorten response time, and reduce the computational burden of the system.

[0069] Step S500: According to the information classification determined by the key information input into the voice service of the intelligent terminal, control the voice recognition and voice application processing according to the independent processing flow of the corresponding classification.

[0070] In the embodiment of this step, the voice input information is further processed based on the key information previously extracted and classified (such as start events, end events, audio data, text data, etc.) to perform specific voice recognition and application response.

[0071] By classifying the key information of the voice service input, the system can identify different types of instructions and requests. Each type of classified information will activate different processing procedures.

[0072] For example, if it is classified as a start event (such as turning on voice, checking the weather), the system will call the relevant voice assistant to turn on the weather query function.

[0073] According to the classification, the system of the present invention can effectively manage multiple parallel processing flows, and each flow is optimized for a specific type of instruction, so that the response is faster and more accurate.

[0074] Regarding the speech recognition process, this process focuses on converting voice data into text, which usually involves complex deep learning models.

[0075] Regarding the application processing flow, after the text is recognized, the corresponding application module will parse the text content and perform corresponding operations.

[0076] For example, when the information determined for the key information input of the smart terminal voice service is classified as the start event information, the start event information processing flow is performed: identifying the start signal of the voice event, such as the moment when the user starts speaking. In order to accurately capture this signal, the present invention detects when to start processing the input through an independent processing flow of the start event.

[0077] For example, when the key information determined by the smart terminal voice service input is classified as the end event information, the control enters the end event information processing flow, which is processed as the end signal of the voice input, such as the moment when the user stops speaking, and an independent processing flow is used to ensure that the system can accurately recognize the end of the voice input.

[0078] For example, when the key information input into the smart terminal voice service is classified as the audio data information, the control enters the processing flow of the audio data information, and is processed as audio data involving the user speaking. The present invention will perform special processing on the audio data information, such as noise reduction, feature extraction, etc., in order to improve the accuracy of speech recognition.

[0079] For example, when the key information input into the smart terminal voice service is classified as the text data information, the control enters the text data information processing flow and is processed into text data recognized from the audio. The present invention will also process the text data information independently through an independent processing flow, and independently process this part of the data, so as to better perform subsequent operations, such as semantic analysis or command recognition.

[0080] The present invention can have the following technical advantages: 1) Improve processing efficiency: Through targeted classification, the system only needs to execute the processing flow related to the user's intention, reducing unnecessary calculation and processing time, thereby improving response speed and efficiency.

[0081] 2) Improve user experience: Effective classification and independent processing enable the system to understand user needs more accurately, provide more intelligent and relevant answers, and improve the quality of interaction.

[0082] 3) Support for complex application scenarios: Through classification management, the system can handle more complex requests, such as managing multiple devices and services at the same time to meet the diverse needs of users.

[0083] In a further embodiment of the present invention, the key information of the TV terminal voice service input is classified, and the input interaction of various voice services currently supported by the industry is abstracted, and the TV terminal voice service input is roughly classified into four categories: start event, end event, audio data, and text data. Based on these four categories, the present invention defines a complete TV terminal voice service input architecture, which provides four highly abstract input methods. The specific architecture information is as follows: Method 1: The abstracted input information is: start event, end event, audio data; Figure 2 As shown, according to the TV manufacturer's voice service terminal (manufacturer), external integrator (voice application), and voice SDK (voice recognition) processing process, the following steps are included: S11, when the key information of the abstracted voice service input is: start event, end event, audio data; S12, according to the start instruction of the start event, control the start of the voice process and control the UI interface processing; S13, sending the input audio data to the background for voice recognition according to the key information of the audio data, the background generates an intermediate recognition result, and notifies the voice application end to perform UI interface processing; S14, and according to the end instruction of the end event, control to stop voice recognition, and notify the UI interface processing of the final voice recognition result, and perform voice instruction processing on the semantic understanding instruction generated in the final voice recognition result.

[0084] refer to Figure 2 As shown, the application scenario of method 1 is: Manufacturer A now needs to quickly access the voice service of the present invention on their tablet. Manufacturer A evaluates the information it can provide and decides to use mode 1 for access; Manufacturer A uses the voice interaction on the tablet computer to start recording and recognize the user's speaking content by long pressing the button, and enter the intention recognition process by releasing the button. Manufacturer A only needs to call the start event of the present invention when the user clicks. After calling the start event, Manufacturer A obtains the audio data in a way supported by itself and inputs it to the service of the present invention for voice recognition. When the user releases the button, the end event is sent to the voice service to complete the voice service access.

[0085] Solution 2: The abstracted input information is: start event, end event, and text data on the screen (including text data recognized in the middle), such as Figure 3 As shown, the processing process includes the following steps: S21, when the key information of the abstracted voice service input is: a start event, an end event, and on-screen text data, the on-screen text data includes the intermediate recognized text data; the on-screen text data refers to the data that is converted into text by voice recognition technology and presented on the screen; S22, according to the start instruction of the start event, control the start of the voice process and control the UI interface processing; S23, converting the input voice content into text according to the audio data, and sending the recognized text, identifying whether the text is the final recognition result, if not, notifying the UI interface for processing; if yes, performing semantic understanding, and generating voice instructions according to the semantic understanding; S24, and according to the end instruction of the end event, control to stop the voice recognition and process the generated voice instruction.

[0086] refer to Figure 3 As shown, the application scenario of method 2 is: Manufacturer A now needs to quickly access the voice service of the present invention on a tablet. Manufacturer A evaluates the information it can provide and decides to access using mode 2. Manufacturer A uses the start event of the present invention to be called when the user clicks and long presses to start voice interaction on the tablet. Then, the manufacturer's own business can obtain the user's audio or text input through its own services or channels. If it is audio input, Manufacturer A automatically completes the text conversion action and converts it into text synchronously. At this time, the voice service of the present invention is called as the on-screen text data to complete the dynamic input. When the user releases the click, the end event of the identification is sent to the voice service following the text to complete the voice service access.

[0087] The mf In another embodiment of the present invention, Figure 4 As shown, the input information abstracted by the third method is: start event, audio data; Figure 4 As shown, the processing process includes the following steps: S31, when the key information of the abstracted voice service input is: start event, audio data; S32, according to the start instruction of the start event, control the start of the voice process and control the UI interface processing; S33. According to the audio data, the audio data is sent to the background for voice recognition, and the voice recognition result is obtained. The UI interface is notified for processing, and it is confirmed whether it is the final recognition result. If it is, the sound pickup is stopped; semantic understanding is performed, and voice commands are generated according to the semantic understanding; if it is not the final recognition result, it is not processed, and the final recognition result is waited for.

[0088] refer to Figure 4As shown, the application scenarios of method 3 are: Manufacturer A now needs to quickly access the voice service of the present invention on the tablet. Manufacturer A decides to use mode three for access by evaluating the information it can provide. Manufacturer A clicks on the tablet to start voice interaction to trigger the voice interaction, and at the same time sends a start event to the voice service of the present invention, and simultaneously starts to input audio data to the system of the present invention. The voice service of the present invention will detect whether there is a silent segment in the audio based on the real-time input audio data, and automatically break the audio into sentences to enter the subsequent intent recognition process.

[0089] In another embodiment of the present invention, Figure 5 As shown, that is, Scheme 4, the abstracted input information is: single recognition text; S41, when the key information of the abstracted voice service input is: single recognition text; S42, the final recognition result is sent to the background for voice recognition, the simulated speech is sent, and the UI interface is notified for processing; when the simulated speech is semantically understood, a voice command is generated based on the semantic understanding.

[0090] Solution 4, the specific application scenario is: When manufacturer A now needs to quickly access the voice service of the present invention on the tablet, manufacturer A decides to use mode 4 to access by evaluating the information it can provide; manufacturer A can only obtain the final text input by the user on the tablet, or adopt the interaction of user typing input. At this time, it only needs to input a single recognition text to the voice service of the present invention according to the interface of the present invention, and the present invention will directly complete the subsequent intent recognition based on this text.

[0091] As can be seen from the above, the present invention has significant advantages over the prior art. First, by classifying and abstracting the key information of the TV terminal voice service input, the adaptation cost can be effectively reduced. It is no longer necessary for TV manufacturers and external integrators to perform high-cost customized adaptation, reducing labor cost investment. Secondly, the TV terminal voice service input is classified into four categories: start event, end event, audio data, and text data, providing a set of standardized terminal voice service input solutions. This enables external integrators to freely dock and import, avoiding repeated investment in labor costs due to platform changes. Furthermore, the scheme of the present invention improves the flexibility of adaptation. No longer limited to a specific platform, the voice service input of different platforms can be efficiently adapted through this set of solutions. The present invention can improve the versatility and compatibility of the TV terminal voice service input, provide convenience for TV manufacturers and external integrators, significantly improve the quality and efficiency of the overall voice service, and provide convenience for users.

[0092] Exemplary Devices like Figure 6As shown, an embodiment of the present invention provides a smart terminal voice service input processing device, the device comprising: The voice input classification presetting module 310 is used to pre-classify the key information of the intelligent terminal voice service input into four categories, namely, start event information, end event information, audio data information, and text data information; A classification processing setting module 320, used to set corresponding independent speech recognition and speech application processing flows for the start event information, the end event information, the audio data information and / or the text data information; The acquisition and extraction module 330 is used to acquire the information input by the voice service of the intelligent terminal and extract the key information of the voice service input; A classification and abstraction module 340 is used to classify and abstract the key information extracted from the intelligent terminal voice service input, wherein the key information classification includes a start event, an end event, audio data and / or text data; The classification processing control module 350 is used to control the processing of voice recognition and voice application according to the independent processing flow of the corresponding classification according to the information classification determined by the key information input to the voice service of the smart terminal, as described above.

[0093] Based on the above embodiments, the present invention further provides an intelligent terminal, whose principle block diagram can be shown as follows: Figure 7 As shown. The intelligent terminal includes a processor, a memory, a network interface, a display screen, and a database connected through a system bus. Among them, the processor of the intelligent terminal is used to provide computing and control capabilities. The memory of the intelligent terminal includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the intelligent terminal is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a method for processing voice service input of an intelligent terminal is implemented. The database of the intelligent terminal is used to store a voice service input processing program for an intelligent terminal.

[0094] Those skilled in the art will understand that Figure 7 The principle block diagram shown in the figure is only a block diagram of a partial structure related to the scheme of the present invention, and does not constitute a limitation on the smart terminal to which the scheme of the present invention is applied. The specific smart terminal may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0095] In one embodiment, a smart terminal is provided, comprising a memory and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by one or more processors, and the one or more programs include instructions for performing the following operations: Obtain information input by the voice service of the smart terminal and extract key information of the voice service input; Classifying and abstracting the key information extracted from the intelligent terminal voice service input, wherein the key information classification includes a start event, an end event, audio data and / or text data; According to the information classification determined by the key information input into the voice service of the intelligent terminal, the voice recognition and voice application processing are controlled according to the independent processing flow of the corresponding classification, as described above.

[0096] The step of obtaining the information input by the voice service of the intelligent terminal and extracting the key information of the voice service input includes: The key information input for the smart terminal voice service is pre-divided into four categories, namely, start event information, end event information, audio data information, and text data information.

[0097] Wherein, before the step of obtaining the information input by the voice service of the intelligent terminal and extracting the key information of the voice service input, the step further includes: Corresponding independent speech recognition and speech application processing flows are respectively set for the start event information, the end event information, the audio data information and / or the text data information.

[0098] The step of obtaining the information input by the voice service of the intelligent terminal and extracting the key information of the voice service input includes: Acquire information input by a voice service of a smart terminal, and pre-process the information input by the voice service; From the pre-processed voice service input information, key information including start events, end events, audio data and / or text data keywords are extracted.

[0099] The step of classifying and abstracting the key information extracted from the intelligent terminal voice service input includes: The key information extracted from the intelligent terminal voice service input is classified according to basic voice interaction; wherein the key information classification includes start event, end event, audio data and / or text data; Extract common features from the key information for classification and ignore unnecessary details.

[0100] The step of determining the information classification based on the key information input to the smart terminal voice service and controlling the voice recognition and voice application processing according to the independent processing flow of the corresponding classification includes: When the key information of the abstracted voice service input is: start event, end event, and audio data; Then, according to the start instruction of the start event, control the start of the voice process and control the UI interface processing; Send the input audio data to the background for speech recognition based on the key information of the audio data. The background generates the intermediate recognition results and notifies the UI interface for processing. And according to the end instruction of the end event, control to stop voice recognition, and notify the UI interface processing of the final voice recognition result, and perform voice instruction processing on the semantic understanding instruction generated in the final voice recognition result.

[0101] The step of determining the information classification based on the key information input to the smart terminal voice service and controlling the voice recognition and voice application processing according to the independent processing flow of the corresponding classification includes: When the key information of the abstracted voice service input is: a start event, an end event, and on-screen text data, the on-screen text data includes the intermediate recognized text data; the on-screen text data refers to the data that is converted into text by voice recognition technology and presented on the screen; Then, according to the start instruction of the start event, control the start of the voice process and control the UI interface processing; Convert the input voice content into text according to the audio data, send the recognized text, identify whether the text is the final recognition result, and notify the UI interface for processing if not; if yes, perform semantic understanding and generate voice instructions based on the semantic understanding; And according to the end instruction of the end event, control to stop the voice recognition and process the generated voice instruction, as described above.

[0102] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided by the present invention can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

Claims

1. A method for processing voice service input of an intelligent terminal, characterized in that: include: Obtain information input by the voice service of the smart terminal and extract key information of the voice service input; Classifying and abstracting the key information extracted from the intelligent terminal voice service input, wherein the classification of the key information includes a start event, an end event, audio data and / or text data; According to the information classification determined by the key information input into the voice service of the intelligent terminal, the voice recognition and voice application processing are controlled according to the independent processing flow of the corresponding classification.

2. The method for processing voice service input of a smart terminal according to claim 1, characterized in that: The step of obtaining the information input by the voice service of the intelligent terminal and extracting the key information of the voice service input includes: The key information input for the smart terminal voice service is pre-divided into four categories, namely, start event information, end event information, audio data information, and text data information.

3. The method for processing voice service input of a smart terminal according to claim 2, characterized in that: The step of obtaining the information input by the voice service of the intelligent terminal and extracting the key information of the voice service input also includes: Corresponding independent speech recognition and speech application processing flows are respectively set for the start event information, the end event information, the audio data information and / or the text data information.

4. The method for processing voice service input of a smart terminal according to claim 1, characterized in that: The steps of obtaining the information input by the voice service of the intelligent terminal and extracting the key information of the voice service input include: Acquire information input by a voice service of a smart terminal, and preprocess the information input by the voice service; From the pre-processed voice service input information, key information including start events, end events, audio data and / or text data keywords are extracted.

5. The method for processing voice service input of a smart terminal according to claim 1, characterized in that: The step of classifying and abstracting the key information extracted from the intelligent terminal voice service input comprises: The key information extracted from the intelligent terminal voice service input is classified according to basic voice interaction; wherein the key information classification includes start event, end event, audio data and / or text data; Extract common features from key information for classification and ignore unnecessary details.

6. The method for processing voice service input of a smart terminal according to claim 1, characterized in that: The step of determining the information classification according to the key information input to the smart terminal voice service and controlling the voice recognition and voice application processing according to the independent processing flow of the corresponding classification includes: When the key information of the abstracted voice service input is: start event, end event, and audio data; Then, according to the start instruction of the start event, control the start of the voice process and control the UI interface processing; According to the key information of the audio data, the input audio data is sent to the background for speech recognition. The background generates intermediate recognition results and notifies the UI interface for processing; And according to the end instruction of the end event, control to stop voice recognition, and notify the UI interface processing of the final voice recognition result, and perform voice instruction processing on the semantic understanding instruction generated in the final voice recognition result.

7. The method for processing voice service input of a smart terminal according to claim 1, characterized in that: The step of determining the information classification according to the key information input to the smart terminal voice service and controlling the voice recognition and voice application processing according to the independent processing flow of the corresponding classification includes: When the key information of the abstracted voice service input is: a start event, an end event, and on-screen text data, the on-screen text data includes the intermediate recognized text data; the on-screen text data refers to the data that is converted into text by voice recognition technology and presented on the screen; Then, according to the start instruction of the start event, control the start of the voice process and control the UI interface processing; Convert the input voice content into text according to the audio data, send the recognized text, identify whether the text is the final recognition result, and notify the UI interface for processing if not; if yes, perform semantic understanding and generate voice instructions based on the semantic understanding; And according to the end instruction of the end event, control to stop the voice recognition and process the generated voice instruction.

8. A voice service input processing device for an intelligent terminal, characterized in that: The device comprises: The voice input classification presetting module is used to pre-classify the key information of the intelligent terminal voice service input into four categories, namely, start event information, end event information, audio data information, and text data information; A classification processing setting module, used to set corresponding independent speech recognition and speech application processing flows for the start event information, end event information, audio data information and / or text data information respectively; The acquisition and extraction module is used to acquire the information input by the intelligent terminal voice service and extract the key information of the voice service input; A classification and abstraction module, used to classify and abstract the key information extracted from the intelligent terminal voice service input, wherein the classification of the key information includes a start event, an end event, audio data and / or text data; The classification processing control module is used to control the processing of voice recognition and voice application according to the independent processing flow of the corresponding classification based on the information classification determined by the key information input to the voice service of the intelligent terminal.

9. An intelligent terminal, characterized in that: The device comprises a memory and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by one or more processors, and the one or more programs include being used to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the method as described in any one of claims 1 to 7.