A real-time voice interaction method based on domestic CPU and operating system

By adopting real-time voice interaction methods on different domestic CPU and operating system platforms, and using voice recognition and real-time communication technology, users' voice-triggered application operations are realized, application errors caused by platform differences are solved, and application universality and consistency are improved.

CN114153627BActive Publication Date: 2025-05-23INSPUR QILU SOFTWARE IND
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111291890.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-03
Publication Date
2025-05-23
Estimated Expiration
2041-11-03

AI Technical Summary

Technical Problem

Between different domestic CPUs and operating systems, applications need to be adapted and modified to avoid errors caused by differences.

Method used

A voice real-time interaction method based on domestic CPUs and operating systems is adopted, through the collaborative work of the client and server, and the voice recognition technology and real-time communication protocols (such as WebSocket) are used to realize user voice-triggered application operations and block the differences between CPUs and operating systems.

Benefits of technology

It realizes that various application-related operations are triggered through user voice on various terminals and servers that use domestic CPUs, avoiding adaptive modifications between different platforms, and improving the universality and consistency of applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114153627B_ABST
    Figure CN114153627B_ABST
Patent Text Reader

Abstract

The present invention particularly relates to a real-time voice interaction method based on a domestic CPU and an operating system. In the real-time voice interaction method based on a domestic CPU and an operating system, the client collects user voices, records them as audio files, and uses Socket communication technology to send the recorded audio files to the server; the server parses the audio files and sends the parsed text content to the client; the client Java application performs logical judgment based on the received text content, executes relevant logic, and sends the corresponding operation instructions to the client desktop application; the client desktop application triggers the preset action of the interface based on the received operation instructions, thereby realizing real-time voice interaction. The real-time voice interaction method based on a domestic CPU and an operating system can shield the CPU differences and operating system differences of different platforms, and realize various application-related operations triggered based on user voice on various servers using domestic CPUs and various terminals using domestic CPUs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of domestic software development, and in particular to a real-time voice interaction method based on a domestic CPU and an operating system. Background Art

[0002] In recent years, domestically produced software and hardware with independent intellectual property rights have developed rapidly, and many basic software and hardware products with independent intellectual property rights have emerged. High-end general-purpose chips with independent intellectual property rights, such as Loongson, Feiteng, and Beida Zhizhi, have flourished, and their technical level has reached the world's advanced level of similar products.

[0003] At the same time, the development of domestic operating system products is also thriving, with domestic operating system products such as Kylin Operating System, Zhongbiao Kylin Operating System, Qidian Operating System, and Phoenix Operating System emerging one after another. These operating systems are almost the same as Windows systems in terms of layout and operation.

[0004] The vigorous development of domestic operating systems has brought unprecedented opportunities for the promotion and use of domestic basic software and hardware. In addition, based on the security and reliability of domestic software and hardware, it is imperative to replace domestic software and hardware in important fields such as government and military industry.

[0005] At present, the CPU and operating system in the domestic environment can integrate multiple applications very well and fully comply with national standards. However, there are still some differences between different CPUs and operating systems, which lead to application errors. Application adaptation and modification need to be made based on different CPUs and operating systems.

[0006] In order to shield the CPU differences and operating system differences of different platforms and realize the triggering of various application-related operations based on user voice, the present invention proposes a real-time voice interaction method based on a domestic CPU and operating system. Summary of the invention

[0007] In order to make up for the defects of the prior art, the present invention provides a simple and efficient real-time voice interaction method based on a domestically produced CPU and an operating system.

[0008] The present invention is achieved through the following technical solutions:

[0009] A real-time voice interaction method based on a domestic CPU and an operating system, characterized in that it comprises the following steps:

[0010] In the first step, the client collects the user's voice and records it as an audio file;

[0011] The client program runs on a domestic operating system of a terminal using a domestic CPU, and includes two parts: a client desktop application and a client Java application. The client continuously monitors the user's voice signal. When the user issues a preset wake-up word, the client Java application can be triggered to record the user's voice and store it as a local audio file.

[0012] In the second step, the client Java application uses Socket communication technology to send the recorded audio file to the server;

[0013] In the third step, after receiving the audio file, the server parses the audio file and transmits it back to the client after identification;

[0014] The server program runs on a server using a domestically produced CPU. The server provides speech recognition functions, parses the audio files sent by the client Java application into text content, and sends the parsed text content to the client.

[0015] In the fourth step, the client Java application performs logical judgment based on the received text content, executes relevant logic, and sends the corresponding operation instructions to the client desktop application;

[0016] In the fifth step, the client desktop application triggers the preset actions of the interface according to the received operation instructions, thereby realizing real-time voice interaction.

[0017] The client desktop is built based on the Electron desktop.

[0018] The client desktop application and the client Java application achieve real-time communication through a WebSocket long connection. The client Java application manages the session by identifying the user. When the client desktop application initiates a WebSocket connection, it records and maintains the user session in order to respond to the voice command issued by the user in real time.

[0019] The preset actions include a voice wake-up success action, a start listening to user voice action, and an execution of user voice command action.

[0020] The data transmission between the client desktop application and the client Java application adopts JSON data body, which includes operation type and parameter value; after receiving the JSON data body sent by the client Java application, the client desktop application parses the JSON data body, obtains the operation type therein, and performs corresponding business logic processing according to the operation type and parameter value.

[0021] The operation types and parameter values ​​include but are not limited to: display text operation, whose parameter value is the value after voice-to-text conversion, and start listening to user voice command operation, whose parameter value is the value corresponding to the interactive action of the voice assistant.

[0022] The client Java application integrates a voice wake-up function so as to continuously monitor the user's voice signal and respond to the user's voice in a timely manner; the voice wake-up function is implemented through an open source voice wake-up SDK toolkit.

[0023] The server uses Kaldi as a speech recognition tool, and the Kaldi model selects the CVTE model.

[0024] The beneficial effect of the present invention is that the real-time voice interaction method based on the domestic CPU and operating system can shield the CPU differences and operating system differences of different platforms, and realize various application-related operations triggered by user voice on various servers using the domestic CPU and various terminals using the domestic CPU. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0026] Attached Figure 1 It is a schematic diagram of the real-time voice interaction method based on the domestic CPU and operating system of the present invention.

[0027] Attached Figure 2 The figure is a timing diagram of the user voice command from the issuance to the completion of the interaction according to the present invention. DETAILED DESCRIPTION

[0028] In order to enable those skilled in the art to better understand the technical solutions in the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0029] The real-time voice interaction method based on a domestically produced CPU and an operating system comprises the following steps:

[0030] In the first step, the client collects the user's voice and records it as an audio file;

[0031] The client program runs on a domestic operating system of a terminal using a domestic CPU, and includes two parts: a client desktop application and a client Java application. The client continuously monitors the user's voice signal. When the user issues a preset wake-up word, the client Java application can be triggered to record the user's voice and store it as a local audio file.

[0032] The recorded audio file is a 16-bit bit depth, 16000Hz sampling rate, mono, and WAV format audio file supported by the server.

[0033] In the second step, the client Java application uses Socket communication technology to send the recorded audio file to the server;

[0034] In the third step, after receiving the audio file, the server parses the audio file and transmits it back to the client after identification;

[0035] The server program runs on a server using a domestically produced CPU. The server provides speech recognition functions, parses the audio files sent by the client Java application into text content, and sends the parsed text content to the client.

[0036] In the fourth step, the client Java application performs logical judgment based on the received text content and executes relevant logic, such as opening a browser, opening a notepad, etc., and sends the corresponding operation instructions to the client desktop application;

[0037] In the fifth step, the client desktop application triggers the preset actions of the interface according to the received operation instructions, thereby realizing real-time voice interaction.

[0038] The client desktop is built based on the Electron desktop. Since Electron has good support for different platforms based on Node and Chromium, it is compatible with different domestic operating system platforms. The basic functions and UI interfaces required for building desktop applications are built through Electron. Electron applications can be quickly built through the scaffolding program provided by Electron. Based on the built Electron application, functions and interfaces are added according to user needs. The client desktop application includes real-time interactive pages and configuration related functions.

[0039] The client desktop application and the client Java application achieve real-time communication through a WebSocket long connection. The client Java application manages the session by identifying the user. When the client desktop application initiates a WebSocket connection, it records and maintains the user session in order to respond to the voice command issued by the user in real time.

[0040] The client desktop application actively initiates a WebSocket connection through a JavaScript script when it starts: it can be initiated through the main process of Electron, or through the rendering process of Electron, and creates a WebSocket long connection with the native client Java application to maintain the connection between the client desktop application and the client Java application for real-time communication, so as to respond to the voice commands issued by the user in real time. And the connection must be maintained during the program startup, otherwise the communication will not be completed, and the voice commands issued by the user will not be responded to in real time.

[0041] The preset actions include a voice wake-up success action, a start listening to user voice action, and an execution of user voice command action.

[0042] The data transmission between the client desktop application and the client Java application adopts JSON data body, which includes operation type and parameter value; after receiving the JSON data body sent by the client Java application, the client desktop application parses the JSON data body, obtains the operation type therein, and performs corresponding business logic processing according to the operation type and parameter value.

[0043] The client Java application implements subsequent business logic as needed. If it can respond directly, the client Java application responds directly. If the client desktop application needs to respond, it is packaged as JSON data and sent to the client desktop application through the WebSocket connection to execute subsequent logic.

[0044] The realization of WebSocket connection is the core of real-time voice interaction. The data transmission format of the message adopts JSON format because the data format of JSON is relatively simple, easy to read and write, the format is compressed, and it occupies little bandwidth, which is very suitable for data transmission in this scenario.

[0045] The operation types and parameter values ​​include but are not limited to: display text operation, whose parameter value is the value after voice-to-text conversion, and start listening to user voice command operation, whose parameter value is the value corresponding to the interactive action of the voice assistant.

[0046] For example, if the received data is: {"oper":"show","value":"parsed content"}, after receiving the data, first get the content show with the key oper in the JSON data, then judge it as a text display operation according to show, then get the value "parsed content" with the key value, and then display the text in the corresponding content area to be displayed according to actual business needs.

[0047] The client Java application integrates a voice wake-up function so as to continuously monitor the user's voice signal and respond to the user's voice in a timely manner; the voice wake-up function is implemented through an open source voice wake-up SDK toolkit.

[0048] The server uses Kaldi as a speech recognition tool, and the Kaldi model selects the CVTE model. After downloading the trained CVTE model, decompress it and place it in the corresponding directory of Kaldi, and then complete the relevant configuration. The audio file to be parsed should be 16-bit bit depth, sampling rate 16000Hz, mono, and wav format.

[0049] The embodiment described above is only one specific implementation of the present invention. Common changes and substitutions made by those skilled in the art within the scope of the technical solution of the present invention should be included in the protection scope of the present invention.

Claims

1. A real-time voice interaction method based on domestic CPU and operating system, It is characterized in that The following steps are involved: In the first step, the client collects the user's voice and records it as an audio file; The client program runs on a domestic operating system on a terminal using a domestic CPU, and includes two parts: a client desktop application and a client Java application; The client continuously monitors the user's voice signal. When the user issues a preset wake-up word, the client Java application is triggered to record the user's voice and store it as a local audio file. The client desktop application and the client Java application realize real-time communication through a WebSocket persistent connection. The client Java application manages the session by identifying the user. When the client desktop application initiates a WebSocket connection, it records and maintains the user session in order to respond to the voice command issued by the user in real time. In the second step, the client Java application uses Socket communication technology to send the recorded audio file to the server; In the third step, after receiving the audio file, the server parses the audio file and transmits it back to the client after identification; The server program runs on a server using a domestically produced CPU. The server provides speech recognition functions, parses the audio files sent by the client Java application into text content, and sends the parsed text content to the client. In the fourth step, the client Java application performs logical judgment based on the received text content, executes relevant logic, and sends the corresponding operation instructions to the client desktop application; Step 5: The client desktop application triggers the preset action of the interface according to the received operation instruction, thereby realizing real-time voice interaction; The data transmission between the client desktop application and the client Java application adopts JSON data body, which includes operation type and parameter value; after receiving the JSON data body sent by the client Java application, the client desktop application parses the JSON data body, obtains the operation type therein, and performs corresponding business logic processing according to the operation type and parameter value.

2. The real-time voice interaction method based on a domestic CPU and an operating system according to claim 1, Features: The client desktop is built based on the Electron desktop.

3. The real-time voice interaction method based on a domestic CPU and an operating system according to claim 1, Features: The preset actions include a voice wake-up success action, a start listening to user voice action, and an execution of user voice command action.

4. The real-time voice interaction method based on a domestic CPU and an operating system according to claim 1, Features: The operation types and parameter values ​​include but are not limited to: display text operation, whose parameter value is the value after voice-to-text conversion, and start listening to user voice command operation, whose parameter value is the value corresponding to the interactive action of the voice assistant.

5. The real-time voice interaction method based on a domestic CPU and an operating system according to claim 1, Features: The client Java application integrates a voice wake-up function so as to continuously monitor the user's voice signal and respond to the user's voice in a timely manner; the voice wake-up function is implemented through an open source voice wake-up SDK toolkit.

6. The real-time voice interaction method based on a domestically produced CPU and an operating system according to claim 1, Features: The server uses Kaldi as a speech recognition tool, and the Kaldi model selects the CVTE model.

Citation Information

Patent Citations

  • Method and device for sending information

    CN108962244A

  • Client server animation system for managing interactive user interface characters

    US5983190A