An offline intelligent device corpus modification method
Patent Information
- Application Number
- CN202010249001.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-03-31
- Publication Date
- 2026-08-18
- Estimated Expiration
- 2040-03-31
AI Technical Summary
[0004]在这一过程中,控制语料和应答语料都是在出厂时就固化在语音控制芯片之中的,如果用户存在对内置的控制或应答语料进行修改的需求,例如将“开启”的控制语音修改为“打开”,或将“设备已开启”的应答语音修改为“设备已打开”,这种需求无法通过自行操作实现,需要以厂家定制的方式进行,难以实现随时更换语音控制或语音应答内容
[0020] 1. The system establishes communication between the smart terminal and the offline smart device, retrieves the pre-stored corresponding corpus, edits and modifies the speech entries in the corpus using the smart terminal's interface, regenerates new speech entry information, and transmits it to the offline smart device for replacement. This enables the modification of the corpus in the offline smart device, meeting the user's need to modify the corpus of the offline smart device during use. Users can modify the response speech or control speech themselves.
Smart Images

Figure CN111459960B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and more specifically to a method for modifying corpora for offline intelligent devices. Background Technology
[0002] Offline voice recognition is now widely used in home appliances, such as smart refrigerators, smart toilets, and smart rice cookers. These appliances typically have control chips and corresponding operating programs pre-stored with several voice commands. When a user interacts with a home appliance that has voice control capabilities, the user's voice information is received by the acquisition module in the appliance and transmitted to the control chip for processing. The chip then determines whether the voice information is a voice command stored in the control chip. If the recognition is successful, the control chip controls the home appliance to perform the operation according to the control command corresponding to the voice command.
[0003] Existing offline voice control systems include audio file management, data storage, voice acquisition, voice playback, and voice recognition modules. In this system architecture, the stored response and control voice data, forming a pre-stored corpus, is generated using specialized tools and then uniformly burned into the production process using specific tools. User voice recognition is performed by the voice recognition module, which has an offline voice recognition model composed of a large amount of training and testing data. After acquiring user voice data, this offline voice recognition model compares and recognizes it. Upon recognizing valid control voice, the controller directs the device to perform the corresponding action, and feedback is also sent to the user through the voice playback module.
[0004] In this process, both the control and response data are embedded in the voice control chip at the factory. If users have a need to modify the built-in control or response data, such as changing the control voice of "on" to "open" or the response voice of "device is on" to "device is open", this need cannot be met by the user and must be done through manufacturer customization. It is difficult to change the voice control or voice response content at any time. Summary of the Invention
[0005] The purpose of this invention is to overcome the aforementioned defects or problems in the prior art and provide a method for modifying offline smart device corpus, which allows users to modify the control or response corpus in the offline smart device themselves, effectively meeting the personalized needs of users.
[0006] To achieve the above objectives, the present invention adopts the following technical solution:
[0007] A method for modifying speech corpus in offline smart devices, used to modify speech corpus in the voice control system of offline smart devices via a smart terminal, including:
[0008] The intelligent terminal establishes communication with the offline intelligent device and obtains feature information for identifying the offline intelligent device. Based on the feature information, it retrieves a pre-stored corpus corresponding to the voice control system in the offline intelligent device, edits the voice entries in the corpus to generate new voice entries, and transmits them to the voice control system.
[0009] The voice control system receives voice entries regenerated by the smart terminal and replaces them with corresponding voice entries pre-stored in the voice control system.
[0010] Furthermore, the smart terminal stores several different corpora, and each corpus corresponds to one or more voice control systems based on the feature information used to identify offline smart devices.
[0011] Furthermore, the corpus in the smart terminal is synchronously updated while the regenerated speech entries are transmitted to the voice control system of the offline smart device, so as to keep the corpus in the smart terminal consistent with the corpus in the corresponding voice control system of the offline smart device.
[0012] Furthermore, the corpus is a response corpus or a control corpus, and the speech entries include semantic content or audio content.
[0013] Furthermore, each of the aforementioned corpora includes several speech entries, and each speech entry has unique file identification information relative to other speech entries in the corpus.
[0014] Furthermore, the regenerated voice entry has the file identification information of its corresponding original voice entry.
[0015] Furthermore, the voice control system determines the corresponding storage location based on the file identifier information of the received regenerated voice entry, and writes the content of the received voice entry to the storage location after erasing the original data content of the storage location, so as to replace the original voice entry in the voice control system.
[0016] Furthermore, editing the speech entries in the response corpus includes editing the original speech text content of the response, editing the pitch, tone, and timbre of the human voice in the audio file generated from the speech text content of the response, and generating new audio files by recording.
[0017] Furthermore, editing the speech entries in the control corpus includes editing the original control speech text content and acquiring new semantic content through recording.
[0018] Furthermore, the smart terminal and offline smart devices are connected via wired or wireless communication; the smart terminal has a human-computer interaction interface for selecting and editing voice entries, performing recording operations, and generating audio files.
[0019] As can be seen from the above description of the present invention, compared with the prior art, the present invention has the following beneficial effects:
[0020] 1. The system establishes communication between the smart terminal and the offline smart device, retrieves the pre-stored corresponding corpus, edits and modifies the speech entries in the corpus using the smart terminal's interface, regenerates new speech entry information, and transmits it to the offline smart device for replacement. This enables the modification of the corpus in the offline smart device, meeting the user's need to modify the corpus of the offline smart device during use. Users can modify the response speech or control speech themselves.
[0021] 2. The smart terminal has multiple different corpora pre-stored, making it usable for different types of offline smart devices.
[0022] 3. The corpus in the smart terminal will be updated synchronously as the voice entries change, ensuring that the content of the corpus in the smart terminal and the offline smart device is consistent every time communication is modified.
[0023] 4. Each speech entry in the corpus has a unique file identifier, which identifies the content, corresponding instructions, storage location, and other relevant information of the speech entry, so that the voice control system of the offline smart device can recognize and call it.
[0024] 5. The editing of voice entries varies depending on the category of the voice entry being modified. For response voice entries, editing can be done by modifying the text content or by recording new voice content. Features in the voice file generated from the text content can also be modified. For control voice entries, editing can be done by directly modifying the text content or by recording voice content for semantic recognition by the smart terminal. The final voice entry includes the corresponding semantic content.
[0025] 6. The smart terminal can communicate with offline smart devices via wired or wireless means, and it also has a human-computer interaction interface to facilitate user operations such as selection and editing. Attached Figure Description
[0026] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments are briefly introduced. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0027] Figure 1 This is a flowchart illustrating an embodiment of an offline intelligent device corpus modification method provided by the present invention. Detailed Implementation
[0028] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are preferred embodiments of the present invention and should not be considered as excluding other embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0029] Unless otherwise expressly defined, the use of terms such as "first," "second," or "third" in the claims, description, and accompanying drawings of this invention is for distinguishing different objects and not for describing a specific order.
[0030] Unless otherwise expressly defined, in the claims, description, and accompanying drawings of this invention, the use of directional terms such as "center," "lateral," "longitudinal," "horizontal," "vertical," "top," "bottom," "inner," "outer," "upper," "lower," "front," "rear," "left," "right," "clockwise," and "counterclockwise" to indicate orientation or positional relationships is based on the orientation and positional relationships shown in the accompanying drawings and is only for the convenience of describing the invention and simplifying the description, and is not intended to indicate or imply that the device or element referred to must have a specific orientation or be constructed and operated in a specific orientation, and therefore should not be construed as limiting the specific scope of protection of this invention.
[0031] Unless otherwise expressly defined, the terms "fixed connection" or "fixed connection" used in the claims, description and drawings of this invention should be interpreted broadly to refer to any connection in which there is no displacement or relative rotation relationship between the two parties, including non-removable fixed connection, detachable fixed connection, integral connection and fixed connection by other means or components.
[0032] In the claims, description and accompanying drawings of this invention, the terms "comprising," "having," and variations thereof are used to mean "including but not limited to."
[0033] See Figure 1 , Figure 1A flowchart illustrating an embodiment of an offline intelligent device corpus modification method provided by the present invention is shown.
[0034] Offline smart devices refer to home smart appliances with voice control functions, such as smart rice cookers, smart toilets, smart air conditioners, and smart water heaters. These appliances typically have a corresponding voice control system. This system collects voice data, performs semantic recognition, executes actions, and provides feedback. Generally, it can communicate with a cloud server via a home network, or it can operate independently of its own operating system. The method provided in this invention targets offline home smart appliances that are not connected to a cloud server. Because the built-in voice control system of offline home smart appliances is fixed at the factory and is not connected to a cloud server, it is difficult to modify the content of their built-in voice control system.
[0035] This invention introduces a smart terminal to meet the modification needs of the speech corpus in the voice control system of an offline smart device. The smart terminal can be a user's smartphone, with an application installed to implement the method. This application uses the smartphone's communication hardware to establish communication with the offline smart device and edits the speech entries that need modification. Simultaneously, the smart terminal has a human-computer interaction interface for selecting and editing speech entries, recording audio, and generating audio files. Users can modify speech entries in the offline smart device's corpus and view relevant information intuitively through this interface.
[0036] Each offline smart device's voice control system stores a corpus, which can be divided into a response corpus and a control corpus. The response corpus provides feedback to the user while the voice control system performs actions upon acquiring control voice commands. Its basic file format is audio files, which are selected and played by the voice control system. The control corpus serves as the basis for comparing the acquired voice commands with the stored control corpus. Its basic file format is text files. The voice recognition model in the voice control system performs semantic recognition on the acquired voice commands and compares it with the stored control corpus to achieve the acquisition of control voice commands.
[0037] The method provided by this invention includes the following steps:
[0038] The smart terminal establishes communication with the offline smart device and obtains the feature information used to identify the offline smart device. Based on the feature information, it retrieves a pre-stored corpus corresponding to the voice control system in the offline smart device, edits the voice entries in the corpus to generate new voice entries, and transmits them to the voice control system of the offline smart device. The voice control system receives the newly generated voice entries from the smart terminal and replaces them with the original voice entries pre-stored in the voice control system.
[0039] Specifically, such as Figure 1 As shown, it includes the following steps:
[0040] S101. The smart terminal establishes communication with the offline smart device through wireless communication (such as Bluetooth);
[0041] S102. The smart terminal obtains the feature information used to identify the offline smart device;
[0042] S103. The smart terminal retrieves the pre-stored corpus of the voice control system corresponding to the offline smart device based on the acquired feature information.
[0043] S104. The user selects the corpus to be modified on the interactive interface of the smart terminal, such as the response corpus or the control corpus.
[0044] S105. The user selects the voice item to be modified from the list of voice items.
[0045] S106. The user edits the selected voice item on the editing interface displayed on the smart terminal;
[0046] S107. After the user finishes editing, the smart terminal generates a new voice entry;
[0047] S108, the smart terminal will transmit the newly generated voice entries to the offline smart device;
[0048] S109. The voice control system of the offline smart device verifies the validity of the new voice entries received.
[0049] S110, The voice control system of the offline smart device determines the storage location of the original voice entry corresponding to the received voice entry based on the file identification information of the voice entry.
[0050] S111 The voice control system of the offline smart device erases the data content of the storage location and writes the content of the received voice entry.
[0051] Specifically, the application's content library on the smart terminal stores several different corpora. Each model of offline smart device also possesses unique identifiers. The corpus of the voice control system for each model of offline smart device corresponds to the corpus on the smart terminal; that is, the corpus of a specific model of offline smart device on the smart terminal is completely identical to the corpus of its voice control system. Furthermore, the smart terminal can communicate with a cloud server. After obtaining the feature information of the offline smart device, if a corpus corresponding to that model is not found in the pre-stored corpus set, the smart terminal requests data from the cloud server to download the corresponding corpus to the corpus set for later use.
[0052] Each corpus consists of several speech entries, and each speech entry has unique file identification information relative to other speech entries in the corpus. This file identification information may specifically include model name, model ID, command word ID, original corpus content, original corpus length, new corpus content, and new corpus length.
[0053] Of course, when a smart terminal establishes communication with an offline smart device, the feature information sent by the offline smart device also includes data content specifically identifying that offline smart device. Based on this feature information, the smart terminal retrieves the corresponding corpus for that offline smart device and displays it as a list on the smart terminal's interface. The user can then select the voice entries they wish to modify for editing.
[0054] Since a corpus typically includes a response corpus and a control corpus, before a user selects a specific voice item, they also need to choose to modify the content in either the response corpus or the control corpus.
[0055] For example, after selecting the response corpus, the user enters the list of response audio entries. The user selects one audio entry and enters the editing interface. At this point, the user can choose the following methods to edit the audio entry:
[0056] ① Modify the text content of the voice entry, and the application will generate the corresponding audio content based on the text content;
[0057] ② Generate new audio content directly by recording.
[0058] Of course, in method ① above, the characteristics of the audio content can be adjusted before generating the corresponding audio content. For example, the pitch, tone, and timbre of the human voice in the audio can be modified and edited, or even a specific human voice can be selected for audio content synthesis.
[0059] The modification of the response corpus ultimately results in an audio file, which is then transmitted to an offline smart device for replacement. When the response voice entry is triggered, the offline smart device plays this audio file.
[0060] For example, after selecting the control corpus, the user enters the list of control speech entries. The user selects one speech entry and enters the editing interface. At this point, the user can choose the following methods to edit the speech entry:
[0061] ① Modify the text content of this audio entry;
[0062] ② Audio content is generated by recording, and the application performs semantic recognition on the audio content to generate new text content.
[0063] The modification of the control corpus ultimately results in a text file, which is then transmitted to an offline smart device for replacement. When the collected speech matches the semantic content contained in the text file, the offline smart device triggers the corresponding action.
[0064] The audio or text content generated above constitutes the new voice entry. This voice entry is transmitted from the smart terminal to the offline smart device. During the transmission, the corpus on the smart terminal is updated synchronously, that is, the new voice entry replaces the original voice entry in the corpus on the smart terminal, so as to ensure that the corpus on the smart terminal is consistent with the corpus in the corresponding voice control system of the offline smart device.
[0065] Furthermore, the regenerated voice entry contains the file identification information of the original voice entry. After receiving a new voice entry, the voice control system of the offline smart device first verifies the validity of the voice entry, confirming that its file format is the required format and contains valid data content. Then, it extracts the file identification information attached to the voice entry, determines the storage location of the original voice entry corresponding to the newly generated voice entry based on the file identification information, erases the data content of the storage location, writes the content of the received voice entry into the storage location, and returns the corresponding result, which is displayed on the smart terminal, thereby completing the modification and replacement of the original voice entry in the voice control system of the offline device.
[0066] In the processing of the above-mentioned voice control system, it needs to determine whether the model name and model ID attached to the voice entry are correct, then determine whether the corresponding command word ID and the original corpus information exist, and then call the corresponding API interface of the database management to replace the corpus.
[0067] This invention provides a method for modifying speech corpus in offline smart devices. By introducing a smart terminal and establishing communication between the smart terminal and the offline smart device, the content of the voice control system in the offline smart terminal can be modified. Users can modify the control or response speech corpus in the offline smart device themselves, effectively meeting the personalized needs of users.
[0068] The foregoing description of the specifications and embodiments is intended to explain the scope of protection of this invention, but does not constitute a limitation on the scope of protection of this invention. Modifications, equivalent substitutions, or other improvements to the embodiments of this invention or a portion thereof that can be obtained by those skilled in the art through logical analysis, reasoning, or limited experimentation, based on the teachings of this invention or the foregoing embodiments, in conjunction with common knowledge, general technical knowledge, and / or existing technology, should all be included within the scope of protection of this invention.
Claims
1. A method for modifying speech corpus in an offline smart device, used to modify speech corpus in the voice control system of an offline smart device via a smart terminal, characterized in that, include: The intelligent terminal establishes communication with the offline intelligent device and obtains feature information for identifying the offline intelligent device. Based on the feature information, it retrieves a pre-stored corpus corresponding to the voice control system in the offline intelligent device, edits the voice entries in the corpus to generate new voice entries, and transmits them to the voice control system. The voice control system receives voice entries regenerated by the smart terminal and replaces them with corresponding voice entries pre-stored in the voice control system. The smart terminal stores several different corpora, and each corpus corresponds to one or more voice control systems based on the feature information used to identify offline smart devices. The corpus includes a response corpus and a control corpus. The response corpus includes the audio content played by the offline smart device, and the control corpus includes a text file. The text file is used to trigger a corresponding action by the offline smart device when the voice collected by the offline smart device matches the text file. When the smart terminal modifies the response corpus, it is configured to generate corresponding new audio content based on the text content of the modified voice entry, or to generate new audio content by recording. When the smart terminal modifies the control corpus, it is configured to modify the text content of the voice entry to generate new text, or to generate audio content by recording and then generate a new text file through semantic recognition. The offline intelligent device plays new audio content when the collected voice triggers a voice entry in the response corpus, and triggers a corresponding action when the collected voice matches the semantic content contained in the new text file.
2. The method for modifying offline intelligent device corpus as described in claim 1, characterized in that, The corpus in the smart terminal is synchronously updated while the regenerated speech entries are transmitted to the voice control system of the offline smart device, so as to keep the corpus in the smart terminal consistent with the corpus in the corresponding voice control system of the offline smart device.
3. The method for modifying offline intelligent device corpus as described in claim 1, characterized in that, The corpus is a response corpus and / or a control corpus, and the speech entries include semantic content and / or audio content.
4. The method for modifying offline intelligent device corpus as described in claim 1, characterized in that, Each of the aforementioned corpora includes several speech entries, and each speech entry has unique file identification information relative to other speech entries in the corpus.
5. The method for modifying offline intelligent device corpus as described in claim 4, characterized in that, The regenerated voice entry has the file identification information of its corresponding original voice entry.
6. The method for modifying offline intelligent device corpus as described in claim 5, characterized in that, The voice control system determines the corresponding storage location based on the file identifier information of the received regenerated voice entry, and writes the content of the received voice entry to the storage location after erasing the original data content of the storage location, so as to replace the original voice entry in the voice control system.
7. The method for modifying offline intelligent device corpus as described in claim 3, characterized in that, Editing the speech entries in the response corpus includes editing the original speech text content of the response, editing the pitch, tone, and timbre of the human voice in the audio file generated from the speech text content of the response, and generating new audio files by recording.
8. The method for modifying offline intelligent device corpus as described in claim 3, characterized in that, Editing the speech entries in the control corpus includes editing the original control speech text content and acquiring new semantic content through recording.
9. The method for modifying offline intelligent device corpus as described in claim 1, characterized in that, The smart terminal and offline smart devices communicate via wired and / or wireless communication; the smart terminal has a human-computer interaction interface for selecting and editing voice entries, performing recording operations, and generating audio files.
Citation Information
Patent Citations
System, method and computer program product for query clarification
CA3048436A1
Methods and systems for speech recognition processing using search query information
CN104854654A
Self-updating semantic understanding system and method
CN107015969A
Corpus generation method, device, electronic equipment and readable storage medium
CN110399499A
Method of corpus structuralization and device
CN102982036A