Formatted data voice program control interaction method

Through the formatted data voice program control interaction method, the complexity and speech recognition complexity of traditional remote control systems are solved, flexible control adaptation and efficient human-computer voice interaction are achieved, and it is suitable for many fields such as home, industrial and medical equipment.

CN120279905APending Publication Date: 2025-07-08NORTHWEST ELECTROMECHANICAL ENG RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510340788.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

Traditional remote control systems need to create specific control buttons or commands in each control interface, which makes the control service complex and not easy to expand. Voice control requires a complex voice recognition system to adapt to various pronunciations.

Method used

The formatted data voice program control interaction method is adopted, and the voice control model is designed by defining object characteristics, behavior elements and actions. The command input bar, the command understanding bar and the task text bar are used to standardize the speech input, and the accuracy confirmation is carried out in combination with the voice recognition library to realize human-computer voice control inside and outside the system.

Benefits of technology

It realizes more flexible adaptation to control needs, reduces development and maintenance costs, improves operating efficiency and system intelligence, and is suitable for multiple fields such as home, industrial and medical equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279905A_ABST
    Figure CN120279905A_ABST
Patent Text Reader

Abstract

The invention discloses a formatted data voice program control interaction method, which comprises the following steps of: determining a control requirement and defining object characteristics, behavior elements and actions in a working scene according to working requirements and capabilities of a logistics and warehousing system based on voice control; designing a voice control model according to control requirements; a task input interface of the voice control model comprises an instruction input field, an instruction understanding field and a task text field, wherein the instruction input field is used for displaying received voice input; after the instruction understanding column receives and recognizes a voice recognition result of an instruction of a phrase type in voice input, generating a corresponding text template; after the action is analyzed, the task text bar is used for carrying out association representation on the filled text template so as to generate a control task; the method specifically comprises the following steps: retrieving information in each row of data in a text template, separating the text template according to a set identifier, performing data extraction and protocol combination according to fields, and generating a code, thereby obtaining a control task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of information software, and particularly to a method for formatted data voice program-controlled interaction. Background Art

[0002] Traditional remote control systems are usually restricted as follows: specific control buttons or commands need to be created in each control interface, which leads to the complexity and poor scalability of control services. In addition, the wide application of voice control usually involves the diversity of voice patterns, so a complex voice recognition system is required to adapt to various pronunciations. Summary of the Invention

[0003] The purpose of the present invention is to provide a method for formatted data voice program-controlled interaction to more flexibly adapt to different control requirements and voice models, thereby improving the user experience.

[0004] To achieve the above task, the present invention adopts the following technical solutions:

[0005] A method for formatted data voice program-controlled interaction, comprising:

[0006] Determine the control requirements according to the working requirements and capabilities of a logistics and warehousing system based on voice control, define the object features, behavior elements and actions in the working scenario, and record the defined object features, behavior elements and actions;

[0007] Design a voice control model according to the control requirements; the task input interface of the voice control model includes an instruction input bar, an instruction understanding bar, and a task text bar, where:

[0008] The instruction input bar is used to display the received voice input. When the warehouse name + warehouse number is recognized in the instruction input bar, the voice control scenario is evoked and the conversation is initiated;

[0009] After the instruction understanding bar receives and recognizes the voice recognition result of the phrase-type instruction in the voice input, a corresponding text template is generated, and the content in the text template is filled according to the parsing result of the free message type in the voice input;

[0010] After the action is parsed, the task text bar is used to perform an associated representation on the filled text template to generate a control task; specifically, retrieve the information in each row of data in the text template, separate the text template according to the set identifier, perform data extraction and protocol combination by field, and generate a code to obtain the control task.

[0011] Further, the object features include name and number; the behavior elements include parameters and actions; the names are "warehouse name", "forklift name", and "material name"; the numbers are "warehouse number", "forklift number", and "material number", the parameters include "quantity of materials"; the instructions include "warehousing", "outbound", "query", and "sorting"; the actions include "cancel", "empty", "withdraw", and "confirm".

[0012] Further, after the instruction input field of the voice control model receives voice input, it first formats the type identifiers in the voice input. The type identifiers include three types: keywords, phrases, and free text messages. Among them:

[0013] The warehouse name and actions are defined as the keyword type; the instructions, including warehousing, outbound, and query, are defined as the phrase type; other information in the voice input is defined as the free text message type;

[0014] The warehouse name, forklift name, material name, and material number are stored in the retrieval database as retrieval terms for matching the corresponding text templates; and retrieval terms are associated from the free text message type recognized from the voice input for element identification and extraction, and the quantity of materials is identified and extracted according to the retrieval term segmentation relationship.

[0015] Further, for the vocabulary of the free text message type in the voice input, an incentive response is used for recognition; first, the parameters in the free text message type are extracted according to the material name and material number in the current context, and then the parameters are associated and displayed in the currently queried text template; data identification and association are automatically performed in this process, and information is automatically matched and understood.

[0016] Further, the names and numbers in the object features, or the representation of names, labels, and parameters, are made in the text template; among them, the task number is automatically identified and generated according to the types "warehousing", "outbound", and "sorting" included in the instructions corresponding to the parsed phrases, and the task type and number are separated by "-"; the "warehouse name", "forklift name", "material name" and "warehouse number", "forklift number", "material number" are separated by "-"; the parameters are automatically filled in with the material name.

[0017] Further, after the text template is called and filled, when receiving "cancel" in the actions, the previous data modification will be cancelled, when receiving "clear", the entire text template will be emptied, when receiving "withdraw", the current operation will be exited, and when receiving "confirm", the control task will be issued.

[0018] Further, the control task is sent to the logistics control system via Ethernet, through the IP address and port, and the logistics control system sends the control task to the corresponding transport vehicle via wifi; during this process, the logistics control system can display the control task on the user terminal to enable the user to confirm and allocate the control task.

[0019] Further, a dedicated voice input recognition library is established to associate and equate the recognition retrieval words in the voice input, and display the results in the instruction understanding column for manual input to confirm or cancel the accuracy of the association.

[0020] Association rules for various retrieval words are set in the voice input recognition library; specifically, a relationship data table for object features, behavior elements, and actions is established to store the equivalent association rules for misspelled words, wrongly-written characters, incorrect word order, disordered words, and modifiers.

[0021] A terminal device includes a processor, a memory, and a computer program stored in the memory; when the processor executes the computer program, the formatted data voice program-controlled interaction method is implemented.

[0022] A computer-readable storage medium stores a computer program; when the computer program is executed by a processor, the formatted data voice program-controlled interaction method is implemented.

[0023] Compared with the prior art, the present invention has the following technical features:

[0024] 1. Expressing machine program control in a literal and standardized manner

[0025] The present invention allows the definition of machine program control to be expressed in a literal and standardized manner, which provides a clearer and more understandable basis for human-machine interaction inside and outside the system. It promotes the standardization in the field of machine program control, reduces the development and maintenance costs, and also improves the technical scalability.

[0026] 2. Implementing human-machine voice control inside and outside the system

[0027] The formatted data voice program-controlled interaction method of the present invention is not limited to a single application field, but also implements human-machine voice control inside and outside the system. This means that users can use voice to interact with various devices and systems, whether in a home environment, an industrial site, or the operation of medical equipment.

[0028] 3. Intelligent operation

[0029] By combining voice recognition and data processing, users no longer need to perform cumbersome manual operations, but can achieve complex control tasks through natural voice input. This not only improves the operation efficiency, but also enhances the intelligence of the system, thus adapting to the changing environment and requirements.

[0030] 4. Promote the practice of software specification

[0031] The present invention is not only a practice of voice program control, but also promotes the practice of software specification. By establishing a control protocol, a data mapping table, and a text template, the present invention leads a new method in the software field and provides a model for standardized software design and development. This has a profound impact on improving software quality, reducing errors, and enhancing maintainability. Description of the Drawings

[0032] Figure 1 is a schematic flow chart in an embodiment of the present invention;

[0033] Figure 2 is the rule for generating the text template of the present invention;

[0034] Figure 3 is the voice recognition application rule of the present invention;

[0035] Figure 4 is the rule for interpreting recognition data of the present invention;

[0036] Figure 5 is the text template representation rule of the present invention;

[0037] Figure 6 is the data coding representation rule of the present invention. Detailed Description of the Invention

[0038] In view of the complex control and refined control of digital program control, the present invention provides a formatted data voice program control interaction method; for program control data, a standardized data set and an element set are associated, and the voice recognition data is specially trained and pattern recognized, and the recognition data input, key point data retrieval, message data association, and remote control are performed. By textually normalizing the machine program control instructions, the matching and correlation input of the internal and external voice recognition data of the system are performed, and then the parsing, encoding, and transmission of the text data are carried out; the uniqueness of this interaction method lies in its wide applicability, which can be used in various application scenarios such as automatic control, remote control, smart home systems, robot control, intelligent transportation systems, industrial automation, etc., and is particularly good at processing digital program control data to achieve the intelligence, automation, and standardization of data processing.

[0039] See the attached Figure 1 , a formatted data voice program control interaction method provided by the present invention includes:

[0040] Step 1, first, according to the working requirements and capabilities of the logistics and warehousing system based on voice control, determine the control requirements, mainly including defining the object characteristics, behavior elements, and actions in the working scenario, and recording the defined object characteristics, behavior elements, and actions.

[0041] The object features include name and number; the behavior elements include main features such as parameters and actions; in this working scenario: the names are "warehouse name", "forklift name", "material name"; the numbers are "warehouse number", "forklift number", "material number", such as numerals like "12306"; the parameters include "quantity of materials"; the instructions include "inbound", "outbound", "query", "sorting"; the actions include verbs such as "cancel", "empty", "cancel", "confirm".

[0042] Step 2, design a voice control model according to the control requirements; the task input interface of the voice control model includes an instruction input bar, an instruction understanding bar, and a task text bar, where:

[0043] (1) Instruction input bar.

[0044] The instruction input bar is used to display the received voice input; the voice input includes the above-mentioned object features and behavior elements; the instruction input bar is also used to display the actions in the voice input; after recognizing "warehouse name" + "warehouse number" in the instruction input bar, the voice control scenario is evoked and the conversation starts.

[0045] In this solution, to avoid confusion in the recognition of voice input, after the instruction input bar of the voice control model receives the voice input, it first formats the type identifiers in the voice input. The type identifiers include three types: keyword, phrase, and free text message, where:

[0046] The warehouse name and actions are defined as the keyword type; the instructions, including inbound, outbound, and query, etc., are defined as the phrase type; other information in the voice input is defined as the free text message type.

[0047]

[0048] (2) Instruction understanding bar.

[0049] After the instruction understanding bar receives and recognizes the instructions of the phrase type in the voice input, including voice recognition results such as "inbound", "outbound", "query", "sorting", etc., it generates the corresponding text template and fills the content in the text template according to the parsing results of the free text message type in the voice input.

[0050] According to the flexible requirements of voice input data, for the vocabulary of the free message type in voice input, an incentive response is used for recognition; first, according to the "material name" and "material number" in the current context, the parameters (i.e., the quantity of materials) in the free message type are extracted, and then the parameters are associated and displayed in the currently queried text template; in this process, data recognition and association are automatically performed, and the information is automatically matched and understood for intermediate process confirmation and verification, as shown in the example:

[0051]

[0052] In this step, background self-organizing data encapsulation and context-based input understanding are adopted, which reflects the advantages of anthropomorphic interactive control. Through the fixed protocol format method, the control protocol is made more accurate. By defining the protocol elements for different application scenarios, a flexible and extensible formatted data representation method is provided for voice programming control.

[0053] Among them, the name and number in the object characteristics, or the representation of the name, label, and parameters are carried out in the text template; among them, the task number is automatically identified and generated according to the types "warehouse entry", "warehouse exit", and "sorting" contained in the instructions corresponding to the parsed phrases, and the task type and number are separated by "-"; the "warehouse name", "forklift name", "material name" and "warehouse number", "forklift number", "material number" are separated by "-"; the parameters are automatically filled or new type supplements are added along with the material name.

[0054] In this example, after "warehouse entry" in the instruction is recognized, the called text template is the "warehouse entry template":

[0055]

[0056] (3) Task text column.

[0057] After parsing the action of the keyword type, the task text column is used to perform an associated representation on the filled text template to generate a control task; specifically, it retrieves the information in each row of data in the text template, separates the text template according to the set identifiers, namely "; " and "-", extracts data by field and combines protocols (in the format of the communication protocol), and generates a code to obtain the control task;

[0058] For example:

[0059]

[0060] After the "text template" is called and filled, when "Undo" in the received action is received, the previous data modification will be undone. When "Clear" is received, the entire text template will be cleared. When "Cancel" is received, the process will exit. When "Confirm" is received, the control task will be issued. The control task is sent to the logistics control system via Ethernet, through the IP address and port. The logistics control system sends the control task to the corresponding transport vehicle via wifi. During this process, the logistics control system can display the control task on the user terminal, enabling the user to confirm and manage the distribution of the control task.

[0061] In this solution, due to the adoption of data screening, data association, and task association processing procedures, a high degree of compatibility is demonstrated for information fault tolerance. At the same time, data confirmation is verified and unified, thereby making voice control more accurate and significantly improving the user's interaction efficiency.

[0062] Based on the above technical solution, for the voice recognition part, the following designs are also carried out:

[0063] When recognizing the voice input, first, each word in the voice input is obtained. Then, in combination with the voice input recognition library, if the word is any one of the defined object features, behavior elements, or actions, the word is considered a retrieval word.

[0064] Combined with the control requirements in the scenario, considering application lightweight and specialization, a dedicated voice input recognition library is established to associate and equate the recognized retrieval words in the voice input, and the results are displayed in the instruction understanding column for manual input to confirm or cancel the accuracy of the association. Specifically:

[0065] Association rules for various retrieval words are set in the voice input recognition library. Specifically, a relational data table corresponding to object features, behavior elements, and actions is established to store the equivalent association rules for misspelled words, near-homophonic words, incorrect word order, disordered word order, and modifiers, as follows:

[0066] Constraint rules for names: Homophonic words, near-homophonic words, easily misspelled words, and misspelled words of names are associated, and equivalent associations are made for incorrect word order, disordered word order, and reversed word order. That is, homophonic words, near-homophonic words, easily misspelled words, and misspelled words of a certain name are all considered to be that name, and incorrect word order, disordered word order, and reversed word order of the word order are all considered to have equivalent meanings.

[0067] Constraint rules for numbers: Homophonic words and near-homophonic words of numbers are associated.

[0068] Constraint rules for instructions: Homophonic words, near-homophonic words, easily misspelled words, and misspelled words of instructions are associated, and equivalent associations are made for incorrect word order, disordered word order, and reversed word order.

[0069] Parameter Constraint Rules: Associate homophones and near-homophones of parameters, define and convert the relationships of numerals, quantifiers such as "ones, tens, hundreds, thousands, ten thousands", and adverbs such as "a little", "a bit", "a little more", and perform multiple-relationship conversions on time, quantifiers, and adverbs.

[0070] Action Constraint Rules: Associate homophones, near-homophones, easily misspelled words, and wrongly written or misused characters of actions.

[0071] In this step, due to the additional design of the relational data table, misspelled words, wrongly written or misused characters, incorrect or disordered orders, and other special features in the recognized voice input are associated and incorporated, with high compatibility and reliability for voice control in the application scenario, ensuring higher user input efficiency.

[0072] Examples:

[0073] The first step, voice wake-up:

[0074] Taking the chemical warehousing and logistics system as a typical object, in the application scenario, after the instruction input bar receives the voice input "Isolated Warehouse" ("warehouse name"), it starts to initiate a conversation.

[0075] The second step, task wake-up:

[0076] Voice input "Warehouse No. 1" ("warehouse number"), and this instruction combines with "Isolated Warehouse" to activate the task input interface of "Isolated Warehouse No. 1".

[0077] The third step, instruction input:

[0078] Voice input "Out of storage" (instruction), and this instruction starts to call the formatted text template in the task input interface of "Isolated Warehouse No. 1", and the text template shows "Task number: Out of storage - Task001; Isolated Warehouse - No. 1;".

[0079] The fourth step, task input:

[0080] Voice input "Robot No. 3 takes out 5 tons of asbestos cloth from storage" (parameters), and this instruction is translated and displayed in the text template of "Isolated Warehouse No. 1" as "Robot - No. 3; Asbestos cloth - 5 tons;", as Figure 2 shown.

[0081] In this step, the possible Sichuan dialect information input is "Sai o hao robot produced Wu Dong's silk cotton cloth". First, retrieve this information from the relationship data table, find the equivalent descriptions of "Sai o hao", "Wu Dong", and "silk cotton", and replace them with "Robot No. 3 produced 5 tons of asbestos cloth"; find the implicit relationships of "produced = out of storage" and "5 tons = 50 boxes", and further understand the input information as "Robot No. 3 out of storage 50 boxes of asbestos cloth"; the understood information is displayed in the instruction understanding column on the task input interface for manual confirmation and cancellation operations of the input instruction accuracy.

[0082] Step 5, Task Understanding and Text Representation:

[0083] After the instruction understanding is completed, match and screen the instruction information with "forklift truck name", "forklift truck number", "material name", "material number", and "material quantity" in sequence, and extract information such as "robot", "No. 3", "asbestos cloth", and "50 boxes" from "Robot No. 3 out of storage 50 boxes of asbestos cloth". This conversion process is as Figure 3 shown, and the extracted information substitution items are displayed in the text template. This display process is as Figure 4 shown.

[0084] Step 6, Control Instruction Generation:

[0085] Voice input "confirm" ("action"), separate the text template by line, ";", and "-", perform data extraction and protocol generation by field. The text template reading process is as Figure 5 shown. The protocol data is encoded as "[Isolation Warehouse][1][Robot][3][Out of Storage][Asbestos Cloth]

[50] ". The data encoding representation rule is as Figure 6 shown. Among them, the "out of storage" information comes from the initial instruction, and other information comes from the text template; when the relationship between "asbestos cloth" and "50 boxes" is confirmed to be valid, the protocol data is sent to Robot No. 3 through the Ethernet protocol.

[0086] Step 7, Robot No. 3 receives the data, performs data parsing, generates tasks and automatically plans for execution.

[0087] In this step, Robot No. 3 can also be an intermediate order dispatching system, and the dispatching system then assigns the receiving robot. After receiving the data, the dispatching system synchronously performs data parsing and display to realize the reproduction of the data in the remote text template, so as to maintain manual secondary inspection.

[0088] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A method for formatted data voice program-controlled interaction, characterized in that Including: Determine the control requirements according to the working requirements and capabilities of the voice-controlled logistics and warehousing system, define the object characteristics, behavior elements and actions in the working scenario, and record the defined object characteristics, behavior elements and actions; Design a voice control model according to the control requirements; the task input interface of the voice control model includes an instruction input bar, an instruction understanding bar, and a task text bar, where: The instruction input bar is used to display the received voice input. When the warehouse name + warehouse number is recognized in the instruction input bar, the voice control scenario is evoked and the conversation is started; After the instruction understanding bar receives and recognizes the speech recognition result of the instruction of the phrase type in the voice input, a corresponding text template is generated, and the content in the text template is filled according to the parsing result of the free message type in the voice input; After the action is parsed, the task text bar is used to perform an associated representation on the filled text template to generate a control task; specifically, retrieve the information in each row of data in the text template, separate the text template according to the set identifier, perform data extraction and protocol combination by field, and generate a code to obtain the control task.

2. The formatted data voice program-controlled interaction method according to claim 1, wherein, The object characteristics include name and number; the behavior elements include parameters and actions; the names are "warehouse name", "forklift name", "material name"; the numbers are "warehouse number", "forklift number", "material number", the parameters include "quantity of materials"; the instructions include "warehousing", "outbound", "query", "sorting"; the actions include "cancel", "empty", "cancel", "confirm".

3. The formatted data voice program-controlled interaction method according to claim 1, wherein After the instruction input bar of the voice control model receives the voice input, first format the type identifier in the voice input. The type identifier includes three types: keyword, phrase, and free message, where: The warehouse name and action are defined as the keyword type; the instructions, including warehousing, outbound, and query, etc., are defined as the phrase type; other information in the voice input is defined as the free message type; Store the warehouse name, forklift name, material name, and material number as search terms in the search database for matching the corresponding text template; and perform element recognition and extraction by associating the search terms from the free message type recognized from the voice input, and the quantity of materials is recognized and extracted according to the segmentation relationship of the search terms.

4. The formatted data voice program-controlled interaction method according to claim 1, characterized in that For the vocabulary of the free message type in the voice input, an incentive response is used for recognition; first, according to the material name and material number in the current context, extract the parameters in the free message type, and then associate and display the parameters to the currently queried text template; in this process, data recognition and association are automatically performed, and information is automatically matched and understood.

5. The formatted data voice program-controlled interaction method according to claim 1, wherein The names and numbers in the object features, or the representation of names, labels, and parameters, are carried out in the text template; among them, the task numbers are automatically identified and generated according to the types "warehousing", "outbound", and "sorting" included in the instructions corresponding to the parsed phrases, and the task types and numbers are separated by "-"; the "warehouse name", "forklift name", "material name" and "warehouse number", "forklift number", "material number" are separated by "-"; the parameters are automatically filled in with the material name.

6. The formatted data voice program-controlled interaction method according to claim 1, characterized in that After the text template is called and filled, when receiving the revocation in the action, the previous data modification will be revoked, when receiving the clear, the entire text template will be emptied, when receiving the cancellation, it will exit, and when receiving the confirmation, the control task will be issued.

7. The formatted data voice program-controlled interaction method according to claim 1, characterized in that The control task is sent to the logistics control system via Ethernet through the IP address and port. The logistics control system sends the control task to the corresponding forklift via wifi; during this process, the logistics control system can display the control task on the user terminal for the user to confirm and manage the distribution of the control task.

8. The formatted data voice program-controlled interaction method according to claim 1, characterized in that A dedicated voice input recognition library is established to associate and equate the recognition search terms in the voice input, and the results are displayed in the instruction understanding column for manual input to confirm or cancel the accuracy of the association. Association rules for various search terms are set in the voice input recognition library; specifically, a relational data table corresponding to object features, behavior elements, and actions is established to store the equivalent association rules for misspelled words, wrongly written characters, incorrect order, disordered order, and modifiers.

9. A terminal device, comprising a processor, a memory, and a computer program stored in the memory; characterized in that, When the processor executes the computer program, it implements the formatted data voice program-controlled interaction method according to any one of claims 1-8.

10. A computer-readable storage medium storing a computer program; characterized in that, When the computer program is executed by the processor, it implements the formatted data voice program-controlled interaction method according to any one of claims 1-8.