Retail terminal man-machine interaction method, system and equipment based on voice recognition and medium

Through real-time voice recognition technology and named entity recognition, effective connection between voice input and business data in retail terminals is achieved, time-consuming and cumbersome problems in the entry and inventory process in the existing technology are solved, and human-computer interaction efficiency and system reliability are improved.

CN119917644APending Publication Date: 2025-05-02SHANDONG INSPUR DIGITAL BUSINESS TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510073585.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-02

AI Technical Summary

Technical Problem

The existing retail terminals have many and complicated products when entering and inventorying products, resulting in a time-consuming and cumbersome process, and lack efficient voice recognition technology to connect with business.

Method used

Through real-time speech recognition technology, voice is converted into text, and combined with named entity recognition, text vector conversion and part-of-speech annotation, the refinement of speech operations and effective docking of business data is achieved.

Benefits of technology

It simplifies a lot of manual operations, improves human-computer interaction efficiency, realizes effective connection between voice input and actual services, saves labor costs and improves the reliability and availability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119917644A_ABST
    Figure CN119917644A_ABST
Patent Text Reader

Abstract

The invention discloses a retail terminal man-machine interaction method, system and device based on voice recognition and a medium, belongs to the technical field of intelligent terminals and voice recognition, and aims to solve the technical problems of how to realize effective connection between voice input and actual services, simplify a large amount of manual operation processes, improve man-machine interaction efficiency and reduce the labor intensity of workers. According to the technical scheme, the method comprises the steps that real-time voice is obtained and converted into characters, the real-time voice is obtained through a bottom layer interface of real-time voice recognition, the real-time voice is converted into the characters through a voice recognition model, high-precision converted characters are used for correcting and outputting at the end of a voice sentence, and the output characters are provided with punctuations; named entity recognition: performing named entity recognition on the characters subjected to voice recognition through a named entity model, and distinguishing product words, modifiers and quantifiers which can be understood by the business; character conversion vectors; identifying part-of-speech tagging; and scheduling logic.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent terminals and speech recognition technology, and in particular to a retail terminal human-computer interaction method, system, device and medium based on speech recognition. Background Art

[0002] Speech recognition technology has also evolved from early statistical model methods (such as hidden Markov model HMM) to deep learning methods (such as recurrent neural network RNN ​​and attention model Transformer), to the current hot end-to-end speech recognition models such as CTC (Connectionist Temporal Classification), RNN-T (Recurrent Neural Network Transducer) and Transformer-based models.

[0003] The retail terminal refers to the end of the product sales channel, the final port where the product reaches the consumer to complete the transaction, and the place where the goods and consumers are displayed and traded face to face. Through this port and place, manufacturers and merchants sell products to consumers, complete the final transaction, and enter into substantive consumption; through this port, consumers buy the products they need and like.

[0004] With the development of artificial intelligence technology, there are more and more application scenarios for operating terminals through voice, motion, touch and other methods. When retail terminals are entering goods into warehouses, taking inventory, and making voice announcements, the process of selecting goods and entering quantities during entry / inventory is time-consuming and cumbersome due to the large number and variety of existing goods. Therefore, how to effectively connect voice entry with actual business, simplify a large number of manual operation processes, and improve the efficiency of human-computer interaction is a technical problem that needs to be solved urgently. Summary of the invention

[0005] The technical task of the present invention is to provide a retail terminal human-computer interaction method, system, device and medium based on voice recognition to solve the problem of how to effectively connect voice input with actual business, simplify a large number of manual operation processes, and improve the efficiency of human-computer interaction.

[0006] The technical task of the present invention is achieved in the following way: a retail terminal human-computer interaction method based on speech recognition, the method is as follows:

[0007] Obtain real-time speech and convert it into text: obtain real-time speech through the underlying interface of real-time speech recognition, convert it into text through the speech recognition model, and correct the output with high-precision transcription at the end of the speech sentence. The output text has punctuation.

[0008] Named entity recognition: The named entity model is used to identify the named entities of the text in speech recognition, and distinguish product words, modifiers, and quantifiers that can be understood by the business.

[0009] Text conversion vector: vectorize the core human-computer interaction operation words and business core words in the business, and vectorize the text data of speech recognition;

[0010] Part-of-speech tagging and recognition: tag the recognized text with parts of speech to achieve refined voice operation and improve the user experience;

[0011] Scheduling logic: After completing the voice input, the interface scheduling recognition is performed, and the parsed text information is displayed on the front end.

[0012] As a preference, in the process of acquiring real-time speech and converting it into text, high-concurrency multiple-channel speech-to-text requests are supported.

[0013] Preferably, the front end includes a functional interface for turning on the start and end buttons of the voice operation and a result confirmation interface for the voice input.

[0014] Preferably, in the process of acquiring real-time voice, a segmented voice recording method is adopted and the voice is stored.

[0015] A retail terminal human-computer interaction system based on speech recognition, the system comprising:

[0016] The real-time speech dictation module is used to obtain real-time speech through the underlying interface of real-time speech recognition, and to convert real-time speech into text through the speech recognition model, and to correct the output with high-precision transcription text at the end of the speech sentence, and the output text has punctuation;

[0017] The named entity recognition module is used to identify the named entities of the text in the speech recognition through the named entity model, and distinguish the product words, modifiers and quantifiers that the business can understand;

[0018] The text vector conversion module is used to vectorize the core human-computer interaction operation words and business core words in the business, and to vectorize the text data of speech recognition;

[0019] The part-of-speech tagging recognition module is used to tag the recognized text with parts of speech, so as to achieve the refinement of voice operation and improve the user experience;

[0020] The scheduling logic module is used to perform interface scheduling recognition after voice input is completed, and to display the parsed text information on the front end.

[0021] As a preferred option, the real-time voice dictation module supports high-concurrency multi-channel voice-to-text requests;

[0022] The real-time voice dictation module uses a segmented voice recording method and stores the voice.

[0023] Preferably, the front end includes a functional interface for turning on the start and end buttons of the voice operation and a result confirmation interface for the voice input.

[0024] Preferably, the working process of the system is as follows:

[0025] (1) The application terminal enters the voice input entrance of the function module by clicking a button;

[0026] (2) Click the record button to record the speech in segments and store the speech;

[0027] (3) transmitting the recorded voice to the real-time voice dictation module, which converts the voice into text;

[0028] (4) transmitting the converted text to a part-of-speech tagging recognition module, which performs part-of-speech tagging on the converted text;

[0029] (5) transmitting the converted text to a named entity recognition module, and the named entity recognition module annotates the converted text with named entities;

[0030] (6) transmitting the converted text to a text vector conversion module, and the text vector conversion module converts the converted text into a vector;

[0031] (7) Extracting the data of specific business from the business database and transmitting it to the text vector conversion module to realize the conversion of business data;

[0032] (8) Transmitting the marked parts of speech, named entities, and vector information to the scheduling logic module, where the voice and business data are aligned and queried;

[0033] (9) Update the data finally produced by the scheduling logic module to the terminal front end for interface display;

[0034] (10) The front end modifies and confirms the results displayed on the interface and completes the business.

[0035] An electronic device comprising: a memory and at least one processor;

[0036] Wherein, the memory stores a computer program;

[0037] The at least one processor executes the computer program stored in the memory, so that the at least one processor executes the above-mentioned retail terminal human-computer interaction method based on speech recognition.

[0038] A computer-readable storage medium stores a computer program, which can be executed by a processor to implement the above-mentioned retail terminal human-computer interaction method based on speech recognition.

[0039] The retail terminal human-computer interaction method, system, device and medium based on speech recognition of the present invention have the following advantages:

[0040] (1) The present invention assists the terminal software operation to realize the transformation of the software from manual control to voice control combined with manual confirmation operation, thus simplifying the software operation process;

[0041] (ii) The present invention is used to understand and convert text based on a speech recognition model through voice input to achieve agreed functions in a specific business environment;

[0042] (III) The present invention realizes the effective connection between voice input and actual business through voice commodity input, voice inventory counting, voice query and broadcast of business indicators, and users can quickly build applications to achieve the purpose of improving labor efficiency;

[0043] (IV) The present invention realizes the conversion from speech to text by accessing the speech recognition model, realizes semantic segmentation of text through the named entity model, maps the recognized language and text to the vector space, vectorizes the commodity information and operation information in the actual business, realizes the conversion from speech to commodity and speech to operation through vector-level comparison, and in the final application link, the voice input personnel are required to select and correct, thereby achieving the goal of simplifying the operation behavior and taking into account the functional availability;

[0044] (V) After configuring the above speech recognition environment, the present invention supports rapid deployment of speech recognition applications, thereby simplifying the original business that requires manual operation;

[0045] (VI) The present invention is initially considered to be applied in the areas of commodity warehousing, inventory counting, and voice broadcasting. The existing commodity categories are numerous and diverse, and the process of selecting the quantity of commodities to be entered during entry / inventory counting is time-consuming and cumbersome. After the introduction of voice entry technology, a large number of manual operations are simplified, and only the final confirmation link is retained, which greatly improves labor efficiency;

[0046] (VII) The present invention can realize the query of retail indicators through voice interaction and voice broadcast, such as real-time inventory query, daily sales query and other core business needs;

[0047] (VIII) The present invention is based on voice recognition technology and terminal business basic data to simplify the manual operation process, convert the cumbersome manual operation process into a voice interaction process, and save labor costs. The entire technical difficulty lies in the combination and matching of business data with the data entered by humans and machines, recognizing the entered voice as valid business data and business operations, and improving system reliability and availability. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] The present invention is further described below in conjunction with the accompanying drawings.

[0049] Attached Figure 1 This is a structural block diagram of a retail terminal human-computer interaction system based on speech recognition. DETAILED DESCRIPTION

[0050] The retail terminal human-computer interaction method, system, device and medium based on speech recognition of the present invention are described in detail below with reference to the accompanying drawings and specific embodiments.

[0051] Embodiment 1:

[0052] This embodiment provides a retail terminal human-computer interaction method based on speech recognition, and the method is specifically as follows:

[0053] S1. Obtain real-time speech and convert it into text: obtain real-time speech through the underlying interface of real-time speech recognition, convert it into text through the speech recognition model, and correct the output with high-precision transcription at the end of the speech sentence. The output text has punctuation.

[0054] S2. Named Entity Recognition: The named entity model is used to identify the named entities of the speech recognition text, and distinguish the product words, modifiers, and quantifiers that the business can understand;

[0055] S3, text conversion vector: vectorize the core human-computer interaction operation words and business core words in the business, and vectorize the text data of speech recognition;

[0056] S4, Part-of-speech tagging recognition: tag the recognized text with parts of speech to achieve refined voice operation and improve the user experience;

[0057] S5, Scheduling logic: After completing the voice input, perform interface scheduling recognition, and display the parsed text information on the front end.

[0058] In the process of acquiring real-time speech and converting it into text in step S1 of this embodiment, high-concurrency multi-channel speech-to-text requests are supported.

[0059] The front end in step S5 of this embodiment includes a function interface for turning on the start and end buttons of the voice operation and a result confirmation interface for the voice input.

[0060] In the process of acquiring real-time voice in step S1 of this embodiment, a segmented voice recording method is adopted and the voice is stored.

[0061] Embodiment 2:

[0062] As attached Figure 1 As shown, this embodiment provides a retail terminal human-computer interaction system based on speech recognition, the system comprising:

[0063] The real-time speech dictation module is used to obtain real-time speech through the underlying interface of real-time speech recognition, and to convert real-time speech into text through the speech recognition model, and to correct the output with high-precision transcription text at the end of the speech sentence, and the output text has punctuation;

[0064] The named entity recognition module is used to identify the named entities of the text in the speech recognition through the named entity model, and distinguish the product words, modifiers and quantifiers that the business can understand;

[0065] The text vector conversion module is used to vectorize the core human-computer interaction operation words and business core words in the business, and to vectorize the text data of speech recognition;

[0066] The part-of-speech tagging recognition module is used to tag the recognized text with parts of speech, so as to achieve the refinement of voice operation and improve the user experience;

[0067] The scheduling logic module is used to perform interface scheduling recognition after voice input is completed, and to display the parsed text information on the front end.

[0068] The real-time voice dictation module in this embodiment supports high-concurrency multi-channel voice-to-text requests.

[0069] The real-time voice dictation module in this embodiment adopts a segmented voice recording method and stores the voice.

[0070] The front end in this embodiment includes a function interface for turning on the start and end buttons of the voice operation and a result confirmation interface for the voice input.

[0071] The working process of the system is as follows:

[0072] (1) The application terminal enters the voice input entrance of the function module by clicking a button;

[0073] (2) Click the record button to record the speech in segments and store the speech;

[0074] (3) transmitting the recorded voice to the real-time voice dictation module, which converts the voice into text;

[0075] (4) transmitting the converted text to a part-of-speech tagging recognition module, which performs part-of-speech tagging on the converted text;

[0076] (5) transmitting the converted text to a named entity recognition module, and the named entity recognition module annotates the converted text with named entities;

[0077] (6) transmitting the converted text to a text vector conversion module, and the text vector conversion module converts the converted text into a vector;

[0078] (7) Extracting the data of specific business from the business database and transmitting it to the text vector conversion module to realize the conversion of business data;

[0079] (8) Transmitting the marked parts of speech, named entities, and vector information to the scheduling logic module, where the voice and business data are aligned and queried;

[0080] (9) Update the data finally produced by the scheduling logic module to the terminal front end for interface display;

[0081] (10) The front end modifies and confirms the results displayed on the interface and completes the business.

[0082] Embodiment 3:

[0083] This embodiment also provides an electronic device, including: a memory and a processor;

[0084] Wherein, the memory stores computer-executable instructions;

[0085] The processor executes the computer-executable instructions stored in the memory, so that the processor executes the retail terminal human-computer interaction method based on speech recognition in any embodiment of the present invention.

[0086] The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor may be a microprocessor or any conventional processor, etc.

[0087] The memory can be used to store computer programs and / or modules. The processor realizes various functions of the electronic device by running or executing the computer programs and / or modules stored in the memory, and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application required for a function, etc.; the data storage area can store data created according to the use of the terminal, etc. In addition, the memory can also include a high-speed random access memory, and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (SMC), a secure digital (SD) card, a flash memory card, at least one disk storage period, a flash memory device, or other volatile solid-state storage devices.

[0088] Embodiment 4:

[0089] This embodiment also provides a computer-readable storage medium, which stores a plurality of instructions, which are loaded by a processor, so that the processor executes the retail terminal human-computer interaction method based on speech recognition in any embodiment of the present invention. Specifically, a system or device equipped with a storage medium can be provided, on which a software program code that implements the functions of any of the above embodiments is stored, and a computer (or CPU or MPU) of the system or device reads and executes the program code stored in the storage medium.

[0090] In this case, the program code itself read from the storage medium can realize the function of any one of the above-mentioned embodiments, and thus the program code and the storage medium storing the program code constitute a part of the present invention.

[0091] The storage medium embodiments for providing the program code include a floppy disk, a hard disk, a magneto-optical disk, an optical disk (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RYM, DVD-RW, DVD+RW), a magnetic tape, a non-volatile memory card, and a ROM. Alternatively, the program code can be downloaded from a server computer via a communication network.

[0092] In addition, it should be clear that the functions of any of the above embodiments can be implemented not only by executing the program code read by the computer, but also by enabling an operating system operating on the computer to complete part or all of the actual operations based on instructions from the program code.

[0093] In addition, it can be understood that the program code read from the storage medium is written to a memory provided in an expansion board inserted into the computer or to a memory provided in an expansion unit connected to the computer, and then based on the instructions of the program code, a CPU installed on the expansion board or the expansion unit is enabled to perform part or all of the actual operations, thereby realizing the functions of any of the above-mentioned embodiments.

[0094] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein by equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A retail terminal human-computer interaction method based on speech recognition, characterized in that: The method is as follows: Obtain real-time speech and convert it into text: obtain real-time speech through the underlying interface of real-time speech recognition, convert it into text through the speech recognition model, and correct the output with high-precision transcription at the end of the speech sentence. The output text has punctuation. Named entity recognition: The named entity model is used to identify the named entities of the text in speech recognition, and distinguish product words, modifiers, and quantifiers that can be understood by the business. Text conversion vector: vectorize the core human-computer interaction operation words and business core words in the business, and vectorize the text data of speech recognition; Part-of-speech tagging and recognition: tag the recognized text with parts of speech to achieve the refinement of voice operation; Scheduling logic: After completing the voice input, the interface scheduling recognition is performed, and the parsed text information is displayed on the front end.

2. The retail terminal human-computer interaction method based on speech recognition according to claim 1, characterized in that: In the process of obtaining real-time voice and converting it into text, high-concurrency multi-channel voice-to-text requests are supported.

3. The retail terminal human-computer interaction method based on speech recognition according to claim 1, characterized in that: The front end includes a function interface for turning on the start and end buttons of voice operation and an interface for confirming the results of voice input.

4. The retail terminal human-computer interaction method based on speech recognition according to any one of claims 1 to 3, characterized in that: In the process of acquiring real-time voice, the voice is recorded in segments and stored.

5. A retail terminal human-computer interaction system based on speech recognition, characterized in that: The system includes: The real-time speech dictation module is used to obtain real-time speech through the underlying interface of real-time speech recognition, and to convert real-time speech into text through the speech recognition model, and to correct the output with high-precision transcription text at the end of the speech sentence, and the output text has punctuation; The named entity recognition module is used to identify the named entities of the text in the speech recognition through the named entity model, and distinguish the product words, modifiers and quantifiers that the business can understand; The text vector conversion module is used to vectorize the core human-computer interaction operation words and business core words in the business, and to vectorize the text data of speech recognition; The part-of-speech tagging recognition module is used to tag the recognized text with parts of speech to achieve the refinement of voice operation; The scheduling logic module is used to perform interface scheduling recognition after voice input is completed, and to display the parsed text information on the front end.

6. The retail terminal human-computer interaction system based on speech recognition according to claim 5, characterized in that: The real-time voice dictation module supports high-concurrency multi-channel voice-to-text requests; The real-time voice dictation module uses a segmented voice recording method and stores the voice.

7. The retail terminal human-computer interaction system based on speech recognition according to claim 5, characterized in that: The front end includes a function interface for turning on the start and end buttons of voice operation and an interface for confirming the results of voice input.

8. The retail terminal human-computer interaction system based on speech recognition according to any one of claims 5 to 7, characterized in that: The working process of the system is as follows: (1) The application terminal enters the voice input entrance of the function module by clicking a button; (2) Click the record button to record the speech in segments and store the speech; (3) transmitting the recorded voice to the real-time voice dictation module, which converts the voice into text; (4) transmitting the converted text to a part-of-speech tagging recognition module, which performs part-of-speech tagging on the converted text; (5) transmitting the converted text to a named entity recognition module, and the named entity recognition module annotates the converted text with named entities; (6) transmitting the converted text to a text vector conversion module, and the text vector conversion module converts the converted text into a vector; (7) Extracting the data of specific business from the business database and transmitting it to the text vector conversion module to realize the conversion of business data; (8) Transmitting the marked parts of speech, named entities, and vector information to the scheduling logic module, where the voice and business data are aligned and queried; (9) Update the data finally produced by the scheduling logic module to the terminal front end for interface display; (10) The front end modifies and confirms the results displayed on the interface and completes the business.

9. An electronic device, characterized in that: include: memory and at least one processor; Wherein, the memory stores a computer program; The at least one processor executes the computer program stored in the memory, so that the at least one processor executes the retail terminal human-computer interaction method based on speech recognition according to any one of claims 1 to 4.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which can be executed by a processor to implement the retail terminal human-computer interaction method based on speech recognition as claimed in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Self-service shopping method, device and equipment and computer readable storage medium

    CN110377812A

  • Service function module pushing method and device

    CN113971949A

  • Barrier-free access method, device and system for application program

    CN117577103A

  • System function page intelligent scheduling method, system, equipment and medium based on voice recognition

    CN119170008A

Cited By

  • Intelligent terminal task scheduling and executing system based on large model

    CN121029346A