Human-computer interaction method, device, electronic device and storage medium for self-service device

By combining the processing of user voice signals with image acquisition modules, the self-service device can accurately identify the user's intentions and output response statements that are consistent with the intentions, solving the problems of inaccurate voice recognition and illegal recording deception in existing technologies, and improving the efficiency and security of human-computer interaction.

CN117037780BActive Publication Date: 2025-09-12BEIJING INTEHEL TECH DEV CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310761359.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-27
Publication Date
2025-09-12
Estimated Expiration
2043-06-27

AI Technical Summary

Technical Problem

The human-computer interaction methods of existing self-service devices are easily interfered by noise when recognizing user voice, cannot accurately identify user intentions, and may be deceived by illegal recordings, resulting in chaotic interactions and reduced user experience.

Method used

By obtaining the user's voice signal, processing and recognizing it to obtain text information, searching the list of text response sentences in the database, and converting them into voice output one by one, combining with the image acquisition module to obtain the user's face information, matching the voice signal to improve recognition accuracy, building a voice-to-text conversion model and database, and realizing the screening and sorting of voice word slots.

Benefits of technology

It improves the recognition accuracy and interaction efficiency of self-service devices for user voice signals, ensures that the output response sentences are consistent with user intentions, and improves the accuracy and security of human-computer interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117037780B_ABST
    Figure CN117037780B_ABST
Patent Text Reader

Abstract

The present application discloses a human-computer interaction method, device, electronic device and storage medium for a self-service device. The method includes: obtaining a user voice signal and processing and identifying it to obtain corresponding text information; searching a list of text response sentences adapted to the text information in a user database based on the text information corresponding to the voice signal; outputting the text response sentences in the text response sentence list to a text-to-speech conversion list in order; converting the text response sentences one by one into voice response sentences and outputting them, and then obtaining the user's response information to the voice response sentences. The present application can accurately identify the voice signals of the interacting person and the user, and output the corresponding response sentences based on the user's voice signal, and then control the execution of the corresponding instructions based on the user's response to the response sentences. The interaction method is more efficient, the information acquisition is more accurate, and the instructions executed by the self-service device are more in line with the user's intentions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to the field of human-computer interaction technology, and in particular to a human-computer interaction method, device, electronic device, and storage medium for a self-service device. Background Art

[0002] With the improvement of medical equipment automation technology, the automated storage, identification, and distribution of some medical equipment have also become possible, such as the self-service surgical gown dispensers and surgical shoe dispensers in hospitals, as well as self-service vending machines and self-service equipment. In the existing technology, during human-computer interaction, external sounds are generally obtained through microphones, and then the sounds are converted into text, and then the corresponding control instructions are obtained to control the self-service equipment to perform corresponding actions. However, since the sound signals collected by the microphones may contain noise, for example, when multiple people speak at the same time, the microphones may obtain the sound signals of multiple people at the same time, thereby collecting the sound information of different people at the same time, resulting in the inability to recognize the user's true semantics. For example, some illegal personnel may use pre-recorded audio to simulate certain sound information, or fail to respond effectively to the replies of the self-service equipment, resulting in the inability to identify whether it is a real human-computer interaction, causing the human-computer interaction to enter a disordered and chaotic state. All of these lead to the inability of self-service equipment to truly understand the user's intentions during human-computer interaction, thereby reducing the user experience. Summary of the Invention

[0003] In view of the above-mentioned defects or deficiencies in the prior art, it is desired to provide a human-computer interaction method, device, electronic device and storage medium for a self-service device that can meet the needs of this field.

[0004] According to one aspect of an embodiment of the present invention, an embodiment of the present application provides a human-computer interaction method for a self-service device, the method comprising:

[0005] Acquire a user voice signal, process and recognize the user voice signal, and obtain text information corresponding to the voice signal, wherein the text information includes a description of the user's voice intention and a voice word slot;

[0006] Searching a user database for a list of textual response sentences that are compatible with the textual response corresponding to the voice signal based on the textual response corresponding to the voice signal;

[0007] Outputting the text reply sentences in the text reply sentence list to the text-to-speech conversion list in order;

[0008] After converting the text reply sentences in the text-to-speech conversion list into voice reply sentences one by one and outputting them, obtaining the user's response information to the voice reply sentences;

[0009] If the user's response information to the voice reply statement is yes, the self-service device executes the control instruction corresponding to the voice reply statement;

[0010] If the user's response information to the voice reply statement is no, outputting the next voice reply statement until the user's response information to the voice reply statement is yes;

[0011] If the user's response information to the last voice reply sentence in the text-to-speech conversion list is no, a prompt is given to re-acquire the user voice signal.

[0012] In another embodiment, before acquiring the user voice signal, the method further includes:

[0013] The image acquisition module of the self-service device obtains facial information of one or more users within a set distance;

[0014] Acquire mouth image change information in the one or more user facial information based on the one or more user facial information;

[0015] Acquire one or more user voice signals based on the mouth image change information in the one or more user facial information, wherein the user voice signals include user voices corresponding to the mouth image changes in the one or more user facial information;

[0016] Obtaining matching information between the voice signal and the user's facial information based on the one or more voice signals and mouth image change information in one or more user's facial information;

[0017] According to the matching information between the voice signal and the user face information, one of the one or more voice signals is obtained according to the set priority as the obtained user voice signal.

[0018] In another embodiment, obtaining mouth image change information in the one or more user facial information based on the one or more user facial information includes:

[0019] Based on the image information collected by the image collection device of the self-service device within a set distance, one or more facial images in the collected image information are obtained according to the face recognition model;

[0020] performing image tracking on one or more of the facial images, respectively, to obtain a facial center point in each of the facial images, the facial center point being a reference point for movement of the facial image in three-dimensional space, and determining motion information of other parts of the facial image by determining a motion trajectory of the facial center point in three-dimensional space;

[0021] According to the movement information of the face center point in the face image in the three-dimensional space, the mouth image change information in the face image is obtained.

[0022] In another embodiment, before acquiring the user voice signal, the method further includes:

[0023] Build a speech-to-text conversion model;

[0024] Based on the speech-to-text conversion model, a speech-to-text conversion database and a semantic text database corresponding to fuzzy speech are constructed and updated in real time;

[0025] Obtaining the speech word slot after speech-to-text conversion based on the speech-to-text conversion database, the semantic text database corresponding to the fuzzy speech, and the speech word slot extraction rules after speech-to-text conversion;

[0026] According to the acquired speech word slot, one or more text reply sentences corresponding to the speech word slot are set;

[0027] A user database is constructed based on one or more text response sentences corresponding to all the voice word slots.

[0028] In another embodiment, the acquiring of a user voice signal, processing and recognizing the user voice signal, and obtaining text information corresponding to the voice signal includes:

[0029] Acquire a user voice signal input by the user through a speech recognition application program interface, and convert the user voice signal input by the user into user input text information;

[0030] According to the user input text information, the user input text information is converted into speech word slot information, and the speech word slot information uses a word vector representation method to perform speech word slot extraction on the user input text information to obtain the semantics of the user input text information;

[0031] According to the voice word slot information converted from the text information input by the user, text information corresponding to the voice signal recognized by the self-service device is obtained.

[0032] In another embodiment, searching a user database for a list of textual response statements adapted to the textual information corresponding to the voice signal based on the textual information corresponding to the voice signal includes:

[0033] Obtaining a speech word slot in the text information according to the text information corresponding to the speech signal;

[0034] Searching, based on the one or more voice word slots in the text message, for one or more textual response sentences corresponding to the one or more voice word slots in the user database;

[0035] Based on the intention description of the user's voice, according to the set rules, one or more text response sentences corresponding to one or more voice word slots are associated to form one or more text response sentences adapted to the text information corresponding to the voice signal;

[0036] The one or more text reply statements are sorted, and the sorted one or more text reply statements are added to the text reply statement list in order.

[0037] In another embodiment, after converting the text reply sentences in the text-to-speech conversion list into voice reply sentences one by one and outputting them, obtaining user response information to the voice reply sentences includes:

[0038] Obtaining the first text reply sentence in the list sequence of the text-to-speech conversion list;

[0039] After converting the first text response sentence into a first voice response sentence, outputting the first voice response sentence, deleting the first text response sentence from the text-to-speech conversion list, using the next text response sentence in the list sequence of the text-to-speech conversion list as the first voice response sentence, and updating the text-to-speech conversion list;

[0040] Obtain the user's response information to the first voice reply statement;

[0041] If the user's response information to the first voice reply statement is yes, the control executes the control instruction pointed to by the first voice reply statement and deletes the updated text-to-speech conversion list;

[0042] If the user's response information to the first voice response statement is no, the first voice response statement converted from the first text response statement in the updated text-to-speech conversion list is output, and the first text response statement is deleted from the updated text-to-speech conversion list, and the next text response statement in the list sequence of the updated text-to-speech conversion list is used as the first voice response statement, and the text-to-speech conversion list is updated again until the user's response information to the first voice response statement is yes, and the control instruction corresponding to the first voice response statement is obtained.

[0043] According to another aspect of an embodiment of the present invention, a human-computer interaction device for a self-service device is disclosed, the device comprising:

[0044] An acquisition module is used to acquire a user voice signal, process and identify the user voice signal, and obtain text information corresponding to the voice signal, wherein the text information includes a description of the user's voice intention and a voice word slot;

[0045] a processing module configured to search a user database for a list of textual response sentences that are compatible with the textual response sentences corresponding to the voice signal based on the textual response sentences corresponding to the voice signal; output the textual response sentences in the list of textual response sentences in order to a text-to-speech conversion list; convert the textual response sentences in the list of textual response sentences into voice response sentences one by one and output them, and obtain user response information to the voice response sentences;

[0046] The judgment module is configured such that if the user's response to the voice reply statement is yes, the self-service device executes the control instruction corresponding to the voice reply statement; if the user's response to the voice reply statement is no, the next voice reply statement is output until the user's response to the voice reply statement is yes; if the user's response to the last voice reply statement in the text-to-speech conversion list is no, a prompt is given to reacquire the user's voice signal.

[0047] Based on another aspect of an embodiment of the present invention, an electronic device is disclosed, which includes one or more processors and a memory, wherein the memory is used to store one or more programs; when the one or more programs are executed by the processor, the processor implements the human-computer interaction method of the self-service device provided by each embodiment of the present invention.

[0048] According to another aspect of the embodiments of the present invention, a computer-readable storage medium storing a computer program is disclosed. When the computer program is executed, the human-computer interaction method of the self-service device provided by each embodiment of the present invention is implemented.

[0049] In an embodiment of the present application, by acquiring a user voice signal and processing and identifying the user voice signal, text information corresponding to the voice signal is obtained; based on the text information corresponding to the voice signal, a list of text reply statements adapted to the text information corresponding to the voice signal in the user database is searched; the text reply statements in the text reply statement list are output sequentially to a text-to-speech conversion list; the text reply statements in the text-to-speech conversion list are converted one by one into voice reply statements and output, and then the user's response information to the voice reply statement is obtained. The present application can accurately identify the voice signals of the interacting person and the user, and output the corresponding reply statement according to the user's voice signal, and then control the execution of the corresponding instruction according to the user's response to the reply statement. The interaction method is more efficient, the information acquisition is more accurate, and the instructions executed by the self-service device are more in line with the user's intention. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Other features, objects and advantages of the present application will become more apparent upon reading the detailed description of non-limiting embodiments made with reference to the following drawings:

[0051] Figure 1 This is a flow chart of a human-computer interaction method for a self-service device provided by an embodiment of the present application;

[0052] Figure 2 A schematic diagram of the structure of a human-computer interaction device of a self-service device provided in one embodiment of the present application;

[0053] Figure 3 This is a diagram of the internal structure of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0054] The present application will be further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are merely for the purpose of explaining the relevant invention and are not intended to limit the invention. It should also be noted that, for ease of description, only portions relevant to the invention are shown in the accompanying drawings.

[0055] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0056] Please refer to Figure 1 , which shows an exemplary process of the human-computer interaction method of the self-service device of the embodiment of the present application.

[0057] like Figure 1 As shown, in step 110, a user voice signal is acquired, and the user voice signal is processed and recognized to obtain text information corresponding to the voice signal. The text information includes an intention description of the user voice and a voice word slot.

[0058] Specifically, the user voice signal can be obtained by setting up a microphone on the self-service device for collection. In other embodiments, other Internet devices, such as mobile phones and tablets, can be used to establish a connection with the self-service device, and then human-computer interaction can be achieved through the app terminal of the self-service device installed on these Internet devices. At this time, human-computer interaction can be achieved through the Internet device's own voice recognition, voice-to-text conversion, image recognition and other functions. The embodiments of this application mainly involve obtaining user voice signals through the microphone of the self-service device and realizing human-computer interaction through the self-service device itself.

[0059] Specifically, after obtaining the user voice signal, the user voice signal can be processed and identified. There are mature technologies in the existing technology for voice signal processing, voice signal recognition, and voice signal conversion to obtain text information corresponding to the voice signal, such as iFlytek's related software and hardware, and Tencent WeChat's voice-to-text conversion function, etc., which can be used as the technical means adopted in the embodiments of this application to process and identify the user voice signal and obtain the text information corresponding to the voice signal. The specific implementation process is not described here.

[0060] Specifically, when processing and identifying the user voice signal and obtaining the text information corresponding to the voice signal, what needs to be paid attention to is the intention description and voice word slot of the user voice. The intention description of the user voice is that after the user voice signal is converted into text information, its semantics conforms to the Chinese expression habits. The voice word slot is the keyword in the user voice signal. By obtaining the intention description of the user voice signal and the voice word slot of the voice signal, the corresponding answer can be obtained through the intention description and the voice word slot.

[0061] Specifically, in one embodiment of the present application, the acquiring of a user voice signal, processing and recognizing the user voice signal, and obtaining text information corresponding to the voice signal includes:

[0062] The user voice signal input by the user is obtained through the application program interface of speech recognition, and the user voice signal input by the user is converted into text information input by the user; specifically, the application program interface first pre-defines some functions, and provides the application program and the developer with the ability to use a set of routines based on certain software or hardware without accessing the source code. The application program interface of speech recognition converts the user voice signal collected by the speech acquisition device into a character string, and then converts the character string into text information in text form, so that the system can recognize it.

[0063] Based on the text information input by the user, the text information input by the user is converted into speech word slot information, and the speech word slot information uses a word vector representation method to extract speech word slots from the text information input by the user to obtain the semantics of the text information input by the user; specifically, by obtaining the speech word slot information, the word slot and the word slot type, the corresponding position of the word slot in the text, etc. can be obtained.

[0064] According to the voice word slot information converted from the text information input by the user, text information corresponding to the voice signal recognized by the self-service device is obtained.

[0065] In step 120, based on the text information corresponding to the voice signal, a list of text reply sentences adapted to the text information corresponding to the voice signal is searched in a user database.

[0066] Specifically, after converting the user's voice signal into text information, the intention description and voice word slot of the user's voice in the text information are obtained. Then, all reply statements corresponding to the current voice word slot can be searched in the user database based on the voice word slot, and then these reply statements can be filtered according to the intention description of the user's voice to obtain reply statements that are more in line with the intention description of the user's voice.

[0067] Specifically, in one embodiment of the present application, searching a user database for a list of textual response statements adapted to the textual information corresponding to the voice signal based on the textual information corresponding to the voice signal includes:

[0068] Obtaining a speech word slot in the text information according to the text information corresponding to the speech signal;

[0069] Based on one or more voice word slots in the text information, search for one or more text reply sentences corresponding to one or more voice word slots in the user database; specifically, there may be multiple voice word slots in the text information converted from the user's voice information. Therefore, by obtaining multiple voice word slots in the text information, more abundant text reply sentences can be found in the user database. It should be noted that these text reply sentences are all text reply sentences for a certain voice word slot. It does not consider whether these text reply sentences express the same meaning, or whether they are a reply to the user's voice information as a whole. These text reply sentences may not be related to each other, and are just independent sentences, words or words.

[0070] Based on the intention description of the user voice and in accordance with the set rules, one or more text reply sentences corresponding to one or more voice word slots are associated to form one or more text reply sentences adapted to the text information corresponding to the voice signal; specifically, after obtaining the text reply sentences corresponding to one or more voice word slots, these text reply sentences can be filtered through the intention description of the user voice to filter out some text reply sentences that meet the user intention description, and then the text reply sentences corresponding to these voice word slots are associated to make the associated text reply sentences become a piece of text, and the description method of the text complies with the description rules of Chinese and whether it conforms to Chinese semantics. For example, three voice word slots are obtained from the text information corresponding to the user's voice signal, and each voice word slot has three corresponding text reply sentences. Then, combined with the intention description of the user's voice, one corresponding text reply sentence is extracted from each of the three voice word slots for association to see whether the associated text reply sentence conforms to the Chinese description rules and whether the associated text reply sentence conforms to the Chinese semantics. If the conditions are not met, the associated text reply sentence is deleted. If it conforms to the Chinese description rules and the Chinese semantics, one or more associated text reply sentences are filtered through the intention description of the user's voice, so that the one or more associated text reply sentences after filtering conform to the intention description of the user's voice.

[0071] The one or more text reply statements are sorted, and the sorted one or more text reply statements are added to the text reply statement list in order. Specifically, since the one or more associated text reply statements have different reply directions or reply methods to the intended description of the user's voice, they may all be consistent with the reply to the user's voice information. However, in combination with the intended description of the user's voice and the different application fields it is aimed at, the one or more text reply statements can be sorted, for example, by prioritizing the text reply statements of a certain voice word slot, or by prioritizing the conciseness of the text reply statements, or by randomly arranging the text reply statements, etc. After sorting the one or more text reply statements, the sorted one or more text reply statements are added to the text reply statement list in order.

[0072] In step 130 , the text response sentences in the text response sentence list are output to the text-to-speech conversion list in sequence.

[0073] Specifically, the text-to-speech conversion list converts text response sentences into speech output. In the embodiments of the present application, speech synthesis is a technology that produces artificial speech through mechanical and electronic methods. It is a technology that converts text information generated by the computer itself or externally input into understandable and fluent Chinese spoken output.

[0074] In step 140, the text reply sentences in the text-to-speech conversion list are converted into voice reply sentences one by one and output, and then the user's response information to the voice reply sentences is obtained.

[0075] Specifically, the text response sentences in the text language conversion list are converted in the order of the list. In order to improve efficiency and save processor work efficiency, the text response sentences in the text-to-speech conversion list of this application are converted into voice response sentences in the order of the list. It first converts the first text response sentence in the list into a voice response sentence, and then outputs the converted voice response sentence. At this time, the text response sentence in the list is deleted, and the next text response sentence is arranged to the first position of the list. If the voice response sentence output at this time obtains the user's affirmative response, it means that the voice response sentence output at this time is a voice response sentence that meets the user's goal, and the result of human-computer interaction has been satisfactory to the user. The self-service device can execute the instruction corresponding to the voice response sentence to The response action is executed, and the human-computer interaction between the current user voice signal and the self-service device is completed. The user controls the self-service device through the voice signal to complete the corresponding work in line with the user's intention. If the voice response statement output at this time does not get a positive response from the user, it means that the voice response statement output at this time does not meet the user's actual intention, then the voice response statement must be discarded and deleted, and the voice response statement corresponding to the next text response statement must be output until the user gets a positive response. If all the text response statements in the list are converted into voice response statements, and all the voice response statements do not get a positive response from the user, then it is necessary to re-acquire the user voice signal, and re-execute the conversion and recognition of the user voice signal, and the search for the corresponding text response statement.

[0076] Specifically, if the user's response information to the voice reply statement is yes, the self-service device executes the control instruction corresponding to the voice reply statement;

[0077] If the user's response information to the voice reply statement is no, outputting the next voice reply statement until the user's response information to the voice reply statement is yes;

[0078] If the user's response information to the last voice reply sentence in the text-to-speech conversion list is no, a prompt is given to re-acquire the user voice signal.

[0079] Specifically, in one embodiment of the present application, after converting the text reply sentences in the text-to-speech conversion list into voice reply sentences one by one and outputting them, obtaining the user's response information to the voice reply sentences includes:

[0080] Obtaining the first text reply sentence in the list sequence of the text-to-speech conversion list;

[0081] After converting the first text response sentence into a first voice response sentence, outputting the first voice response sentence, deleting the first text response sentence from the text-to-speech conversion list, using the next text response sentence in the list sequence of the text-to-speech conversion list as the first voice response sentence, and updating the text-to-speech conversion list;

[0082] Obtain the user's response information to the first voice reply statement;

[0083] If the user's response information to the first voice reply statement is yes, the control executes the control instruction pointed to by the first voice reply statement and deletes the updated text-to-speech conversion list;

[0084] If the user's response information to the first voice response statement is no, the first voice response statement converted from the first text response statement in the updated text-to-speech conversion list is output, and the first text response statement is deleted from the updated text-to-speech conversion list, and the next text response statement in the list sequence of the updated text-to-speech conversion list is used as the first voice response statement, and the text-to-speech conversion list is updated again until the user's response information to the first voice response statement is yes, and the control instruction corresponding to the first voice response statement is obtained.

[0085] In one embodiment of the present application, before obtaining the user voice signal, the human-computer interaction method of the self-service device further includes:

[0086] The image acquisition module of the self-service device obtains facial information of one or more users within a set distance; specifically, the image acquisition module of the self-service device uses all facial images within its acquisition range as user facial information, and each facial image corresponds to one user facial information. The user facial information collected by the image acquisition module of the self-service device in the embodiment of the present application is a dynamic facial image.

[0087] Based on the one or more user facial information, obtain the mouth image change information in the one or more user facial information; specifically, in collecting the user facial information, based on each tracked user facial image, monitor the mouth image changes of the user facial image in the user facial information in real time. Since the mouth image change can determine that there is a sound signal at this time, it is possible to determine which user facial information the sound collected at this time corresponds to based on the mouth image change in the user facial information. For example, among multiple user facial information, the sound signal can be matched with the user facial information based on the timing of the mouth image change and the length of the change, the timing of the sound signal, the length of the sound signal, the timbre, scale, tone, etc. of the sound signal, etc.

[0088] Acquire one or more user voice signals based on the mouth image change information in the one or more user facial information, wherein the user voice signals include user voices corresponding to the mouth image changes in the one or more user facial information;

[0089] Obtaining matching information between the voice signal and the user's facial information based on the one or more voice signals and mouth image change information in one or more user's facial information;

[0090] According to the matching information between the voice signal and the user face information, one of the one or more voice signals is obtained according to the set priority as the obtained user voice signal.

[0091] Specifically, in one embodiment of the present application, obtaining mouth image change information in the one or more user facial information based on the one or more user facial information includes:

[0092] Based on the image information collected by the image collection device of the self-service device within a set distance, one or more facial images in the collected image information are obtained according to the face recognition model;

[0093] performing image tracking on one or more of the facial images, respectively, to obtain a facial center point in each of the facial images, the facial center point being a reference point for movement of the facial image in three-dimensional space, and determining motion information of other parts of the facial image by determining a motion trajectory of the facial center point in three-dimensional space;

[0094] Based on the motion information of the facial center point in the facial image in three-dimensional space, mouth image change information in the facial image is obtained. Specifically, the facial center point corresponds to a three-dimensional coordinate in the three-dimensional space, and the mouth image change information in the facial image is determined by determining the combined displacement vector of the facial center point's motion trajectory along the X, Y, and Z axes in the three-dimensional coordinate system.

[0095] Specifically, in one embodiment of the present application, obtaining matching information between the voice signal and the user facial information based on the one or more voice signals and mouth image change information in one or more user facial information includes:

[0096] Acquiring voice parameter information of one or more voice signals, wherein the voice parameter information includes timbre, scale, loudness, pitch, and acoustic characteristics of the voice;

[0097] Acquire each independent voice signal of the one or more voice signals based on voice parameter information of the one or more voice signals;

[0098] Obtain the time node of voice transformation of each independent voice signal and the duration between two adjacent time nodes;

[0099] Obtaining mouth image transformation information in each independent user face information based on mouth image transformation information in one or more user face information;

[0100] Obtain the time nodes of mouth image transformation in each independent user's facial information, as well as the duration between two adjacent time nodes;

[0101] Query the time node of the voice transformation of each independent voice signal, and the time node of the mouth image transformation in each independent user face information that matches the duration between two adjacent time nodes, and the duration between two adjacent time nodes to obtain matching information between the voice signal and the user face information.

[0102] Specifically, in one embodiment of the present application, before obtaining the user voice signal, the human-computer interaction method of the self-service device further includes:

[0103] Build a speech-to-text conversion model;

[0104] Based on the speech-to-text conversion model, a speech-to-text conversion database and a semantic text database corresponding to fuzzy speech are constructed and updated in real time;

[0105] Obtaining the speech word slot after speech-to-text conversion based on the speech-to-text conversion database, the semantic text database corresponding to the fuzzy speech, and the speech word slot extraction rules after speech-to-text conversion;

[0106] According to the acquired speech word slot, one or more text reply sentences corresponding to the speech word slot are set;

[0107] A user database is constructed based on one or more text response sentences corresponding to all the voice word slots.

[0108] In an embodiment of the present application, a user voice signal is obtained by processing and identifying the user voice signal to obtain text information corresponding to the voice signal; based on the text information corresponding to the voice signal, a list of text reply statements adapted to the text information corresponding to the voice signal in the user database is searched; the text reply statements in the text reply statement list are output sequentially to a text-to-speech conversion list; the text reply statements in the text-to-speech conversion list are converted one by one into voice reply statements and output, and then the user's response information to the voice reply statement is obtained. The present application can accurately identify the voice signals of the interacting person and the user, and output the corresponding reply statement according to the user's voice signal, and then control the execution of the corresponding instruction according to the user's response to the reply statement. The interaction method is more efficient, the information acquisition is more accurate, and the instructions executed by the self-service device are more in line with the user's intention.

[0109] It should be understood that although Figure 1 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. In addition, Figure 1 At least part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.

[0110] Figure 2 This is a structural diagram of a human-computer interaction device of a self-service device provided by an embodiment of the present application. Figure 2 As shown, the human-computer interaction device of the self-service device includes:

[0111] Acquisition module, processing module, judgment module;

[0112] An acquisition module is used to acquire a user voice signal, process and identify the user voice signal, and obtain text information corresponding to the voice signal, wherein the text information includes a description of the user's voice intention and a voice word slot;

[0113] a processing module configured to search a user database for a list of textual response sentences that are compatible with the textual response sentences corresponding to the voice signal based on the textual response sentences corresponding to the voice signal; output the textual response sentences in the list of textual response sentences in order to a text-to-speech conversion list; convert the textual response sentences in the list of textual response sentences into voice response sentences one by one and output them, and obtain user response information to the voice response sentences;

[0114] The judgment module is configured such that if the user's response to the voice reply statement is yes, the self-service device executes the control instruction corresponding to the voice reply statement; if the user's response to the voice reply statement is no, the next voice reply statement is output until the user's response to the voice reply statement is yes; if the user's response to the last voice reply statement in the text-to-speech conversion list is no, a prompt is given to reacquire the user's voice signal.

[0115] Specifically, in another embodiment of the present application, the acquisition module is used by the image acquisition module of the self-service device to acquire one or more user facial information within a set distance; based on the one or more user facial information, obtain the mouth image change information in the one or more user facial information; based on the mouth image change information in the one or more user facial information, obtain one or more user voice signals, and the user voice signal includes the user voice of the mouth image change in one or more user facial information; based on the one or more voice signals and the mouth image change information in one or more user facial information, obtain the matching information between the voice signal and the user facial information; based on the matching information between the voice signal and the user facial information, obtain one of the one or more voice signals as the acquired user voice signal according to the set priority.

[0116] Specifically, in another embodiment of the present application, the acquisition module is used to acquire one or more facial images in the collected image information based on the image information collected by the image acquisition device of the self-service device within a set distance according to the face recognition model; perform image tracking on one or more of the facial images respectively, and obtain the facial center point in each of the facial images respectively, the facial center point is the reference base point for the movement of the facial image in three-dimensional space, and the motion information of other parts of the facial image is determined by judging the movement trajectory of the facial center point in the three-dimensional space; based on the movement information of the facial center point in the facial image in the three-dimensional space, the mouth image change information in the facial image is obtained.

[0117] Specifically, in another embodiment of the present application, the acquisition module is used to construct a speech-to-text conversion model; based on the speech-to-text conversion model, a speech-to-text conversion database and a semantic text database corresponding to the fuzzy speech are constructed and updated in real time; based on the speech-to-text conversion database and the semantic text database corresponding to the fuzzy speech, the speech word slot after speech-to-text conversion is extracted according to the speech word slot extraction rules; based on the acquired speech word slot, one or more text reply sentences corresponding to the speech word slot are set; based on the one or more text reply sentences corresponding to all the speech word slots, a user database is constructed.

[0118] Specifically, in another embodiment of the present application, the processing module is used to obtain a user voice signal input by a user through an application program interface of voice recognition, and convert the user voice signal input by the user into user-input text information; based on the user-input text information, the user-input text information is converted into voice word slot information, and the voice word slot information uses a word vector representation method to perform voice word slot extraction on the user-input text information to obtain the semantics of the user-input text information; based on the voice word slot information converted from the user-input text information, the text information corresponding to the voice signal recognized by the self-service device is obtained.

[0119] Specifically, in another embodiment of the present application, the processing module is used to obtain the voice word slots in the text information corresponding to the voice signal; based on the one or more voice word slots in the text information, search for one or more text reply sentences corresponding to one or more of the voice word slots in the user database; based on the intention description of the user voice, according to the set rules, associate the one or more text reply sentences corresponding to one or more of the voice word slots to form one or more text reply sentences adapted to the text information corresponding to the voice signal; sort the one or more text reply sentences, and add the sorted one or more text reply sentences to the text reply sentence list in order.

[0120] Specifically, in another embodiment of the present application, the judgment module is used to obtain the first text response sentence in the list sequence of the text-to-speech conversion list; after converting the first text response sentence into the first voice response sentence, the first voice response sentence is output, and the first text response sentence is deleted from the text-to-speech conversion list, the next text response sentence in the list sequence of the text-to-speech conversion list is used as the first voice response sentence, and the text-to-speech conversion list is updated; obtain the user's response information to the first voice response sentence; if the user's response information to the first voice response sentence is yes, then control the execution of the control instruction pointed to by the first voice response sentence, and delete the updated text-to-speech conversion list; if the user's response information to the first voice response sentence is no, then output the first voice response sentence converted from the first text response sentence in the updated text-to-speech conversion list, delete the first text response sentence from the updated text-to-speech conversion list, use the next text response sentence in the list sequence of the updated text-to-speech conversion list as the first voice response sentence, and update the text-to-speech conversion list again, until the user's response information to the first voice response sentence is yes, and the control instruction corresponding to the first voice response sentence is obtained.

[0121] In an embodiment of the present application, a user voice signal is acquired through an acquisition module, and the user voice signal is processed and identified to obtain text information corresponding to the voice signal; the processing module searches for a list of text reply statements that are adapted to the text information corresponding to the voice signal in the user database based on the text information corresponding to the voice signal; the text reply statements in the text reply statement list are output sequentially to a text-to-speech conversion list; the text reply statements in the text-to-speech conversion list are converted one by one into voice reply statements and output, and then the user's response information to the voice reply statement is acquired. The present application can accurately identify the voice signals of the interacting person and the user, and output the corresponding reply statement according to the user's voice signal, and then control the execution of the corresponding instruction according to the user's response to the reply statement. The interaction method is more efficient, the information acquisition is more accurate, and the instructions executed by the self-service device are more in line with the user's intention.

[0122] The specific definition of the human-computer interaction device of the self-service device can be found in the definition of the human-computer interaction method of the self-service device above and will not be repeated here. The various modules in the human-computer interaction device of the self-service device described above can be implemented in whole or in part through software, hardware, or a combination thereof. Each of the above modules can be embedded in or independent of the processor of the computer device in hardware form, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each of the above modules.

[0123] In particular, according to the embodiments of the present disclosure, Figure 3 As shown, the present invention discloses an electronic device, which includes one or more processors and a memory, wherein the memory is used to store one or more programs; when the one or more programs are executed by the processor, the processor implements the human-computer interaction method of the self-service device described in an embodiment of the present invention.

[0124] In particular, according to embodiments of the present disclosure, the human-computer interaction method for a self-service device described in any of the above embodiments can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product comprising a computer program tangibly embodied on a machine-readable medium, the computer program containing program code for executing the human-computer interaction method for a self-service device. In such embodiments, the computer program can be downloaded and installed from a network via a communication component and / or installed from removable media.

[0125] The one or more programs are stored in a read-only memory (ROM) or a random access memory (RAM) to perform various appropriate actions and processes. The RAM contains software programs that the server uses to perform its services, as well as various programs and data required for vehicle driving operations. The server, its controlled hardware devices, the ROM, and the RAM are connected to each other via a bus, and various input / output interfaces are also connected to the bus.

[0126] The following components are connected to the input / output interface: an input section including a keyboard, mouse, and the like; an output section including a cathode ray tube (CRT), a liquid crystal display (LCD), and speakers; and a communication section including a network interface card (NIC) such as a LAN card and a modem. The communication section performs communication processing via a network such as the Internet. A drive is also connected to the input / output interface as needed. Removable media such as magnetic disks, optical disks, magneto-optical disks, and semiconductor memories are installed in the drive as needed, so that computer programs read from the media can be installed into the memory as needed.

[0127] In particular, according to embodiments of the present disclosure, the human-computer interaction method for a self-service device described in any of the above embodiments can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product comprising a computer program tangibly embodied on a machine-readable medium, the computer program containing program code for executing the human-computer interaction method for a self-service device. In such embodiments, the computer program can be downloaded and installed from a network via a communication component and / or installed from removable media.

[0128] The units or modules involved in the embodiments described in this application may be implemented by software or hardware. The units or modules described may also be provided in a processor. The names of these units or modules do not, in certain circumstances, constitute limitations on the units or modules themselves.

[0129] The above description is merely a preferred embodiment of the present application and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of the invention herein is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also encompasses other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the inventive concept. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features having similar functions disclosed in this application.

Claims

1. A human-computer interaction method for a self-service device, characterized in that: The method comprises: Acquire a user voice signal, process and recognize the user voice signal, and obtain text information corresponding to the voice signal, wherein the text information includes a description of the user's voice intention and a voice word slot; Searching a user database for a list of textual response sentences that are compatible with the textual response corresponding to the voice signal based on the textual response corresponding to the voice signal; Outputting the text reply sentences in the text reply sentence list to the text-to-speech conversion list in order; After converting the text reply sentences in the text-to-speech conversion list into voice reply sentences one by one and outputting them, obtaining the user's response information to the voice reply sentences; If the user's response information to the voice reply statement is yes, the self-service device executes the control instruction corresponding to the voice reply statement; If the user's response information to the voice reply statement is no, outputting the next voice reply statement until the user's response information to the voice reply statement is yes; If the user's response to the last voice reply statement in the text-to-speech conversion list is negative, prompting the user to reacquire the voice signal; The step of searching a user database for a list of textual response statements adapted to the textual information corresponding to the voice signal based on the textual information corresponding to the voice signal includes: Obtaining a speech word slot in the text information according to the text information corresponding to the speech signal; Searching for multiple text reply sentences corresponding to the multiple voice word slots in the user database according to the multiple voice word slots in the text information; Based on the intention description of the user's voice, and in accordance with set rules, the multiple text reply sentences corresponding to the multiple voice word slots are associated to form multiple text reply sentences adapted to the text information corresponding to the voice signal; Sorting the plurality of text reply statements, and adding the sorted plurality of text reply statements into a text reply statement list in order; Before obtaining the user voice signal, the method further includes: The image acquisition module of the self-service device obtains facial information of one or more users within a set distance; Acquire mouth image change information in the one or more user facial information based on the one or more user facial information; Acquire one or more user voice signals based on the mouth image change information in the one or more user facial information, wherein the user voice signals include user voices corresponding to the mouth image changes in the one or more user facial information; Obtaining matching information between the voice signal and the user's facial information based on the one or more voice signals and mouth image change information in one or more user's facial information; According to the matching information between the voice signal and the user face information, one of the one or more voice signals is obtained according to the set priority as the obtained user voice signal.

2. The method according to claim 1, characterized in that The acquiring, based on the one or more user facial information, mouth image change information in the one or more user facial information includes: Based on the image information collected by the image collection device of the self-service device within a set distance, one or more facial images in the collected image information are obtained according to the face recognition model; performing image tracking on one or more of the facial images, respectively, to obtain a facial center point in each of the facial images, the facial center point being a reference point for movement of the facial image in three-dimensional space, and determining motion information of other parts of the facial image by determining a motion trajectory of the facial center point in three-dimensional space; According to the movement information of the face center point in the face image in the three-dimensional space, the mouth image change information in the face image is obtained.

3. The method according to claim 1, characterized in that Before acquiring the user voice signal, the method further includes: Build a speech-to-text conversion model; Based on the speech-to-text conversion model, a speech-to-text conversion database and a semantic text database corresponding to fuzzy speech are constructed and updated in real time; Obtaining the speech word slot after speech-to-text conversion based on the speech-to-text conversion database, the semantic text database corresponding to the fuzzy speech, and the speech word slot extraction rules after speech-to-text conversion; According to the acquired speech word slot, one or more text reply sentences corresponding to the speech word slot are set; A user database is constructed based on one or more text response sentences corresponding to all the voice word slots.

4. The method according to claim 1, wherein The acquiring of the user voice signal, processing and recognizing the user voice signal, and obtaining text information corresponding to the voice signal includes: Acquire a user voice signal input by the user through a speech recognition application program interface, and convert the user voice signal input by the user into user input text information; According to the user input text information, the user input text information is converted into speech word slot information, and the speech word slot information uses a word vector representation method to perform speech word slot extraction on the user input text information to obtain the semantics of the user input text information; According to the voice word slot information converted from the text information input by the user, text information corresponding to the voice signal recognized by the self-service device is obtained.

5. The method according to claim 1, wherein After converting the text reply sentences in the text-to-speech conversion list into voice reply sentences one by one and outputting them, obtaining the user's response information to the voice reply sentences includes: Obtaining the first text reply sentence in the list sequence of the text-to-speech conversion list; After converting the first text response sentence into a first voice response sentence, outputting the first voice response sentence, deleting the first text response sentence from the text-to-speech conversion list, using the next text response sentence in the list sequence of the text-to-speech conversion list as the first voice response sentence, and updating the text-to-speech conversion list; Obtain the user's response information to the first voice reply statement; If the user's response information to the first voice reply statement is yes, the control executes the control instruction pointed to by the first voice reply statement and deletes the updated text-to-speech conversion list; If the user's response information to the first voice response statement is no, the first voice response statement converted from the first text response statement in the updated text-to-speech conversion list is output, and the first text response statement is deleted from the updated text-to-speech conversion list, and the next text response statement in the list sequence of the updated text-to-speech conversion list is used as the first voice response statement, and the text-to-speech conversion list is updated again until the user's response information to the first voice response statement is yes, and the control instruction corresponding to the first voice response statement is obtained.

6. A human-computer interaction device for a self-service device, characterized in that: The device comprises: An acquisition module is configured to acquire user voice signals, process and identify the user voice signals, and obtain text information corresponding to the voice signals, wherein the text information includes a description of the user voice's intention and a voice word slot; acquire one or more user face information within a set distance through an image acquisition module of the self-service device; acquire mouth image change information in the one or more user face information based on the one or more user face information; acquire one or more user voice signals based on the mouth image change information in the one or more user face information, wherein the user voice signal includes the user voice with the mouth image change in the one or more user face information; acquire matching information between the voice signal and the user face information based on the one or more voice signals and the mouth image change information in the one or more user face information; acquire one of the one or more voice signals as the acquired user voice signal according to a set priority based on the matching information between the voice signal and the user face information; A processing module is used to search for a list of text reply sentences that are adapted to the text information corresponding to the voice signal in a user database based on the text information corresponding to the voice signal; output the text reply sentences in the text reply sentence list to a text-to-speech conversion list in order; convert the text reply sentences in the text-to-speech conversion list into voice reply sentences one by one and output them, and then obtain the user's response information to the voice reply sentences; obtain the voice word slots in the text information based on the text information corresponding to the voice signal; search for multiple text reply sentences corresponding to multiple voice word slots in the user database based on multiple voice word slots; according to the intention description of the user's voice, and in accordance with set rules, associate the multiple text reply sentences corresponding to the multiple voice word slots to form multiple text reply sentences that are adapted to the text information corresponding to the voice signal; sort the multiple text reply sentences, and add the sorted multiple text reply sentences to the text reply sentence list in order; The judgment module is configured such that if the user's response to the voice reply statement is yes, the self-service device executes the control instruction corresponding to the voice reply statement; if the user's response to the voice reply statement is no, the next voice reply statement is output until the user's response to the voice reply statement is yes; if the user's response to the last voice reply statement in the text-to-speech conversion list is no, a prompt is given to reacquire the user's voice signal.

7. An electronic device, characterized in that: The device includes one or more processors and a memory, wherein the memory is used to store one or more programs; When the one or more programs are executed by the processor, the processor is caused to implement the method according to any one of claims 1 to 5.

8. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed, the method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Method and device for outputting information

    CN110263142A

  • Method for correcting a speech response and natural language dialogue system

    US20140188477A1