Live voice interaction method, device, storage medium and electronic device

By introducing voice interaction mode in live e-commerce applications, the problem of poor vision or blind users cannot interact friendly is solved, and a full voice-based shopping experience is achieved, suitable for users with limited vision and expanded application scenarios.

CN115050361BActive Publication Date: 2025-06-06BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210592020.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-27
Publication Date
2025-06-06
Estimated Expiration
2042-05-27

AI Technical Summary

Technical Problem

The existing live e-commerce technology cannot effectively provide friendly interactive services to people with poor eyesight or blindness, which limits its application scenarios.

Method used

By implementing the voice interaction mode in the live broadcast application, users can interact through the voice input device. The live broadcast application converts the voice input into liquid information interaction content to realize a full voice-based shopping experience.

Benefits of technology

No need for user manual input, shopping can be completed in the live broadcast room by voice interaction alone, suitable for users with poor vision or blindness, significantly expanding the application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115050361B_ABST
    Figure CN115050361B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a live broadcast voice interaction method, device, storage medium and electronic device. The above method includes, when the live broadcast application is opened, in response to receiving voice interaction mode wake-up information, configuring the working state of the live broadcast application to a voice interaction state, the above voice interaction state is a state of determining an object by parsing the information received by the voice input device and performing liquidity information interaction for the above object; when the above working state is the above voice interaction state, in response to receiving liquidity information, generating interaction management information for the target object based on the above liquidity information; when the above working state is the above voice interaction state, in response to receiving interaction trigger information, generating liquidity information interaction content for the above target object based on the above interaction trigger information and the above interaction management information. The present disclosure allows shopping to be completed in the live broadcast room by voice throughout the process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of Internet technology, and in particular to a live voice interaction method, device, storage medium and electronic device. Background Art

[0002] With the development of Internet technology, live broadcast applications have played an increasingly important role in people's lives. It is a general trend to develop e-commerce in live broadcast applications. However, live broadcast e-commerce in related technologies mostly relies on character input devices such as keyboards and mice to interact with users, which is not friendly enough for people with poor eyesight. In other words, related technologies cannot provide interactive services for some specific groups of people, and the application scenarios of interactive services are obviously limited. Summary of the invention

[0003] In order to solve at least one of the above technical problems, the present disclosure provides a live voice interaction method, device, storage medium and electronic device. The technical solution of the present disclosure is as follows:

[0004] According to a first aspect of an embodiment of the present disclosure, a live voice interaction method is provided, comprising:

[0005] When the live broadcast application is opened, in response to receiving voice interaction mode wake-up information, the working state of the live broadcast application is configured as a voice interaction state, wherein the voice interaction state is a state of determining an object by parsing information received by a voice input device and performing liquidity information interaction for the object;

[0006] In a case where the working state is the voice interaction state, in response to receiving the circulation information, generating the interaction management information for the target object based on the circulation information;

[0007] When the working state is the voice interaction state, in response to receiving interaction trigger information, the circulation information interaction content of the target object is generated based on the interaction trigger information and the interaction management information.

[0008] In an exemplary embodiment, the method further comprises:

[0009] When the working state is the voice interaction state, the voice input device is maintained in an awake state.

[0010] In an exemplary embodiment, the method further comprises:

[0011] When the working state is the voice interaction state, in response to receiving voice interaction mode closing information, the working state is configured as a character interaction state, and the voice input device is configured as a non-awakening state. The character interaction state is a state in which an object is determined through information received by the character input device and liquidity information interaction is performed for the object.

[0012] In an exemplary embodiment, when the live broadcast application is opened, in response to receiving voice interaction mode wake-up information, before configuring the working state of the live broadcast application to the voice interaction state, the method further includes:

[0013] In response to the live broadcast application being opened, triggering the voice output device to broadcast first information and wake up the voice input device, wherein the first information is used to inquire the user whether to configure the live broadcast application to the voice interaction state;

[0014] parsing feedback information received by the voice input device for the first information to obtain a first parsing result;

[0015] When the first analysis result indicates that the user has a willingness for voice interaction, it is determined that the voice interaction mode wake-up information is received.

[0016] In an exemplary embodiment, when the working state is the voice interaction state, in response to receiving the liquidity information, before generating the interaction management information for the target object based on the liquidity information, the method further includes:

[0017] In response to the live broadcast process reaching the liquidity information interaction period corresponding to the target object, triggering the voice output device to broadcast second information, where the second information is used to inquire whether the user is associated with the target object;

[0018] parsing feedback information received by the voice input device for the second information to obtain a second parsing result;

[0019] When the second analysis result indicates that the user has an association intention, the liquidity information is received.

[0020] In an exemplary embodiment, when the working state is the voice interaction state, in response to receiving the liquidity information, before generating the interaction management information for the target object based on the liquidity information, the method further includes:

[0021] In response to receiving a first target voice during the live broadcast, parsing the first target voice;

[0022] In response to the existence of information corresponding to the target object in the parsing result, the circulation information is received.

[0023] In an exemplary embodiment, the liquidity information includes at least one liquidity information item, and the receiving the liquidity information includes:

[0024] Obtain at least one attribute item corresponding to the target object;

[0025] For each attribute item, trigger the voice output device to broadcast third information corresponding to the attribute item, where the third information is used to request the user to select at least one attribute value corresponding to the attribute item;

[0026] The feedback information received by the voice input device for the third information is analyzed to obtain the circulation information item corresponding to the attribute item.

[0027] In an exemplary embodiment, the liquidity information includes at least one liquidity information item, and the receiving the liquidity information includes:

[0028] receiving a second target speech, and parsing the second target speech;

[0029] Obtain at least one attribute item corresponding to the target object;

[0030] The content corresponding to each of the attribute items is determined according to the analysis result, and the circulation information item corresponding to the attribute item is obtained according to the content.

[0031] In an exemplary embodiment, when the working state is the voice interaction state, in response to receiving the interaction trigger information, before generating the circulation information interaction content of the target object based on the interaction trigger information and the interaction management information, the method further includes:

[0032] After generating the interaction management information, triggering the voice output device to broadcast fourth information, the fourth information is used to request the user's interaction trigger parameter, the interaction trigger parameter includes at least one of the following: interaction trigger mode, interaction trigger account, interaction trigger password, interaction trigger password, interaction trigger biometric information;

[0033] The interaction trigger information is obtained according to the received interaction trigger parameter.

[0034] In an exemplary embodiment, the generating the target object's circulation information interaction content based on the interaction trigger information and the interaction management information includes:

[0035] triggering the voice output device to broadcast fifth information, where the fifth information is used to request the user's address information;

[0036] Parsing the received third target voice to obtain the address information;

[0037] The circulation information interaction content is generated according to the interaction management information, the interaction trigger information and the address information.

[0038] In an exemplary implementation, parsing the received third target voice to obtain the address information includes:

[0039] Get address template information;

[0040] The parsing result of the third target voice is matched with the address template information to obtain the address information.

[0041] In an exemplary embodiment, in response to receiving the circulation information, before generating the interaction management information for the target object based on the circulation information, the method further includes: triggering the voice output device to broadcast each circulation information item in the circulation information;

[0042] The response to receiving the liquidity information, generating the interactive management information for the target object based on the liquidity information, includes: generating the interactive management information for the target object based on the liquidity information when a confirmation instruction for each of the liquidity information items in the liquidity information is received.

[0043] In an exemplary embodiment, in response to receiving the circulation information, before generating the interaction management information for the target object based on the circulation information, the method further includes:

[0044] In response to receiving the fourth target speech, parsing the fourth target speech;

[0045] The target content related to the target object is determined according to the analysis result of the fourth target voice, and the voice output device is triggered to broadcast the associated information corresponding to the target content, where the associated information is used to explain or give suggestions for the target content.

[0046] According to a second aspect of an embodiment of the present disclosure, a live voice interaction device is provided, comprising:

[0047] The switching module is configured to execute, when the live broadcast application is opened, in response to receiving voice interaction mode wake-up information, configuring the working state of the live broadcast application to a voice interaction state, wherein the voice interaction state is a state of determining an object by analyzing information received by a voice input device and performing liquidity information interaction for the object;

[0048] The voice processing module is configured to execute, when the working state is the voice interaction state, in response to receiving the liquidity information, generating interaction management information for the target object based on the liquidity information; and, when the working state is the voice interaction state, in response to receiving the interaction trigger information, generating the liquidity information interaction content of the target object based on the interaction trigger information and the interaction management information.

[0049] In an exemplary embodiment, the voice processing module is configured to maintain the wake-up state of the voice input device when the working state is the voice interaction state.

[0050] In an exemplary embodiment, the switching module is configured to execute, when the working state is the voice interaction state, in response to receiving information that the voice interaction mode is turned off, configuring the working state to a character interaction state and configuring the voice input device to a non-awakening state, wherein the character interaction state is a state in which an object is determined through information received from the character input device and liquidity information interaction is performed on the object.

[0051] In an exemplary embodiment, the switching module is configured to execute:

[0052] In response to the live broadcast application being opened, triggering the voice output device to broadcast first information and wake up the voice input device, wherein the first information is used to inquire the user whether to configure the live broadcast application to the voice interaction state;

[0053] parsing feedback information received by the voice input device for the first information to obtain a first parsing result;

[0054] When the first analysis result indicates that the user has a willingness for voice interaction, it is determined that the voice interaction mode wake-up information is received.

[0055] In an exemplary embodiment, the speech processing module is configured to perform:

[0056] In response to the live broadcast process reaching the liquidity information interaction period corresponding to the target object, triggering the voice output device to broadcast second information, where the second information is used to inquire whether the user is associated with the target object;

[0057] parsing feedback information received by the voice input device for the second information to obtain a second parsing result;

[0058] When the second analysis result indicates that the user has an association intention, the liquidity information is received.

[0059] In an exemplary embodiment, the speech processing module is configured to perform:

[0060] In response to receiving a first target voice during the live broadcast, parsing the first target voice;

[0061] In response to the existence of information corresponding to the target object in the parsing result, the circulation information is received.

[0062] In an exemplary embodiment, the speech processing module is configured to perform:

[0063] Obtain at least one attribute item corresponding to the target object;

[0064] For each attribute item, trigger the voice output device to broadcast third information corresponding to the attribute item, where the third information is used to request the user to select at least one attribute value corresponding to the attribute item;

[0065] The feedback information received by the voice input device for the third information is analyzed to obtain the circulation information item corresponding to the attribute item.

[0066] In an exemplary embodiment, the speech processing module is configured to perform:

[0067] receiving a second target speech, and parsing the second target speech;

[0068] Obtain at least one attribute item corresponding to the target object;

[0069] The content corresponding to each of the attribute items is determined according to the analysis result, and the circulation information item corresponding to the attribute item is obtained according to the content.

[0070] In an exemplary embodiment, the speech processing module is configured to perform:

[0071] After generating the interaction management information, triggering the voice output device to broadcast fourth information, the fourth information is used to request the user's interaction trigger parameter, the interaction trigger parameter includes at least one of the following: interaction trigger mode, interaction trigger account, interaction trigger password, interaction trigger password, interaction trigger biometric information;

[0072] The interaction trigger information is obtained according to the received interaction trigger parameter.

[0073] In an exemplary embodiment, the speech processing module is configured to perform:

[0074] triggering the voice output device to broadcast fifth information, where the fifth information is used to request the user's address information;

[0075] Parsing the received third target voice to obtain the address information;

[0076] The circulation information interaction content is generated according to the interaction management information, the interaction trigger information and the address information.

[0077] In an exemplary embodiment, the speech processing module is configured to perform:

[0078] Get address template information;

[0079] The parsing result of the third target voice is matched with the address template information to obtain the address information.

[0080] In an exemplary embodiment, the speech processing module is configured to perform:

[0081] Triggering the voice output device to broadcast each circulation information item in the circulation information;

[0082] When a confirmation instruction for each of the circulation information items in the circulation information is received, interaction management information for a target object is generated based on the circulation information.

[0083] In an exemplary embodiment, the speech processing module is configured to perform:

[0084] In response to receiving the fourth target speech, parsing the fourth target speech;

[0085] The target content related to the target object is determined according to the analysis result of the fourth target voice, and the voice output device is triggered to broadcast the associated information corresponding to the target content, where the associated information is used to explain or give suggestions for the target content.

[0086] According to a third aspect of an embodiment of the present disclosure, there is provided an electronic device, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to execute instructions to implement a method as described in any one of the first aspects above.

[0087] According to a fourth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided. When instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute any one of the methods in the first aspect of the embodiment of the present disclosure.

[0088] According to the fifth aspect of an embodiment of the present disclosure, a computer program product is provided, the computer program product comprising a computer program, the computer program being stored in a readable storage medium, at least one processor of a computer device reading and executing the computer program from the readable storage medium, so that the computer device performs any one of the methods in the first aspect of the embodiment of the present disclosure.

[0089] The technical solution provided by the embodiments of the present disclosure brings at least the following beneficial effects:

[0090] The disclosed embodiment completes the acquisition of liquidity information, generation of interactive management information, interaction triggering, and generation of interactive content of liquidity information through full voice, without the need to use the eyes, which significantly expands the application scenarios. Taking the e-commerce scenario as an example, shopping can be completed in the live broadcast room by voice interaction alone without the need for the user to manually input any information. This technical solution does not require the user to use the eyes, but can complete shopping in the live broadcast room by hearing and issuing voice commands, which is very suitable for users who need to rest their eyes or have poor eyesight. At present, the scale of the blind or visually impaired population is showing an increasing trend. There are also some elderly people with poor eyesight. Because of the inconvenience of life, these users may have a stronger demand for online shopping. The disclosed embodiment supports voice shopping during live shopping, which can well meet the specific needs of such users.

[0091] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0092] The drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute improper limitations on the present disclosure.

[0093] Figure 1 It is a schematic diagram of an implementation environment of a live voice interaction method according to an exemplary embodiment.

[0094] Figure 2 is a flow chart of a live voice interaction method according to an exemplary embodiment;

[0095] Figure 3 is a schematic diagram of a scenario of a live voice interaction method according to an exemplary embodiment;

[0096] Figure 4 According to an exemplary embodiment, Figure 2 A schematic diagram of a scenario for requesting confirmation of the circulation information of the interactive management information obtained by the example;

[0097] Figure 5Another schematic diagram of a live voice interaction method according to an exemplary embodiment;

[0098] Figure 6 A schematic diagram of an address information adaptation method according to an exemplary embodiment;

[0099] Figure 7 A specific voice interaction scenario flow chart according to an exemplary embodiment;

[0100] Figure 8 A block diagram of a live voice interaction device according to an exemplary embodiment;

[0101] Fig. 9 It is a block diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION

[0102] In order to enable ordinary persons in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings.

[0103] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar first objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0104] All information involved in this disclosure is authorized by the user or fully authorized by all parties.

[0105] Figure 1 FIG. 1 is a schematic diagram of an implementation environment of a live voice interaction method according to an exemplary embodiment. Figure 1 , the implementation environment specifically includes: a terminal 101 and a server 102.

[0106] Terminal 101 may be at least one of a smart phone, a smart watch, a desktop computer, a laptop computer and a portable computer. An application providing a live broadcast service may be installed and run on terminal 101, and a user may log in to the application through terminal 101 to obtain the live broadcast service provided by the application, and participate in e-commerce shopping while using the live broadcast service. Terminal 101 may generally refer to one of a plurality of terminals, and this embodiment only takes terminal 101 as an example. Those skilled in the art may know that the number of the above-mentioned terminals may be more or less. For example, the above-mentioned terminals may be only a few, or the above-mentioned terminals may be dozens or hundreds, or a greater number. The embodiment of the present disclosure does not limit the number of terminals and device types.

[0107] The server 102 may be at least one of a single server, multiple servers, a cloud computing platform, and a virtualization center. The server 102 may be connected to the terminal 101 and other terminals via a wireless network or a wired network. The server 102 may receive voice information sent by the terminal 101, and provide live broadcast and e-commerce shopping services to the terminal 101 based on the received voice information. Of course, the server 102 may also include other functional servers to provide more comprehensive and diversified services. Figure 2 is a flow chart of a live voice interaction method according to an exemplary embodiment. Figure 2 As shown, the above method at least includes the following steps S10-S30.

[0108] In step S10, when the live broadcast application is opened, in response to receiving the voice interaction mode wake-up information, the working state of the live broadcast application is configured as a voice interaction state, and the voice interaction state is a state of determining the object by analyzing the information received by the voice input device and performing liquidity information interaction on the object.

[0109] The embodiments of the present disclosure do not limit the objects and liquidity information interactions, which may have corresponding meanings according to different scenarios. For example, in an e-commerce scenario, the object may refer to an item in an e-commerce transaction, and the liquidity information interaction may refer to the transaction of the item.

[0110] The present disclosure does not limit the live broadcast application. The voice interaction mode wake-up information can be received when the live broadcast application is opened, so that the user can conduct e-commerce shopping based on voice when visiting any live broadcast room. The voice interaction mode wake-up information can also be received when the live broadcast application is opened and the user enters a specific live broadcast room. After receiving the voice interaction mode wake-up information, the voice interaction state can be entered. In this state, the voice input device is the only device for receiving information input by the user. The live broadcast application determines the object and performs liquidity information interaction for the above object by parsing the information received by the voice input device. In the embodiment of the present disclosure, by setting the voice interaction state, the user can be provided with a full-process live broadcast e-commerce service based on voice interaction in this state. The embodiment of the present disclosure does not limit the specific type of voice input device, which can be various microphones or other devices that can record voice.

[0111] In the embodiment of the present disclosure, when the working state is the voice interaction state, the voice input device is maintained in the awake state. That is, when the working state of the live broadcast application is configured as the voice interaction state, the voice input device is in the awake state throughout the process, and the user can participate in the e-commerce activities in the live broadcast room by voice, for example, by placing orders, selecting objects, triggering liquidity information interaction, etc.

[0112] In step S20, when the working state is the voice interaction state, in response to receiving the circulation information, the interaction management information for the target object is generated based on the circulation information.

[0113] The embodiments of the present disclosure do not limit the liquidity information. In the e-commerce scenario, it can refer to the order information of the item. The embodiments of the present disclosure do not limit its specific content. It can include various information of the target object that is about to exchange liquidity information, such as the type, size, color, price, quantity, mode of transportation, etc. of the target object. The target object can be understood as any object listed in the live broadcast room. The user's voice can be obtained through the voice input device, and the liquidity information can be obtained according to the received voice.

[0114] The interactive management information is not limited in the embodiments of the present disclosure, and may refer to an item order in an e-commerce scenario. Specifically, the interactive management information refers to an information set formed by the information obtained related to the target object. Based on the interactive management information, combined with the user's address information and liquidity information, liquidity information interaction can be performed to obtain liquidity information interaction content. The liquidity information interaction content is the credential for completing the liquidity information interaction of the target object. Based on the liquidity information interaction content, operations required for other e-commerce links such as receipt operations, refund operations, rights protection operations, logistics inquiries, etc. can also be implemented.

[0115] In step S30, when the working state is the voice interaction state, in response to receiving the interaction trigger information, the liquidity information interaction content of the target object is generated based on the interaction trigger information and the interaction management information.

[0116] The embodiments of the present disclosure do not limit the specific meaning of the interaction trigger information. Taking the e-commerce scenario as an example, it can refer to the payment information generated for purchasing e-commerce items. Specifically, the interaction trigger information can be understood as the information required by the user to trigger the interaction for the target object in the interaction management information. Of course, the embodiments of the present disclosure do not limit the specific meaning of the liquidity information interaction content. In the e-commerce scenario, it can represent the transaction order generated for the target object.

[0117] After the interaction trigger information is obtained, the interaction trigger operation can be performed on the above-mentioned interaction management information, and the liquidity information interaction content is generated after the interaction trigger is successful. The liquidity information interaction content is the certificate obtained by the liquidity information interaction. So far, the acquisition of liquidity information, generation of interaction management information, interaction triggering, and generation of liquidity information interaction content are completed by voice throughout the process. Taking the e-commerce scenario as an example, shopping can be completed in the live broadcast room by voice interaction alone without the need for users to manually enter any information. This technical solution does not require users to use their eyes, but can complete shopping in the live broadcast room by hearing and issuing voice commands alone. It is very suitable for users who need to rest their eyes or have poor eyesight. At present, the scale of the blind or visually impaired population is showing an increasing trend. There are also some elderly people with poor eyesight. Because of the inconvenience of life, these users may have a stronger demand for online shopping. The disclosed embodiment supports voice shopping during live shopping, which can well meet the specific needs of such users.

[0118] In an exemplary embodiment, when the above-mentioned working state is the above-mentioned voice interaction state, in response to receiving the voice interaction mode closing information, the above-mentioned working state is configured as the character interaction state, and the above-mentioned voice input device is configured as the non-awakening state, and the above-mentioned character interaction state is a state of determining the object through the information received by the character input device and conducting liquidity information interaction for the above-mentioned object.

[0119] That is to say, when the working state is the voice interaction state, the user can switch the working state to the character interaction state by sending a voice interaction mode closing message. The character interaction state can be understood as a shopping state that ordinary users often use. In this state, information can be input through character input devices such as keyboards, mice, and touch screens, thereby completing the selection of objects and the exchange of circulation information of objects. In other words, the embodiment of the present disclosure provides an e-commerce shopping solution in voice mode, and also provides an ordinary e-commerce shopping solution. By being compatible with the two shopping solutions and allowing users to switch at will, the different shopping needs of different types of users are fully met.

[0120] In an exemplary embodiment, when the live broadcast application is opened, in response to receiving the voice interaction mode wake-up information, before configuring the working state of the live broadcast application to the voice interaction state, the method further includes: in response to the live broadcast application being opened, triggering the voice output device to broadcast the first information and wake up the voice input device, wherein the first information is used to inquire the user whether to configure the live broadcast application to the voice interaction state; parsing the feedback information received by the voice input device for the first information to obtain a first parsing result; and when the first parsing result indicates that the user has the willingness to interact with voice, determining that the voice interaction mode wake-up information has been received.

[0121] The present disclosure does not limit the voice output device, for example, it can be a speaker, sound box and other peripherals. When the live broadcast application is opened, the voice output device can be triggered to play a voice (first information) so that the user can choose whether to set the live broadcast application to a voice interaction state. For example, the first information can be "Do you want to enter the voice live broadcast room" or "Please select the interaction mode. If you need to enter the voice live broadcast state, please issue the 'voice live broadcast' command, otherwise, please issue the 'normal live broadcast' command." After hearing the first information, the user can give feedback in the form of voice. For example, the feedback voice can be "yes" or "voice live broadcast."

[0122] In the case where the first information is "Do you want to enter the voice live broadcast room?" and the voice feedback from the user is "Yes", or in the case where the first information is "Please select the interactive mode. If you need to enter the voice live broadcast state, please issue the 'voice live broadcast' command, otherwise, please issue the 'normal live broadcast' command" and the voice feedback from the user is "voice live broadcast", it can be determined that the user has the intention to interact with the voice. Therefore, it can be considered that the above-mentioned voice interaction mode wake-up information has been received, and the voice interaction state can be automatically entered, supporting the provision of e-commerce shopping experience to users in a voice-based manner throughout the process. Under this implementation method, users can be provided with an entry to trigger the live broadcast application to enter the voice interaction state in a timely manner, quickly enter the voice interaction state, and provide a better experience for users with poor eyesight.

[0123] In an exemplary embodiment, when the working state is the voice interaction state, in response to receiving the liquidity information, before generating the interactive management information for the target object based on the liquidity information, the method further includes: in response to the live broadcast process reaching the liquidity information interaction period corresponding to the target object, triggering the voice output device to broadcast the second information, the second information is used to ask the user whether to associate the target object; parsing the feedback information received by the voice input device for the second information to obtain the second parsing result; receiving the liquidity information when the second parsing result indicates that the user has the intention to associate. Specifically, in the e-commerce scenario, the user is associated with the target object, which means that the user purchases the target object.

[0124] As the live broadcast progresses, the anchor may put many objects on the shelves in turn. When each object is put on the shelves, a period of time can be reserved for the user to place an order for the object. The live broadcast application can trigger the above-mentioned voice output device to broadcast a piece of content (second information) during this time period to ask the user whether to place an order to associate with the object. For example, the anchor puts an electric toothbrush on the shelves and prompts the audience to place an order. In this case, the second information "The current anchor has put an electric toothbrush on the shelves, please ask if it is associated" can be broadcast. If the voice information fed back by the user is "willing", "I buy" or "I want" and other content, it can be determined that the user has the intention to associate. In this case, voice interaction with the user can continue to be carried out to obtain subsequent liquidity information. This solution can accompany the live broadcast process, actively ask the user whether there is a need for association for each object on the shelves, and provide users with the opportunity to place an order in a timely manner, ensuring that users will not miss the opportunity to place an order due to reasons such as inability to see, unclear vision or inconvenience in using the mouse and keyboard, thereby improving the user's shopping experience.

[0125] In another exemplary embodiment, when the working state is the voice interaction state, in response to receiving the liquidity information, before generating the interaction management information for the target object based on the liquidity information, the method further includes: in response to receiving the first target voice during the live broadcast, parsing the first target voice; in response to the existence of information corresponding to the target object in the analysis result, receiving the liquidity information.

[0126] Of course, at any point in the live broadcast process, the user can trigger the order process for the objects put on the shelves during the live broadcast. That is to say, the user can output the voice for placing an order (the first target voice) at any time. For example, the current live broadcast room will have a total of four objects on the shelves: "windbreaker", "lipstick", "fan" and "drinks". At any time in the live broadcast process, the user can place an order for the above four objects. For example, you can input "I want to buy a windbreaker", "I want a fan", etc. to trigger the live broadcast application to interact with the user by voice, and then obtain specific liquidity information.

[0127] The disclosed embodiments do not limit the specific content of the first target voice. For example, if it is determined that the received voice contains objects listed in the live broadcast room and words expressing the user's willingness to associate, it can be determined that the first target voice is received. Or if the received voice contains preset instruction words, such as "place an order", the voice output device can be triggered to play the corresponding response "Please place an order". After the response is played, the voice received again through the voice input device can be considered as the first target voice. This solution supports users to place orders and shop at any time during the live broadcast, further enhancing the shopping experience.

[0128] In an exemplary embodiment, the above-mentioned liquidity information includes at least one liquidity information item, and the above-mentioned receiving the above-mentioned liquidity information includes: obtaining at least one attribute item corresponding to the above-mentioned target object; for each attribute item, triggering the above-mentioned voice output device to broadcast the third information corresponding to the above-mentioned attribute item, and the above-mentioned third information is used to request the user to select at least one attribute value corresponding to the above-mentioned attribute item; parsing the feedback information received by the above-mentioned voice input device for the above-mentioned third information, and obtaining the liquidity information item corresponding to the above-mentioned attribute item.

[0129] For example, if a user associates a windbreaker, the windbreaker object includes at least four attributes, namely size, color, quantity and pattern. Accordingly, the circulation information items that need to be filled in include at least windbreaker size, windbreaker color, windbreaker quantity and windbreaker pattern. The user can be guided to select the specific attribute values ​​of the above four attributes in an item-by-item interactive manner, and the corresponding circulation information items are obtained according to the selection results, thereby obtaining the overall circulation information.

[0130] like Figure 3As shown, the voice output device can announce "Please select the size of the windbreaker, S, M, L" and wait for user input. After the voice input device receives the voice "M", the specific attribute value of "size" is determined to be "M". Then, the voice output device can announce "Please select the color of the windbreaker, black, gray, white" and wait for user input. After the voice input device receives the voice "black", the specific attribute value of "color" is determined to be "black". Then, the voice output device can announce "Please determine the quantity of windbreaker" and wait for user input. After the voice input device receives the voice "2", the specific attribute value of "quantity" is determined to be "2". Then, the voice output device can announce "Please select the style of the windbreaker, loose or slim" and wait for user input. After the voice input device receives the voice "slim", the specific attribute value of "style" is determined to be "slim".

[0131] Please refer to Figure 4 , which shows that according to Figure 3 Schematic diagram of the exchange management information obtained by the example. Figure 3 In the example, the liquidity information corresponding to the target object can be obtained:

[0132] Target audience: Windbreaker;

[0133] Windbreaker size: M;

[0134] Windbreaker color: black;

[0135] Number of windbreakers: 2;

[0136] Trench coat fit: Slim fit.

[0137] The liquidity information can be broadcasted to request confirmation, and after the user confirms, interactive management information for the target object, the windbreaker, can be generated based on the liquidity information. In this implementation, the user's determination result for the attribute value of each attribute item is determined by guiding the user item by item, without missing the content of important attribute items, so as to obtain interactive management information that fully meets the user's wishes, ensure that there are no errors in the ordering process, and improve the user's shopping experience.

[0138] In another embodiment, the above-mentioned circulation information includes at least one circulation information item, and the above-mentioned receiving the above-mentioned circulation information includes: receiving a second target voice, and parsing the above-mentioned second target voice; obtaining at least one attribute item corresponding to the above-mentioned target object; determining the content corresponding to each of the above-mentioned attribute items according to the parsing result, and obtaining the circulation information item corresponding to the above-mentioned attribute item according to the above-mentioned content.

[0139] In this embodiment, the user can utter a voice (second target voice) after entering the ordering state. The ordering state can be entered by the first target voice uttered in the previous text. The second target voice may include content related to one or more attribute items of the target object. By parsing the second target voice, the content corresponding to the relevant attribute items can be obtained. Please refer to Figure 5 After entering the ordering state, the user outputs a voice (second target voice), and the second target voice may be "I want to buy a black slim windbreaker", and the following three contents in the liquidity information can be determined:

[0140] Target audience: Windbreaker;

[0141] Windbreaker color: black;

[0142] Trench coat fit: Slim fit.

[0143] As can be seen from the above, the liquidity information is not enough. In this case, a voice prompt can be given to the user to guide the user to issue a new second target voice again. For example, the voice output device is triggered to play "Please continue to say the number and size of windbreakers". The user continues to issue a new second target voice "I want to buy two size M windbreakers". In this case, the following two contents of the order can be determined:

[0144] Windbreaker size: M;

[0145] Number of windbreakers:2.

[0146] At this point, the liquidity information is complete. This implementation method can quickly determine the entire content of the liquidity information through one or a limited number of voice interactions, with high interaction efficiency, fast order placement, and no information loss or error.

[0147] At least two methods of obtaining liquidity information are given in the foregoing text. In the process of obtaining the liquidity information items in the liquidity information, the voice output device can also request the user to confirm each determined liquidity information item. For example, after obtaining the complete liquidity information, the voice output device can be triggered to broadcast the specific content of the liquidity information "Please confirm the following: Target object: Windbreaker; Windbreaker size: M; Windbreaker color: black; Windbreaker quantity: 2; Windbreaker version: slim fit." If the user feedback is "all correct", interactive management information can be generated. If the user feedback indicates that there are erroneous liquidity information items, voice interaction with the user can continue until all the liquidity information is correct. That is to say, in the embodiment of the present disclosure, when a confirmation instruction is received for each of the above-mentioned liquidity information items in the above-mentioned liquidity information, interactive management information for the target object is generated based on the above-mentioned liquidity information. By confirming the various contents of the liquidity information with the user, it is ensured that the generated interactive management information meets the user's expectations, the probability of placing an erroneous order is reduced, and the user experience is improved.

[0148] In one embodiment, when the above-mentioned working state is the above-mentioned voice interaction state, in response to receiving the interaction trigger information, before generating the liquidity information interaction content of the above-mentioned target object based on the above-mentioned interaction trigger information and the above-mentioned interaction management information, the above-mentioned method also includes: after generating the above-mentioned interaction management information, triggering the above-mentioned voice output device to broadcast fourth information, and the above-mentioned fourth information is used to request the user's interaction trigger parameters, and the above-mentioned interaction trigger parameters include at least one of the following: interaction trigger mode, interaction trigger account, interaction trigger password, interaction trigger password, interaction trigger biometric information; the above-mentioned interaction trigger information is obtained according to the above-mentioned interaction trigger parameters received.

[0149] The disclosed embodiments do not limit the interaction triggering method. For example, it can support interaction triggering of multiple verification methods such as fingerprint, password, iris, face, etc., and does not limit the interaction triggering channel. For example, the interaction triggering channel can be provided by a live broadcast application, or it can be connected to any third-party interaction triggering channel. The information related to the interaction triggering is encapsulated as interaction triggering parameters, and the user is triggered to give the interaction triggering parameters through voice interaction with the user, thereby generating interaction triggering information.

[0150] In one embodiment, after obtaining the interactive management information and the interactive trigger information, it is also necessary to obtain the shipping address of the target object. In this case, the above-mentioned target object's liquidity information interactive content generated based on the above-mentioned interactive trigger information and the above-mentioned interactive management information includes: triggering the above-mentioned voice output device to broadcast the fifth information, and the above-mentioned fifth information is used to request the user's address information; parsing the received third target voice to obtain the above-mentioned address information; generating the above-mentioned liquidity information interactive content according to the above-mentioned interactive management information, the above-mentioned interactive trigger information and the above-mentioned address information. The present disclosure embodiment does not limit the address information, which may include the user's name, telephone number, address, zip code and other content, and may also include optional information for shipping, such as selecting weekday delivery or weekend delivery, morning delivery or afternoon delivery, whether to deliver cash on delivery, whether to associate delivery insurance and other content. The present disclosure embodiment supports fully automatic voice acquisition of target object information, fully automatic voice acquisition of interactive trigger information and fully automatic voice acquisition of address information, thereby realizing the full voice of live shopping and improving user experience.

[0151] The embodiments of the present disclosure do not limit the address information parsing method. Figure 6 As shown, the address template information can be obtained; the parsing result for the third target voice is matched with the address template information to obtain the address information. That is to say, the disclosed embodiment can generate address information according to the address template by intelligent matching. For example, according to the specific address in the parsing result, the corresponding content can be automatically adapted for the province, city, district, house number, floor, building number and other items in the address template, thereby improving the standardization of address information and reducing the probability of delivery errors to the target object. Of course, the address information can also be confirmed by the user through voice broadcast.

[0152] In one embodiment, in response to receiving the liquidity information, before generating the interactive management information for the target object based on the liquidity information, the method further includes: in response to receiving a fourth target voice, parsing the fourth target voice; determining the target content related to the target object according to the parsing result of the fourth target voice, triggering the voice output device to broadcast the associated information corresponding to the target content, wherein the associated information is used to explain or give suggestions for the target content.

[0153] The embodiments of the present disclosure do not limit the specific content and issuance opportunity of the fourth target voice. For example, the user can issue the fourth target voice at any time during the live broadcast process. The fourth target voice can request an explanation or give suggestions on the relevant content of the target object. For example, the live broadcast platform has a total of four objects, "towel", "soap", "electric toothbrush", and "wool coat". The user can issue the following voice "I want to know how long an electric toothbrush can be used after a charge", and the information related to the power consumption of the electric compressor can be automatically broadcast to the user. For another example, the user can also issue "how to wash a woolen coat", and the precautions for washing a woolen coat can be automatically broadcast to the user. The embodiments of the present disclosure can automatically answer the questions that users have during the live broadcast and give relevant suggestions through voice interaction, further enhancing the live broadcast experience.

[0154] Please refer to Figure 7 , which shows a specific voice interaction scenario flow chart in an embodiment of the present disclosure. The voice interaction process includes at least the following steps:

[0155] First, after the user enters the live broadcast room, the user is asked about the interaction mode of the live broadcast room. If the user chooses the voice interaction mode, the corresponding voice live broadcast room is provided to the user. In this voice live broadcast room, the microphone (voice input device) is in an awake state throughout the whole process. If the microphone receives the first target voice issued by the user, the order process can be entered. In the order process, the order information is obtained through voice interaction with the user, and an item order is generated according to the order information. Then, the payment information is obtained through voice interaction with the user, and the payment is completed and a transaction order is generated according to the payment information, thereby successfully completing an e-commerce shopping with full voice.

[0156] Figure 8 is a block diagram of a live voice interaction device according to an exemplary embodiment. Figure 8 , the device comprises:

[0157] The switching module 10 is configured to execute, when the live broadcast application is opened, in response to receiving the voice interaction mode wake-up information, to configure the working state of the live broadcast application to the voice interaction state, wherein the voice interaction state is a state of determining an object by analyzing the information received by the voice input device and performing the circulation information interaction for the object;

[0158] The voice processing module 20 is configured to execute, when the above-mentioned working state is the above-mentioned voice interaction state, in response to receiving the liquidity information, generating interaction management information for the target object based on the above-mentioned liquidity information; and, when the above-mentioned working state is the above-mentioned voice interaction state, in response to receiving the interaction trigger information, generating the liquidity information interaction content for the above-mentioned target object based on the above-mentioned interaction trigger information and the above-mentioned interaction management information.

[0159] In an exemplary embodiment, the voice processing module is configured to maintain the wake-up state of the voice input device when the working state is the voice interaction state.

[0160] In an exemplary embodiment, the switching module is configured to execute, when the working state is the voice interaction state, in response to receiving information that the voice interaction mode is turned off, configuring the working state to a character interaction state and configuring the voice input device to a non-awakening state, wherein the character interaction state is a state in which an object is determined through information received from the character input device and liquidity information interaction is performed on the object.

[0161] In an exemplary embodiment, the switching module is configured to execute:

[0162] In response to the live broadcast application being opened, triggering the voice output device to broadcast first information and wake up the voice input device, wherein the first information is used to inquire the user whether to configure the live broadcast application to the voice interaction state;

[0163] Analyze the feedback information received by the voice input device for the first information to obtain a first analysis result;

[0164] When the first analysis result indicates that the user has a willingness for voice interaction, it is determined that the voice interaction mode wake-up information is received.

[0165] In an exemplary embodiment, the speech processing module is configured to execute:

[0166] In response to the live broadcast process reaching the liquidity information interaction period corresponding to the target object, triggering the voice output device to broadcast second information, where the second information is used to inquire the user whether to associate with the target object;

[0167] Analyze the feedback information received by the voice input device for the second information to obtain a second analysis result;

[0168] When the second analysis result indicates that the user has an intention to associate, the liquidity information is received.

[0169] In an exemplary embodiment, the speech processing module is configured to execute:

[0170] In response to receiving a first target voice during the live broadcast, parsing the first target voice;

[0171] In response to the presence of information corresponding to the target object in the analysis result, the circulation information is received.

[0172] In an exemplary embodiment, the speech processing module is configured to execute:

[0173] Obtain at least one attribute item corresponding to the target object;

[0174] For each attribute item, trigger the voice output device to broadcast third information corresponding to the attribute item, where the third information is used to request the user to select at least one attribute value corresponding to the attribute item;

[0175] The feedback information received by the voice input device for the third information is analyzed to obtain the liquidity information item corresponding to the attribute item.

[0176] In an exemplary embodiment, the speech processing module is configured to execute:

[0177] receiving a second target speech, and parsing the second target speech;

[0178] Obtain at least one attribute item corresponding to the target object;

[0179] The content corresponding to each of the above attribute items is determined according to the analysis result, and the circulation information item corresponding to the above attribute item is obtained according to the above content.

[0180] In an exemplary embodiment, the speech processing module is configured to execute:

[0181] After the interaction management information is generated, the voice output device is triggered to broadcast fourth information, where the fourth information is used to request interaction trigger parameters of the user, where the interaction trigger parameters include at least one of the following: interaction trigger mode, interaction trigger account, interaction trigger password, interaction trigger password, and interaction trigger biometric information;

[0182] The interaction trigger information is obtained according to the received interaction trigger parameter.

[0183] In an exemplary embodiment, the speech processing module is configured to execute:

[0184] triggering the voice output device to broadcast fifth information, where the fifth information is used to request the user's address information;

[0185] Parsing the received third target voice to obtain the above address information;

[0186] The above-mentioned circulation information interaction content is generated according to the above-mentioned interaction management information, the above-mentioned interaction trigger information and the above-mentioned address information.

[0187] In an exemplary embodiment, the speech processing module is configured to execute:

[0188] Get address template information;

[0189] The analysis result of the third target voice is matched with the address template information to obtain the address information.

[0190] In an exemplary embodiment, the speech processing module is configured to execute:

[0191] Triggering the above-mentioned voice output device to broadcast each item of the above-mentioned circulation information;

[0192] When a confirmation instruction for each of the above-mentioned circulation information items in the above-mentioned circulation information is received, interactive management information for the target object is generated based on the above-mentioned circulation information.

[0193] In an exemplary embodiment, the speech processing module is configured to execute:

[0194] In response to receiving the fourth target speech, parsing the fourth target speech;

[0195] The target content related to the target object is determined according to the analysis result of the fourth target voice, and the voice output device is triggered to broadcast the associated information corresponding to the target content, where the associated information is used to explain or give suggestions for the target content.

[0196] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0197] Fig. 9 It is a block diagram of an electronic device 600 for live voice interaction according to an exemplary embodiment.

[0198] The electronic device may be a server or a terminal device, and its internal structure diagram may be as follows: Fig. 9As shown. The electronic device includes a processor, a memory and a network interface connected via a system bus. The processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the electronic device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a live voice interaction method is implemented.

[0199] Those skilled in the art will understand that Fig. 9 The structure shown in the figure is merely a block diagram of a partial structure related to the scheme of the present disclosure, and does not constitute a limitation on the electronic device to which the scheme of the present disclosure is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0200] In an exemplary embodiment, an electronic device is also provided, including: a processor; a memory for storing instructions executable by the processor; wherein the processor is configured to execute the instructions to implement the live voice interaction method as in the embodiment of the present disclosure.

[0201] In an exemplary embodiment, a computer-readable storage medium is also provided. When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device can execute the live voice interaction method in the embodiment of the present disclosure.

[0202] In an exemplary embodiment, a computer program product is also provided, which includes a computer program, the computer program is stored in a readable storage medium, and at least one processor of a computer device reads and executes the computer program from the readable storage medium, so that the computer device executes the live voice interaction method of the embodiment of the present disclosure.

[0203] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0204] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present disclosure are indicated by the following claims.

[0205] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A live voice interaction method, It is characterized in that include: When the live broadcast application is opened, triggering the voice output device to broadcast first information and wake up the voice input device, wherein the first information is used to inquire the user whether to configure the live broadcast application to the voice interaction state; parsing feedback information received by the voice input device for the first information to obtain a first parsing result; In the case where the first analysis result indicates that the user has a willingness to interact with the voice, in response to receiving a voice interaction mode wake-up message, the working state of the live broadcast application is configured as a voice interaction state, wherein the voice interaction state is a state in which an object is determined by analyzing information received by a voice input device and a liquidity information interaction is performed for the object; When the working state is the voice interaction state, in response to the live broadcast process reaching the liquidity information interaction period corresponding to the target object, the voice output device is triggered to broadcast second information, where the second information is used to inquire whether the user is associated with the target object; parsing feedback information received by the voice input device for the second information to obtain a second parsing result; In the case where the second analysis result indicates that the user has an association intention, receiving the circulation information, and generating interaction management information for the target object based on the circulation information; When the working state is the voice interaction state, in response to receiving the interaction trigger information, the liquidity information interaction content of the target object is generated based on the interaction trigger information and the interaction management information to provide a full-process live e-commerce service based on voice interaction.

2. The live voice interaction method according to claim 1, It is characterized in that The method further comprises: When the working state is the voice interaction state, the voice input device is maintained in an awake state.

3. The live voice interaction method according to claim 1, It is characterized in that The method further comprises: When the working state is the voice interaction state, in response to receiving voice interaction mode closing information, the working state is configured as a character interaction state, and the voice input device is configured as a non-awakening state. The character interaction state is a state in which an object is determined through information received by the character input device and liquidity information interaction is performed for the object.

4. The live voice interaction method according to claim 1, It is characterized in that Before generating the interaction management information for the target object based on the circulation information, the method further includes: In response to receiving a first target voice during the live broadcast, parsing the first target voice; In response to the existence of information corresponding to the target object in the parsing result, the circulation information is received.

5. The live voice interaction method according to claim 1 or 4, It is characterized in that The liquidity information includes at least one liquidity information item, and the receiving the liquidity information includes: Obtain at least one attribute item corresponding to the target object; For each attribute item, trigger the voice output device to broadcast third information corresponding to the attribute item, where the third information is used to request the user to select at least one attribute value corresponding to the attribute item; The feedback information received by the voice input device for the third information is analyzed to obtain the circulation information item corresponding to the attribute item.

6. The live voice interaction method according to claim 5, It is characterized in that The liquidity information includes at least one liquidity information item, and the receiving the liquidity information includes: receiving a second target speech, and parsing the second target speech; Obtain at least one attribute item corresponding to the target object; The content corresponding to each of the attribute items is determined according to the analysis result, and the circulation information item corresponding to the attribute item is obtained according to the content.

7. The live voice interaction method according to claim 1, It is characterized in that In the case where the working state is the voice interaction state, in response to receiving the interaction trigger information, before generating the circulation information interaction content of the target object based on the interaction trigger information and the interaction management information, the method further includes: After generating the interaction management information, triggering the voice output device to broadcast fourth information, the fourth information is used to request the user's interaction trigger parameter, the interaction trigger parameter includes at least one of the following: interaction trigger mode, interaction trigger account, interaction trigger password, interaction trigger password, interaction trigger biometric information; The interaction trigger information is obtained according to the received interaction trigger parameter.

8. The live voice interaction method according to claim 7, It is characterized in that The generating the target object's circulation information interaction content based on the interaction trigger information and the interaction management information includes: triggering the voice output device to broadcast fifth information, where the fifth information is used to request the user's address information; Parsing the received third target voice to obtain the address information; The circulation information interaction content is generated according to the interaction management information, the interaction trigger information and the address information.

9. The live voice interaction method according to claim 8, It is characterized in that The step of parsing the received third target voice to obtain the address information includes: Get address template information; The parsing result of the third target voice is matched with the address template information to obtain the address information.

10. The live voice interaction method according to claim 1, It is characterized in that Before generating the interactive management information for the target object based on the circulation information, the method further includes: triggering the voice output device to broadcast each circulation information item in the circulation information; The generating the interactive management information for the target object based on the circulation information includes: upon receiving a confirmation instruction for each of the circulation information items in the circulation information, generating the interactive management information for the target object based on the circulation information.

11. The live voice interaction method according to claim 1, It is characterized in that Before generating the interaction management information for the target object based on the circulation information, the method further includes: In response to receiving the fourth target speech, parsing the fourth target speech; The target content related to the target object is determined according to the analysis result of the fourth target voice, and the voice output device is triggered to broadcast the associated information corresponding to the target content, where the associated information is used to explain or give suggestions for the target content.

12. A live voice interaction device, It is characterized in that include: The switching module is configured to execute, when the live broadcast application is opened, triggering the voice output device to broadcast the first information and wake up the voice input device, wherein the first information is used to inquire the user whether to configure the live broadcast application to the voice interaction state; parsing the feedback information received by the voice input device for the first information to obtain a first parsing result; In the case where the first analysis result indicates that the user has a willingness to interact with the voice, in response to receiving a voice interaction mode wake-up message, the working state of the live broadcast application is configured as a voice interaction state, wherein the voice interaction state is a state in which an object is determined by analyzing information received by a voice input device and a liquidity information interaction is performed for the object; The voice processing module is configured to execute, when the working state is the voice interaction state, in response to the live broadcast process entering a period of liquidity information interaction corresponding to the target object, triggering the voice output device to broadcast second information, wherein the second information is used to inquire whether the user is associated with the target object; parsing feedback information received by the voice input device for the second information to obtain a second parsing result; In the case where the second analysis result indicates that the user has an association intention, receiving the circulation information, and generating interaction management information for the target object based on the circulation information; When the working state is the voice interaction state, in response to receiving the interaction trigger information, the liquidity information interaction content of the target object is generated based on the interaction trigger information and the interaction management information to provide a full-process live e-commerce service based on voice interaction.

13. The live voice interaction device according to claim 12, It is characterized in that The voice processing module is configured to maintain the wake-up state of the voice input device when the working state is the voice interaction state.

14. The live voice interaction device according to claim 12, Features: The switching module is configured to execute, when the working state is the voice interaction state, in response to receiving voice interaction mode closing information, configuring the working state to a character interaction state and configuring the voice input device to a non-awakening state. The character interaction state is a state in which an object is determined through information received from the character input device and liquidity information interaction is performed on the object.

15. The live voice interaction device according to claim 12, It is characterized in that The speech processing module is configured to execute: In response to receiving a first target voice during the live broadcast, parsing the first target voice; In response to the existence of information corresponding to the target object in the parsing result, the circulation information is received.

16. The live voice interaction device according to claim 12 or 15, It is characterized in that The speech processing module is configured to execute: Obtain at least one attribute item corresponding to the target object; For each attribute item, trigger the voice output device to broadcast third information corresponding to the attribute item, where the third information is used to request the user to select at least one attribute value corresponding to the attribute item; The feedback information received by the voice input device for the third information is analyzed to obtain the circulation information item corresponding to the attribute item.

17. The live voice interaction device according to claim 16, It is characterized in that The speech processing module is configured to execute: receiving a second target speech, and parsing the second target speech; Obtain at least one attribute item corresponding to the target object; The content corresponding to each of the attribute items is determined according to the analysis result, and the circulation information item corresponding to the attribute item is obtained according to the content.

18. The live voice interaction device according to claim 12, It is characterized in that The speech processing module is configured to execute: After generating the interaction management information, triggering the voice output device to broadcast fourth information, the fourth information is used to request the user's interaction trigger parameter, the interaction trigger parameter includes at least one of the following: interaction trigger mode, interaction trigger account, interaction trigger password, interaction trigger password, interaction trigger biometric information; The interaction trigger information is obtained according to the received interaction trigger parameter.

19. The live voice interaction device according to claim 18, It is characterized in that The speech processing module is configured to execute: triggering the voice output device to broadcast fifth information, where the fifth information is used to request the user's address information; Parsing the received third target voice to obtain the address information; The circulation information interaction content is generated according to the interaction management information, the interaction trigger information and the address information.

20. The live voice interaction device according to claim 19, It is characterized in that The speech processing module is configured to execute: Get address template information; The parsing result of the third target voice is matched with the address template information to obtain the address information.

21. The live voice interaction device according to claim 12, It is characterized in that The speech processing module is configured to execute: Triggering the voice output device to broadcast each circulation information item in the circulation information; When a confirmation instruction for each of the circulation information items in the circulation information is received, interaction management information for a target object is generated based on the circulation information.

22. The live voice interaction device according to claim 12, It is characterized in that The speech processing module is configured to execute: In response to receiving the fourth target speech, parsing the fourth target speech; The target content related to the target object is determined according to the analysis result of the fourth target voice, and the voice output device is triggered to broadcast the associated information corresponding to the target content, where the associated information is used to explain or give suggestions for the target content.

23. An electronic device, It is characterized in that include: processor; a memory for storing instructions executable by the processor; Wherein, the processor is configured to execute the instructions to implement the live voice interaction method as described in any one of claims 1 to 11.

24. A computer-readable storage medium, It is characterized in that When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the live voice interaction method as described in any one of claims 1 to 11.

25. A computer program product, It is characterized in that The computer program product includes a computer program, which is stored in a readable storage medium. At least one processor of a computer device reads and executes the computer program from the readable storage medium, so that the computer device performs the live voice interaction method as described in any one of claims 1 to 11.

Citation Information

Patent Citations

  • Voice meal ordering method and device, electronic equipment and storage medium

    CN111242721A

  • Interactive video display and generation method and device, equipment and storage medium

    CN111741368A

  • Method and device for providing commodity object information and electronic equipment

    CN113298585A