End-to-end cloud voice interaction methods, devices, electronic devices and storage media

By establishing an end-to-cloud communication link between the client and the cloud, the conversion and execution of voice data are realized, solving the problem of insufficient voice interaction capabilities of third-party applications in devices such as Internet TVs and set-top boxes, and improving the user experience.

CN118968998BActive Publication Date: 2025-10-31CHINA MOBILE GRP FUJIAN CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411033040.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-30
Publication Date
2025-10-31
Estimated Expiration
2044-07-30

AI Technical Summary

Technical Problem

In existing technologies, third-party applications for Internet TVs and set-top boxes cannot achieve effective end-to-end voice interaction functions due to insufficient cloudification and voice interaction capabilities.

Method used

An end-to-cloud communication link is established between the client and the cloud. User voice data is acquired through a voice acquisition unit, converted into text data, and transmitted to the cloud through the end-to-cloud communication suite and the cloud communication server. The service unit in the cloud performs corresponding operations based on the text data, thereby realizing the cloudification of applications.

Benefits of technology

It enables voice data conversion and operation execution between the client and the cloud, improves user experience, and meets the voice interaction needs of cloud applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118968998B_ABST
    Figure CN118968998B_ABST
Patent Text Reader

Abstract

This disclosure proposes an edge-cloud voice interaction method, device, electronic device, and storage medium. Executed by a client, the method includes: a voice acquisition unit acquiring first voice data from a user and sending the first voice data to a first service unit; the first service unit sending the first voice data to a voice cloud platform to obtain first text data generated by the voice cloud platform based on the first voice data; the first service unit sending the first text data to a first proxy service unit; and the first proxy service unit sending the first text data to a second proxy service unit in the cloud based on a communication link between the edge-cloud communication suite and the edge-cloud communication server in the cloud. Thus, the client can convert the acquired user voice data into text data and forward the text data to the cloud, enabling the cloud to perform corresponding business operations based on the text data, thereby realizing application cloudification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of communication technology, and in particular to an edge-cloud voice interaction method, device, electronic device, and storage medium. Background Technology

[0002] Currently, to simplify user operation processes and improve user experience, various internet TVs and set-top boxes provide voice interaction functions through remote controls with voice capabilities. When a user presses the voice button on the remote control and speaks the desired operation, the client collects the voice data and interacts with the voice cloud platform. Through technologies such as evaluation, speech recognition, speech synthesis, and semantic parsing, the user's voice is semantically recognized. The recognized semantics are then sent back to the third-party application through the voice assistant integrated within the application. The third-party application processes the semantics to perform operations such as skipping, playing, pausing, and exiting.

[0003] In related technologies, due to the large number of cloud-based third-party application businesses, and the fact that some third-party applications do not have voice interaction capabilities, it is impossible to realize the end-to-cloud voice interaction function. Summary of the Invention

[0004] This disclosure aims to at least partially address one of the technical problems in the related art.

[0005] Therefore, this disclosure proposes a cloud-based voice interaction method, device, electronic device, storage medium, and computer program product.

[0006] The edge-cloud voice interaction method proposed in the first aspect of this disclosure is executed by a client, wherein the client is provided with a first service unit, a voice acquisition unit, an edge-cloud communication suite, and a first proxy service unit, and the method includes:

[0007] The voice acquisition unit acquires the user's first voice data and sends the first voice data to the first service unit;

[0008] The first service unit sends the first voice data to the voice cloud platform to obtain the first text data generated by the voice cloud platform based on the first voice data;

[0009] The first service unit sends the first text data to the first agent service unit;

[0010] The first proxy service unit sends the first text data to the second proxy service unit in the cloud based on the communication link between the end-to-cloud communication suite and the end-to-cloud communication server in the cloud.

[0011] The end-to-cloud voice interaction method proposed in the second aspect of this disclosure is executed in the cloud, wherein the cloud provides an end-to-cloud communication server, a second proxy service unit, and a second service unit, and the method includes:

[0012] The second proxy service unit receives the first text data sent by the first proxy service unit of the client based on the communication link between the end-cloud communication server and the end-cloud communication suite of the client. The first text data is obtained by the client from the voice cloud platform based on the user's first voice data.

[0013] The second agent service unit sends the first text data to the second service unit;

[0014] The second service unit determines the target cloud application from multiple initial cloud applications in the cloud based on the first text data;

[0015] The second service unit generates and sends target operation instructions to the target cloud application based on the first text data.

[0016] The target cloud application responds to the target operation command, executes the target operation, and obtains the target operation result.

[0017] The edge-cloud voice interaction device proposed in the third aspect of this disclosure is executed by a client, wherein the client is provided with a first service unit, a voice acquisition unit, an edge-cloud communication suite, and a first proxy service unit, and the device includes:

[0018] The first acquisition module is used to control the voice acquisition unit to acquire the user's first voice data and send the first voice data to the first service unit.

[0019] The second acquisition module is used to control the first service unit to send the first voice data to the voice cloud platform in order to obtain the first text data generated by the voice cloud platform based on the first voice data.

[0020] The first sending module is used to control the first service unit to send the first text data to the first proxy service unit;

[0021] The second sending module is used to control the first agent service unit to send the first text data to the second agent service unit in the cloud based on the communication link between the end-to-cloud communication suite and the end-to-cloud communication server in the cloud.

[0022] The edge-cloud voice interaction device proposed in the fourth aspect embodiment of this disclosure is executed by the cloud, wherein the cloud provides an edge-cloud communication server, a second proxy service unit, and a second service unit, and the device includes:

[0023] The receiving module is used to control the second agent service unit to receive the first text data sent by the first agent service unit of the client based on the communication link between the end-cloud communication server and the end-cloud communication suite of the client. The first text data is obtained by the client from the voice cloud platform based on the user's first voice data.

[0024] The third sending module is used to control the second proxy service unit to send the first text data to the second service unit;

[0025] The determination module is used to control the second service unit to determine the target cloud application from multiple initial cloud applications in the cloud based on the first text data;

[0026] The generation module is used to control the second service unit to generate and send target operation instructions to the target cloud application based on the first text data.

[0027] The execution module is used to control the target cloud application to respond to the target operation instructions, execute the target operation, and obtain the target operation result.

[0028] The fifth aspect of this disclosure provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the end-to-cloud voice interaction method as proposed in the first aspect of this disclosure, or implements the end-to-cloud voice interaction method as proposed in the second aspect of this disclosure.

[0029] A sixth aspect of this disclosure provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the end-to-cloud voice interaction method as proposed in the first aspect of this disclosure, or implements the end-to-cloud voice interaction method as proposed in the second aspect of this disclosure.

[0030] The seventh aspect of this disclosure provides a computer program product that, when executed by an instruction processor, performs the end-to-cloud voice interaction method as proposed in the first aspect of this disclosure, or implements the end-to-cloud voice interaction method as proposed in the second aspect of this disclosure.

[0031] The edge-cloud voice interaction method, apparatus, electronic device, storage medium, and computer program product proposed in this disclosure have at least the following beneficial effects: the voice acquisition unit acquires the user's first voice data and sends the first voice data to the first service unit; the first service unit sends the first voice data to the voice cloud platform to obtain the first text data generated by the voice cloud platform based on the first voice data; the first service unit sends the first text data to the first proxy service unit; and the first proxy service unit sends the first text data to the second proxy service unit in the cloud based on the communication link between the edge-cloud communication suite and the edge-cloud communication server in the cloud. Thus, the client can convert the acquired user's voice data into text data and forward the text data to the cloud, so that the cloud can perform corresponding business operations based on the text data, thereby realizing application cloudification.

[0032] Additional aspects and advantages of this disclosure will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this disclosure. Attached Figure Description

[0033] The above and / or additional aspects and advantages of this disclosure will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, in which:

[0034] Figure 1 This is a schematic flowchart of an embodiment of the edge-cloud voice interaction method proposed in this disclosure;

[0035] Figure 2 This is a schematic diagram of the system architecture for edge-cloud voice interaction proposed in one embodiment of this disclosure;

[0036] Figure 3 This is a flowchart illustrating another embodiment of the edge-cloud voice interaction method proposed in this disclosure;

[0037] Figure 4 This is a schematic flowchart of an embodiment of the edge-cloud voice interaction method proposed in this disclosure;

[0038] Figure 5 This is a flowchart illustrating another embodiment of the edge-cloud voice interaction method proposed in this disclosure;

[0039] Figure 6 This is an interactive schematic diagram of an edge-cloud voice interaction method proposed in an embodiment of this disclosure.

[0040] Figure 7 This is a schematic diagram of the structure of an edge-cloud voice interaction device according to an embodiment of this disclosure;

[0041] Figure 8 This is a schematic diagram of the structure of an edge-cloud voice interaction device according to an embodiment of this disclosure;

[0042] Figure 9 A block diagram of an exemplary electronic device suitable for implementing embodiments of the present disclosure is shown. Detailed Implementation

[0043] Embodiments of this disclosure are described in detail below, with examples of embodiments illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are used only to explain this disclosure, and should not be construed as limiting this disclosure. Rather, embodiments of this disclosure include all variations, modifications, and equivalents falling within the spirit and scope of the appended claims.

[0044] The technical solutions provided in this disclosure are applicable to a variety of systems, especially 5G systems. For example, applicable systems may include Global System for Mobile Communication (GSM), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA) General Packet Radio Service (GPRS), Long Term Evolution (LTE), LTE Frequency Division Duplex (FDD), LTE Time Division Duplex (TDD), Long Term Evolution Advanced (LTE-A), Universal Mobile Telecommunication System (UMTS), and 5G New Radio (NR). All of these systems include terminals and network equipment. The systems may also include a core network component, such as Evolved Packet System (EPS) or 5G system (5GS).

[0045] Figure 1 This is a schematic flowchart of an embodiment of the edge-cloud voice interaction method proposed in this disclosure.

[0046] It should be noted that the execution subject of the end-to-cloud voice interaction method in this embodiment is the end-to-cloud voice interaction device. This device can be implemented by software and / or hardware, and it can be configured in a network device. There are no restrictions on this.

[0047] like Figure 1 As shown, the cloud-based voice interaction method includes:

[0048] S101: The voice acquisition unit acquires the user's first voice data and sends the first voice data to the first service unit.

[0049] Among them, see Figure 2 , Figure 2 This is a schematic diagram of the system architecture for end-to-cloud voice interaction proposed in an embodiment of this disclosure. The end-to-cloud voice interaction method described in this embodiment is for voice interaction between the client and the cloud. The client is equipped with a first service unit, a voice acquisition unit, an end-to-cloud communication suite, and a first proxy service unit. The cloud is equipped with an end-to-cloud communication server, a second proxy service unit, and a second service unit.

[0050] The client can be, for example, a set-top box; there are no restrictions on this.

[0051] The client-side cloud communication suite and the cloud-side cloud communication server establish a communication link, which enables communication between the user client and the cloud.

[0052] The voice acquisition unit can be, for example, an Android application package (APK), the first service unit can be, for example, a Software Development Kit (SDK), and the first proxy service unit can be, for example, a Speech SDK proxy; there are no restrictions on this.

[0053] The second proxy service unit can be, for example, a Yudian APK proxy, and the second service unit can be, for example, a Yudian SDK; there are no restrictions on this.

[0054] In this embodiment of the disclosure, the voice acquisition unit may be set in the client and is used to acquire the user's first voice data.

[0055] The first voice data can be voice data generated by the user to an external voice acquisition device (e.g., a remote control, a smart speaker, etc.).

[0056] In this embodiment of the disclosure, after the external voice acquisition device acquires the user's first voice data, it sends the first voice data to the voice acquisition unit. After the voice acquisition unit acquires the user's first voice data, it forwards the first voice data to the first service unit.

[0057] S102: The first service unit sends the first voice data to the voice cloud platform to obtain the first text data generated by the voice cloud platform based on the first voice data.

[0058] The voice cloud platform is used to parse and process voice data to obtain corresponding text data. It can also convert text data to obtain corresponding voice data.

[0059] In other words, in this embodiment of the present disclosure, the first service unit may send the first voice data to the voice cloud platform, and the voice cloud platform processes the first voice data to obtain the first text data corresponding to the first voice data.

[0060] S103: The first service unit sends the first text data to the first agent service unit.

[0061] In this embodiment of the disclosure, the first service unit sends the first voice data to the voice cloud platform to obtain the first text data generated by the voice cloud platform based on the first voice data, and then sends the first text data to the first agent service unit.

[0062] S104: The first proxy service unit sends the first text data to the second proxy service unit in the cloud based on the communication link between the end-to-cloud communication suite and the end-to-cloud communication server in the cloud.

[0063] In this embodiment of the disclosure, after the first service unit sends the first text data to the first proxy service unit, the first proxy service unit sends the first text data to the second proxy service unit in the cloud based on the communication link between the end-to-cloud communication suite and the end-to-cloud communication server in the cloud. Thus, the client can convert the collected voice data into text data through the voice cloud platform and send the text data to the cloud, so that the cloud application in the cloud can perform the target operation based on the text data.

[0064] The first proxy service unit in the client is responsible for data forwarding in the client. That is, when data communication is required between the client and the cloud, the first proxy service unit usually forwards the data to be forwarded to the second proxy service unit in the cloud through the communication link between the end-to-cloud communication suite and the end-to-cloud communication server in the cloud after obtaining the data to be forwarded. After receiving the data to be forwarded, the second proxy service unit in the cloud is responsible for sending the data to be forwarded to the data requester in the cloud.

[0065] In this embodiment, the voice acquisition unit acquires the user's first voice data and sends the first voice data to the first service unit. The first service unit sends the first voice data to the voice cloud platform to obtain the first text data generated by the voice cloud platform based on the first voice data. The first service unit sends the first text data to the first proxy service unit. The first proxy service unit sends the first text data to the second proxy service unit in the cloud based on the communication link between the end-to-cloud communication suite and the end-to-cloud communication server in the cloud. Thus, the client can convert the acquired user's voice data into text data and forward the text data to the cloud, so that the cloud can perform corresponding business operations based on the text data, thereby realizing application cloudification.

[0066] Figure 3 This is a flowchart illustrating another embodiment of the edge-cloud voice interaction method proposed in this disclosure.

[0067] like Figure 3 As shown, the cloud-based voice interaction method includes:

[0068] S301: Upon receiving the first startup command in the first display window, obtain user information related to the target cloud application.

[0069] In this embodiment of the disclosure, the client also has a first display window corresponding to the target cloud application. Taking the target cloud application as a video playback application as an example, the first display window may be, for example, the main interface of the video playback application, a poster displaying recommended playback content, the icon of the video playback application, etc., and there are no restrictions on this.

[0070] The first launch command is a command given by the user to launch the target cloud application. This first launch command can be, for example, a touch command given by the user to the first display window, or a voice command given by the user to instruct the target cloud application to launch, etc., and there are no restrictions on this.

[0071] Among them, user information related to the target cloud application can be, for example, the user's account information used to log in to the target cloud application, and there are no restrictions on this.

[0072] S302: The first agent service unit sends user information to the cloud based on the communication link.

[0073] In other words, in this embodiment of the present disclosure, when the first display window in the client receives the first start command, the client can obtain user information related to the target cloud application. When the first display window receives the first start command, it can obtain user information related to the target cloud application, and then the first agent service unit can send the user information to the cloud based on the communication link.

[0074] S303: The voice acquisition unit acquires the user's first voice data and sends the first voice data to the first service unit.

[0075] S304: The first service unit sends the first voice data to the voice cloud platform to obtain the first text data generated by the voice cloud platform based on the first voice data.

[0076] S305: The first service unit sends the first text data to the first agent service unit.

[0077] S306: The first proxy service unit sends the first text data to the second proxy service unit in the cloud based on the communication link between the end-to-cloud communication suite and the end-to-cloud communication server in the cloud.

[0078] For detailed descriptions of S303-S306, please refer to the above embodiments, which will not be repeated here.

[0079] S307: The first agent service unit receives the target operation result sent by the second agent service unit based on the communication link, wherein the target operation result is the operation result made by the target cloud application in the cloud based on the first text data.

[0080] The target operation result is the operation result made by the target cloud application in the cloud based on the first text data. Taking the target cloud application as a video playback application as an example, if the first text data is "play video A", the target operation result can be the video stream data of video A. If the first text data is "pause playback of video A", the target operation result can be the pause playback interface of video A. There are no restrictions on this.

[0081] In other words, in this embodiment of the present disclosure, the first agent service unit receives the target operation result sent by the second agent service unit based on the communication link. Then, the client can trigger the execution of the subsequent end-to-cloud voice interaction method based on the target operation result. For details, please refer to the following embodiments, which will not be repeated here.

[0082] S308: The first agent service unit sends the target operation result to the first service unit.

[0083] In this embodiment of the disclosure, after the first proxy service unit receives the target operation result sent by the second proxy service unit based on the communication link, the first proxy service unit sends the target operation result to the first service unit.

[0084] S309: The first service unit generates a second display window based on the target operation result, wherein the second display window includes: operation prompt information of the target cloud application, and / or at least one content of the target cloud application to be displayed.

[0085] The second display window includes: operation prompts for the target cloud application, and / or at least one piece of content to be displayed in the target cloud application. Taking a video playback application as an example, the operation prompts for the target cloud application may be, for example, an interface to be operated related to the target cloud application, such as pause playback, fast forward, or speed playback. The at least one piece of content to be displayed in the target cloud application may be, for example, the video stream to be played in the target cloud application, which includes all search content of a certain program to be played (e.g., B Variety Season 1, B Variety Season 2, etc.), without limitation.

[0086] In other words, in this embodiment of the present disclosure, after receiving the target operation result sent by the first proxy service unit, the first service unit will generate a second display window based on the target operation result.

[0087] S310: The first service unit controls the display device to display the second display window.

[0088] In this embodiment of the present disclosure, the first service unit generates a second display window based on the target operation result and displays the second display window on the display device.

[0089] S311: The first agent service unit receives the text description information corresponding to the target operation result sent by the second agent service unit based on the communication link.

[0090] The text description information is used to describe the result of the target operation. For example, if the target operation result is the video stream data of video A, the text description information could be "Playing video A". If the target operation result is to pause playing video A, the text description information could be "Pausing playing video A". There are no restrictions on this.

[0091] In other words, in this embodiment of the present disclosure, the first proxy service unit receives text description information corresponding to the target operation result sent by the second proxy service unit based on the communication link.

[0092] S312: The first agent service unit sends text description information to the first service unit.

[0093] In this embodiment of the disclosure, after receiving the text description information corresponding to the target operation result sent by the second agent service unit based on the communication link, the first agent service unit can send the text description information corresponding to the target operation result to the first service unit.

[0094] S313: The first service unit sends the text description information to the voice cloud platform to obtain the second voice data generated by the voice cloud platform based on the text description information.

[0095] Among them, the voice data corresponding to the text description information is the second voice data.

[0096] In other words, in this embodiment of the present disclosure, the first service unit may send text description information to the voice cloud platform, and the voice cloud platform may perform conversion processing on the text description information to convert the text description information in text form into voice form, which is a second voice processing.

[0097] S314: The first service unit sends the second voice data to the first agent service unit.

[0098] In this embodiment of the disclosure, the first service unit sends the text description information to the voice cloud platform to obtain the second voice data generated by the voice cloud platform based on the text description information. After that, the first service unit may send the second voice data to the first proxy service unit.

[0099] S315: The first agent service unit plays the second voice data.

[0100] In this embodiment of the disclosure, after the first service unit sends the second voice data to the first proxy service unit, the first proxy service unit plays the second voice data, thereby realizing end-to-end cloud voice interaction.

[0101] In this embodiment of the disclosure, when the first display window receives the first start command, user information related to the target cloud application is obtained. The first agent service unit sends the user information to the cloud via a communication link. The voice acquisition unit obtains the user's first voice data and sends the first voice data to the first service unit. The first service unit sends the first voice data to the voice cloud platform to obtain the first text data generated by the voice cloud platform based on the first voice data. The first service unit sends the first text data to the first agent service unit. The first agent service unit sends the first text data to the second agent service unit in the cloud via a communication link between the end-to-cloud communication suite and the end-to-cloud communication server in the cloud. The first agent service unit receives the target operation result sent by the second agent service unit via the communication link. The target operation result is the operation result made by the target cloud application in the cloud based on the first text data. The first agent service unit sends the target operation result to the first service unit. The first service unit generates a second display window based on the target operation result. The display window includes: operation prompts for the target cloud application, and / or at least one content to be displayed for the target cloud application. The first service unit controls the display device to display the second display window. The first proxy service unit receives text description information corresponding to the target operation result sent by the second proxy service unit based on the communication link. The first proxy service unit sends the text description information to the first service unit. The first service unit sends the text description information to the voice cloud platform to obtain the second voice data generated by the voice cloud platform based on the text description information. The first service unit sends the second voice data to the first proxy service unit. The first proxy service unit plays the second voice data. The client can convert the obtained user's voice data into text data and forward the text data to the cloud so that the cloud can perform corresponding business operations based on the text data, realize application cloudification, and receive the target operation result sent by the cloud and play the second voice data corresponding to the target operation result, thereby realizing interactive voice feedback, realizing voice interaction between the client and the cloud, and improving the user experience.

[0102] Figure 4 This is a schematic flowchart of an embodiment of the edge-cloud voice interaction method proposed in this disclosure.

[0103] It should be noted that the execution subject of the end-to-cloud voice interaction method in this embodiment is the end-to-cloud voice interaction device. This device can be implemented by software and / or hardware, and it can be configured in a network device. There are no restrictions on this.

[0104] like Figure 4 As shown, the cloud-based voice interaction method includes:

[0105] S401: The second proxy service unit receives the first text data sent by the first proxy service unit of the client based on the communication link between the end-cloud communication server and the end-cloud communication suite of the client. The first text data is obtained by the client from the voice cloud platform based on the user's first voice data.

[0106] The explanations of the same terms in this disclosure and the above embodiments can be found in the above embodiments, and will not be repeated here.

[0107] In this embodiment of the disclosure, see the above. Figure 2 The edge-cloud voice interaction method described in this embodiment is executed by the cloud.

[0108] In this embodiment of the disclosure, the second proxy service unit in the cloud receives the first text data sent by the first proxy service unit of the client based on the communication link between the end-cloud communication server and the end-cloud communication suite of the client.

[0109] S402: The second agent service unit sends the first text data to the second service unit.

[0110] In this embodiment of the disclosure, the second agent service unit can send first text data to the second service unit so that the second service unit can trigger the execution of subsequent end-to-cloud voice interaction methods based on the first text data. For details, please refer to the following embodiments, which will not be repeated here.

[0111] S403: The second service unit determines the target cloud application from multiple initial cloud applications in the cloud based on the first text data.

[0112] Among them, there are several initial cloud applications in the cloud.

[0113] In this embodiment of the disclosure, after receiving the first text data sent by the second agent service unit, the second service unit can determine the target cloud application that matches the first text data from multiple initial cloud applications in the cloud.

[0114] In this embodiment of the disclosure, the target cloud application is determined from multiple initial cloud applications in the cloud based on the first text data. This can be done by parsing the application identification information from the first text data and using the initial cloud application corresponding to the application identification information as the target application, or by parsing the user requirement information from the first text data and using the initial cloud application that can meet the user requirement information as the target cloud application. There are no restrictions on this.

[0115] S404: The second service unit generates and sends the target operation instruction to the target cloud application based on the first text data.

[0116] Among them, the target operation instruction is a control instruction issued by the second service unit to control the target cloud application to perform the target operation.

[0117] In this embodiment of the disclosure, after the second service unit determines the target cloud application from multiple initial cloud applications in the cloud based on the first text data, the second service unit generates and sends a target operation instruction to the target cloud application based on the first text data.

[0118] S405: The target cloud application responds to the target operation command, executes the target operation, and obtains the target operation result.

[0119] In this embodiment of the disclosure, the second service unit generates and sends a target operation instruction to the target cloud application based on the first text data.

[0120] In this embodiment of the disclosure, after receiving the target operation instruction, the target application can respond to the target operation instruction and execute the target operation to obtain the target operation result.

[0121] In this embodiment, the second proxy service unit receives first text data sent by the first proxy service unit of the client based on the communication link between the end-cloud communication server and the end-cloud communication suite of the client. The first text data is obtained by the client from the voice cloud platform based on the user's first voice data. The second proxy service unit sends the first text data to the second service unit. Based on the first text data, the second service unit determines the target cloud application from multiple initial cloud applications in the cloud. Based on the first text data, the second service unit generates and sends a target operation instruction to the target cloud application. The target cloud application responds to the target operation instruction and executes the target operation to obtain the target operation result. Thus, the cloud can execute the target operation based on the text data sent by the client corresponding to the user's voice data, obtaining a target operation result that meets the user's needs. This allows cloud applications in the cloud to execute target operations based on the user's voice without introducing voice interaction functionality, thereby improving the user experience and demonstrating high applicability.

[0122] Figure 5 This is a flowchart illustrating another embodiment of the edge-cloud voice interaction method proposed in this disclosure.

[0123] like Figure 5 As shown, the cloud-based voice interaction method includes:

[0124] S501: Receive user information sent by the first agent service unit of the client via the communication link.

[0125] In this embodiment of the disclosure, the cloud can receive user information sent by the first agent service unit of the client via the communication link.

[0126] S502: Authenticate user information and, if authentication is successful, determine the virtual machine corresponding to the client, whereby the virtual machine provides the runtime environment for the target cloud application.

[0127] In this process, the cloud can allocate a corresponding virtual machine to each client, and the same target cloud application can run simultaneously in multiple virtual machines to meet the needs of multiple clients at the same time.

[0128] In this embodiment of the disclosure, after the cloud receives the user information sent by the first agent service unit of the client via the communication link, the management platform in the cloud can perform authentication processing on the user information, and if the authentication is successful, determine the virtual machine corresponding to the client to provide a running environment for the target cloud application.

[0129] S503: The second proxy service unit receives the first text data sent by the first proxy service unit of the client based on the communication link between the end-cloud communication server and the end-cloud communication suite of the client. The first text data is obtained by the client from the voice cloud platform based on the user's first voice data.

[0130] S504: The second agent service unit sends the first text data to the second service unit.

[0131] S505: The second service unit determines the target cloud application from multiple initial cloud applications in the cloud based on the first text data.

[0132] S506: The second service unit generates and sends target operation instructions to the target cloud application based on the first text data.

[0133] S507: The target cloud application responds to the target operation command, executes the target operation, and obtains the target operation result.

[0134] For a detailed description of S503-S507, please refer to the above embodiments, which will not be repeated here.

[0135] S508: The target cloud application sends the target operation result to the second service unit.

[0136] In this embodiment of the disclosure, the target cloud application responds to the target operation instruction, executes the target operation, and after obtaining the target operation result, sends the target operation result to the second service unit.

[0137] S509: The second service unit sends the target operation result to the second agent service unit.

[0138] In this embodiment of the disclosure, after the target cloud application sends the target operation result to the second service unit, the second service unit sends the target operation result to the second proxy service unit so that the second proxy server can forward the target operation result to the client.

[0139] S510: The second agent service unit sends the target operation result to the first agent service unit based on the communication link.

[0140] In this embodiment of the disclosure, after receiving the target operation result sent by the second service unit, the second proxy service unit can send the target operation result to the first proxy service unit based on the communication link.

[0141] S511: The second service unit generates text description information based on the target operation result.

[0142] In this embodiment of the disclosure, after the second agent service unit sends the target operation result to the first agent service unit based on the communication link, the second service unit generates text description information based on the target operation result.

[0143] S512: The second service unit sends the text description information to the second agent service unit.

[0144] In this embodiment of the disclosure, after the second service unit generates text description information based on the target operation result, the second service unit sends the text description information to the second proxy service unit.

[0145] S513: The second agent service unit sends text description information to the first agent service unit based on the communication link.

[0146] In this embodiment of the disclosure, after the second service unit sends the text description information to the second proxy service unit, the second proxy service unit sends the text description information to the first proxy service unit based on the communication link.

[0147] In this embodiment, the cloud receives user information sent by the first proxy service unit of the client via a communication link, authenticates the user information, and determines the virtual machine corresponding to the client if the authentication is successful. The virtual machine provides the runtime environment for the target cloud application. The second proxy service unit receives first text data sent by the first proxy service unit of the client based on the communication link between the end-cloud communication server and the end-cloud communication suite of the client. The first text data is obtained by the client from the voice cloud platform based on the user's first voice data. The second proxy service unit sends the first text data to the second service unit, which then determines the target cloud application from multiple initial cloud applications based on the first text data. Based on the first text data, the system generates and sends a target operation instruction to the target cloud application. The target cloud application responds to the instruction, executes the target operation, and obtains the result. The target cloud application then sends the result to the second service unit, which in turn sends it to the second proxy service unit. The second proxy service unit then sends the result to the first proxy service unit via a communication link. Based on the result, the second service unit generates text description information and sends it to the second proxy service unit. The second proxy service unit then sends the text description information to the first proxy service unit via the communication link. This enables interactive voice feedback, facilitating voice interaction between the client and the cloud, and improving the user experience.

[0148] Figure 6 This is an interactive schematic diagram of an end-to-cloud voice interaction method proposed in an embodiment of this disclosure.

[0149] like Figure 6 As shown, the cloud-based voice interaction method includes:

[0150] S601: When the client receives the first startup command in the first display window, it obtains user information related to the target cloud application.

[0151] S602: The first agent service unit sends user information to the second agent service unit in the cloud based on the communication link.

[0152] S603: The second agent service unit sends user information to the management platform.

[0153] S604: The management platform authenticates user information and, if the authentication is successful, determines the virtual machine corresponding to the client. The virtual machine provides the runtime environment for the target cloud application.

[0154] S605: The client's voice acquisition unit acquires the user's first voice data and sends the first voice data to the first service unit.

[0155] S606: The first service unit sends the first voice data to the voice cloud platform to obtain the first text data generated by the voice cloud platform based on the first voice data.

[0156] S607: The first service unit sends the first text data to the first agent service unit.

[0157] S608: The first agent service unit sends the first text data to the second agent service unit in the cloud based on the communication link between the end-to-cloud communication suite and the end-to-cloud communication server in the cloud.

[0158] S609: The second agent service unit sends the first text data to the second service unit.

[0159] S610: The second service unit determines the target cloud application from multiple initial cloud applications in the cloud based on the first text data.

[0160] S611: The second service unit generates and sends target operation instructions to the target cloud application based on the first text data.

[0161] S612: The target cloud application responds to the target operation command, executes the target operation, and obtains the target operation result.

[0162] S613: The second service unit sends the target operation result to the second agent service unit.

[0163] S614: The second agent service unit sends the target operation result to the first agent service unit based on the communication link.

[0164] S615: The first agent service unit sends the target operation result to the first service unit.

[0165] S616: The first service unit generates a second display window based on the result of the target operation.

[0166] S617: The first service unit controls the display device to display the second display window.

[0167] S618: The second service unit generates text description information based on the target operation result.

[0168] S619: The second service unit sends the text description information to the second agent service unit.

[0169] S620: The second agent service unit sends text description information to the first agent service unit based on the communication link.

[0170] S621: The first agent service unit sends text description information to the first service unit.

[0171] S622: The first service unit sends the text description information to the voice cloud platform to obtain the second voice data generated by the voice cloud platform based on the text description information.

[0172] S623: The first service unit sends the second voice data to the first agent service unit.

[0173] S624: The first agent service unit plays the second voice data.

[0174] Figure 7 This is a schematic diagram of the structure of an edge-cloud voice interaction device proposed in an embodiment of this disclosure.

[0175] like Figure 7 As shown, the edge-cloud voice interaction device 70 is executed by a client. The client includes a first service unit, a voice acquisition unit, an edge-cloud communication suite, and a first proxy service unit. The device includes:

[0176] The first acquisition module 701 is used to control the voice acquisition unit to acquire the user's first voice data and send the first voice data to the first service unit.

[0177] The second acquisition module 702 is used to control the first service unit to send the first voice data to the voice cloud platform in order to acquire the first text data generated by the voice cloud platform based on the first voice data;

[0178] The first sending module 703 is used to control the first service unit to send the first text data to the first proxy service unit;

[0179] The second sending module 704 is used to control the first agent service unit to send the first text data to the second agent service unit in the cloud based on the communication link between the end-to-cloud communication suite and the end-to-cloud communication server in the cloud.

[0180] In some embodiments of this disclosure, the client also includes a first display window corresponding to the target cloud application;

[0181] Among them, the edge-cloud voice interaction device 70 is also used for:

[0182] Upon receiving the first startup command in the first display window, obtain user information related to the target cloud application;

[0183] The first agent service unit sends user information to the cloud based on the communication link.

[0184] In some embodiments of this disclosure, the edge-cloud voice interaction device 70 is further configured to:

[0185] The first agent service unit receives the target operation result sent by the second agent service unit based on the communication link. The target operation result is the operation result made by the target cloud application in the cloud based on the first text data.

[0186] The first agent service unit sends the target operation result to the first service unit;

[0187] The first service unit generates a second display window based on the target operation result. The second display window includes: operation prompt information of the target cloud application, and / or at least one content of the target cloud application to be displayed.

[0188] The first service unit controls the display device to display the second display window.

[0189] In some embodiments of this disclosure, the edge-cloud voice interaction device 70 is further configured to:

[0190] The first agent service unit receives text description information corresponding to the target operation result sent by the second agent service unit based on the communication link;

[0191] The first agent service unit sends text description information to the first service unit;

[0192] The first service unit sends the text description information to the voice cloud platform to obtain the second voice data generated by the voice cloud platform based on the text description information;

[0193] The first service unit sends the second voice data to the first agent service unit;

[0194] The first agent service unit plays the second voice data.

[0195] With the above Figures 1 to 4 Corresponding to the edge-cloud voice interaction method provided in the embodiments, this disclosure also provides an edge-cloud voice interaction device. Because the edge-cloud voice interaction device provided in the embodiments of this disclosure is similar to the one described above... Figures 1 to 2 The implementation of the edge-cloud voice interaction method provided in the embodiments corresponds to the edge-cloud voice interaction device proposed in the embodiments of this disclosure, and will not be described in detail in the embodiments of this disclosure.

[0196] In this embodiment, the voice acquisition unit acquires the user's first voice data and sends the first voice data to the first service unit. The first service unit sends the first voice data to the voice cloud platform to obtain the first text data generated by the voice cloud platform based on the first voice data. The first service unit sends the first text data to the first proxy service unit. The first proxy service unit sends the first text data to the second proxy service unit in the cloud based on the communication link between the end-to-cloud communication suite and the end-to-cloud communication server in the cloud. Thus, the client can convert the acquired user's voice data into text data and forward the text data to the cloud, so that the cloud can perform corresponding business operations based on the text data, thereby realizing application cloudification.

[0197] Figure 8 This is a schematic diagram of the structure of an edge-cloud voice interaction device proposed in an embodiment of this disclosure.

[0198] like Figure 8 As shown, the edge-cloud voice interaction device 80 is executed by the cloud, wherein the cloud includes an edge-cloud communication server, a second proxy service unit, and a second service unit. The device includes:

[0199] The receiving module 801 is used to control the second proxy service unit to receive the first text data sent by the first proxy service unit of the client based on the communication link between the end-cloud communication server and the end-cloud communication suite of the client. The first text data is obtained by the client from the voice cloud platform based on the user's first voice data.

[0200] The third sending module 802 is used to control the second proxy service unit to send the first text data to the second service unit;

[0201] The determination module 803 is used to control the second service unit to determine the target cloud application from multiple initial cloud applications in the cloud based on the first text data;

[0202] The generation module 804 is used to control the second service unit to generate and send target operation instructions to the target cloud application based on the first text data.

[0203] The execution module 805 is used to control the target cloud application to respond to the target operation command, execute the target operation, and obtain the target operation result.

[0204] In some embodiments of this disclosure, the edge-cloud voice interaction device 80 is further configured to:

[0205] After the target cloud application responds to the target operation instruction, executes the target operation, and obtains the target operation result, the target cloud application sends the target operation result to the second service unit.

[0206] The second service unit sends the target operation result to the second agent service unit;

[0207] The second agent service unit sends the target operation result to the first agent service unit based on the communication link.

[0208] In some embodiments of this disclosure, the edge-cloud voice interaction device 80 is further configured to:

[0209] Receive user information sent by the client's first proxy service unit via the communication link;

[0210] User information is authenticated, and if the authentication is successful, the corresponding virtual machine is determined, whereby the virtual machine provides the runtime environment for the target cloud application.

[0211] In some embodiments of this disclosure, the edge-cloud voice interaction device 80 is further configured to:

[0212] After the second agent service unit sends the target operation result to the first agent service unit based on the communication link, the second service unit generates text description information based on the target operation result;

[0213] The second service unit sends the text description information to the second agent service unit;

[0214] The second agent service unit sends text description information to the first agent service unit based on the communication link.

[0215] With the above Figures 4 to 5 Corresponding to the edge-cloud voice interaction method provided in the embodiments, this disclosure also provides an edge-cloud voice interaction device. Because the edge-cloud voice interaction device provided in the embodiments of this disclosure is similar to the one described above... Figures 2 to 4 The implementation of the edge-cloud voice interaction method provided in the embodiments corresponds to the edge-cloud voice interaction device proposed in the embodiments of this disclosure, and will not be described in detail in the embodiments of this disclosure.

[0216] In this embodiment, the second proxy service unit receives first text data sent by the first proxy service unit of the client based on the communication link between the end-cloud communication server and the end-cloud communication suite of the client. The first text data is obtained by the client from the voice cloud platform based on the user's first voice data. The second proxy service unit sends the first text data to the second service unit. Based on the first text data, the second service unit determines the target cloud application from multiple initial cloud applications in the cloud. Based on the first text data, the second service unit generates and sends a target operation instruction to the target cloud application. The target cloud application responds to the target operation instruction and executes the target operation to obtain the target operation result. Thus, the cloud can execute the target operation based on the text data sent by the client corresponding to the user's voice data, obtaining a target operation result that meets the user's needs. This allows cloud applications in the cloud to execute target operations based on the user's voice without introducing voice interaction functionality, thereby improving the user experience and demonstrating high applicability.

[0217] To implement the above embodiments, this disclosure also proposes an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the end-to-cloud voice interaction method proposed in the foregoing embodiments of this disclosure.

[0218] To implement the above embodiments, this disclosure also proposes a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the end-to-cloud voice interaction method as proposed in the foregoing embodiments of this disclosure.

[0219] To implement the above embodiments, this disclosure also proposes a computer program product that, when the instruction processor in the computer program product is executed, performs the end-to-cloud voice interaction method as proposed in the foregoing embodiments of this disclosure.

[0220] Figure 9 A block diagram of an exemplary electronic device suitable for implementing embodiments of the present disclosure is shown. Figure 9 The electronic device 12 shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments disclosed herein.

[0221] like Figure 9 As shown, the electronic device 12 is represented in the form of a general-purpose computing device. The components of the electronic device 12 may include, but are not limited to: one or more processors or processing units 16, system memory 28, and bus 18 connecting different system components (including system memory 28 and processing unit 16).

[0222] Bus 18 represents one or more of several bus architectures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the various bus architectures. Examples of these architectures include, but are not limited to, the Industry Standard Architecture (ISA) bus, the Micro Channel Architecture (MAC) bus, the Enhanced ISA bus, the Video Electronics Standards Association (VESA) local bus, and the Peripheral Component Interconnect (PCI) bus.

[0223] Electronic device 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by electronic device 12, including volatile and non-volatile media, removable and non-removable media.

[0224] Memory 28 may include computer system readable media in the form of volatile memory, such as Random Access Memory (RAM) 30 and / or cache memory 32. Electronic device 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 34 may be used to read and write non-removable, non-volatile magnetic media (… Figure 9 Not shown; usually referred to as a "hard drive".

[0225] although Figure 9 Not shown, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk") and an optical disc drive for reading and writing to a removable non-volatile optical disc (e.g., a compact disc read-only memory (CD-ROM), a digital video disc read-only memory (DVD-ROM), or other optical media) may be provided. In these cases, each drive may be connected to bus 18 via one or more data media interfaces. Memory 28 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the embodiments of this disclosure.

[0226] A program / utility 40 having a set (at least one) of program modules 42 may be stored, for example, in memory 28. Such program modules 42 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. Program modules 42 typically perform the functions and / or methods described in the embodiments of this disclosure.

[0227] Electronic device 12 can also communicate with one or more external devices 14 (e.g., keyboard, pointing device, display 24, etc.), and with one or more devices that enable a user to interact with electronic device 12, and / or with any device that enables electronic device 12 to communicate with one or more other computing devices (e.g., network card, modem, etc.). This communication can be performed via input / output (I / O) interface 22. Furthermore, electronic device 12 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 20. As shown, network adapter 20 communicates with other modules of electronic device 12 via bus 18. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with electronic device 12, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0228] The processing unit 16 executes various functional applications and edge-cloud voice interaction by running programs stored in the system memory 28, such as implementing the edge-cloud voice interaction method mentioned in the foregoing embodiments.

[0229] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.

[0230] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

[0231] It should be noted that in the description of this disclosure, the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. Furthermore, in the description of this disclosure, unless otherwise stated, "a plurality of" means two or more.

[0232] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing a particular logical function or process, and the scope of preferred embodiments of this disclosure includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the function involved, as will be understood by those skilled in the art to which embodiments of this disclosure pertain.

[0233] It should be understood that various parts of this disclosure can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0234] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.

[0235] Furthermore, the functional units in the various embodiments of this disclosure can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0236] The storage media mentioned above can be read-only memory, disk, or optical disk, etc.

[0237] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this disclosure. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0238] Although embodiments of the present disclosure have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present disclosure. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present disclosure.

Claims

1. A cloud-edge voice interaction method, characterized in that, Executed by a client, wherein the client is configured with a first service unit, a voice acquisition unit, an end-to-cloud communication suite, and a first proxy service unit, the method comprising: The voice acquisition unit acquires the user's first voice data and sends the first voice data to the first service unit. The first service unit sends the first voice data to the voice cloud platform to obtain the first text data generated by the voice cloud platform based on the first voice data; The first service unit sends the first text data to the first proxy service unit; The first proxy service unit sends the first text data to the second proxy service unit in the cloud based on the communication link between the end-to-cloud communication suite and the end-to-cloud communication server in the cloud.

2. The method as described in claim 1, characterized in that, The client also includes a first display window corresponding to the target cloud application; The method further includes: Upon receiving a first launch command in the first display window, obtain user information related to the target cloud application. The first agent service unit sends the user information to the cloud based on the communication link.

3. The method as described in claim 2, characterized in that, The method further includes: The first proxy service unit receives the target operation result sent by the second proxy service unit based on the communication link, wherein the target operation result is the operation result made by the target cloud application in the cloud based on the first text data; The first proxy service unit sends the target operation result to the first service unit; The first service unit generates a second display window based on the target operation result, wherein the second display window includes: operation prompt information of the target cloud application, and / or at least one content to be displayed of the target cloud application; The first service unit controls the display device to display the second display window.

4. The method as described in claim 3, characterized in that, The method further includes: The first proxy service unit receives text description information corresponding to the target operation result sent by the second proxy service unit based on the communication link; The first proxy service unit sends the text description information to the first service unit; The first service unit sends the text description information to the voice cloud platform to obtain the second voice data generated by the voice cloud platform based on the text description information; The first service unit sends the second voice data to the first proxy service unit; The first agent service unit plays the second voice data.

5. A cloud-based voice interaction method, characterized in that, The method is executed in the cloud, wherein the cloud includes an end-to-cloud communication server, a second proxy service unit, and a second service unit. The second proxy service unit receives the first text data sent by the first proxy service unit of the client based on the communication link between the end-cloud communication server and the end-cloud communication suite of the client. The first text data is obtained by the client from the voice cloud platform based on the user's first voice data. The second proxy service unit sends the first text data to the second service unit; The second service unit determines the target cloud application from multiple initial cloud applications in the cloud based on the first text data; The second service unit generates and sends a target operation instruction to the target cloud application based on the first text data; The target cloud application responds to the target operation instruction, executes the target operation, and obtains the target operation result.

6. The method as described in claim 5, characterized in that, After the target cloud application responds to the target operation instruction, executes the target operation, and obtains the target operation result, the method further includes: The target cloud application sends the target operation result to the second service unit; The second service unit sends the target operation result to the second proxy service unit; The second proxy service unit sends the target operation result to the first proxy service unit based on the communication link.

7. The method as described in claim 5, characterized in that, The method further includes: Receive user information sent by the first proxy service unit of the client via the communication link; The user information is authenticated, and if the authentication is successful, the virtual machine corresponding to the client is determined, wherein the virtual machine provides the runtime environment for the target cloud application.

8. The method as described in claim 6, characterized in that, After the second proxy service unit sends the target operation result to the first proxy service unit based on the communication link, the process further includes: The second service unit generates text description information based on the target operation result; The second service unit sends the text description information to the second proxy service unit; The second proxy service unit sends the text description information to the first proxy service unit based on the communication link.

9. A cloud-based voice interaction device, characterized in that, Executed by a client, wherein the client is equipped with a first service unit, a voice acquisition unit, an end-to-cloud communication suite, and a first proxy service unit, and the device includes: The first acquisition module is used to control the voice acquisition unit to acquire the user's first voice data and send the first voice data to the first service unit. The second acquisition module is used to control the first service unit to send the first voice data to the voice cloud platform in order to obtain the first text data generated by the voice cloud platform based on the first voice data. The first sending module is used to control the first service unit to send the first text data to the first proxy service unit; The second sending module is used to control the first proxy service unit to send the first text data to the second proxy service unit in the cloud based on the communication link between the end-to-cloud communication suite and the end-to-cloud communication server in the cloud.

10. A cloud-based voice interaction device, characterized in that, Executed in the cloud, wherein the cloud includes an end-to-cloud communication server, a second proxy service unit, and a second service unit, and the device includes: The receiving module is used to control the second proxy service unit to receive the first text data sent by the first proxy service unit of the client based on the communication link between the end-cloud communication server and the end-cloud communication suite of the client, wherein the first text data is obtained by the client from the voice cloud platform based on the user's first voice data; The third sending module is used to control the second proxy service unit to send the first text data to the second service unit; The determining module is used to control the second service unit to determine the target cloud application from multiple initial cloud applications in the cloud based on the first text data; The generation module is used to control the second service unit to generate and send target operation instructions to the target cloud application based on the first text data; The execution module is used to control the target cloud application to respond to the target operation instruction, execute the target operation, and obtain the target operation result.

11. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-8.

12. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-8.

13. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Voice interaction method, device, system and equipment for protecting privacy and storage medium

    CN113472806A

  • Human-computer interaction method, device and equipment and storage medium

    CN113674742A