Device operation method, apparatus, device, medium, and program product
By recognizing user intent through the multimodal gateway and voice service central control of the robot platform, and generating device operation commands, the problem of cumbersome operation of IoT devices is solved, enabling device operation without the need for an application and improving the user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JD DIGITS HAIYI INFORMATION TECHNOLOGY CO LTD
- Filing Date
- 2025-01-26
- Publication Date
- 2026-07-28
AI Technical Summary
In existing technologies, the operation process of IoT devices is mostly based on applications or touch interaction, which is cumbersome and complex.
The robot platform uses a multimodal gateway, voice service control center, and model library to recognize voice intent, determine device identification and type, generate operation instructions, and enable device operation without requiring users to install or operate any applications.
It simplifies the operation process of IoT devices and improves the user experience, especially for users who are unfamiliar with the operation process, providing an intuitive and easy-to-interact operation method.
Smart Images

Figure CN122474048A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the fields of artificial intelligence, intelligent robots and the Internet of Things, and more specifically, to methods of operating devices, apparatus, equipment, media and program products. Background Technology
[0002] With the widespread application of IoT devices, controlling and managing these devices is a crucial means of realizing intelligent applications in the IoT field. To achieve control over IoT devices, relevant operations must first be performed on them, such as network configuration, activation, binding, and unbinding.
[0003] In the process of realizing the present invention, the inventors discovered that the related technologies have at least the following problems: the operation process of IoT devices in the related technologies is mostly based on applications or touch interaction, which is long and cumbersome. Summary of the Invention
[0004] In view of the above, this disclosure provides a method, apparatus, device, medium and program product for operating a device.
[0005] One aspect of this disclosure provides a device operation method, comprising: sending first voice information collected from a user to a robot platform, so that the robot platform can access the first voice information based on a multimodal gateway, and call a target large model from a model library through a voice service central control to perform intent recognition on the first voice information to obtain a first recognition result, wherein the robot platform includes a multimodal gateway, a voice service central control, and a model library; when it is determined that the first recognition result indicates that the user needs to operate the device, sending second voice information to the user, wherein the second voice information is used to represent a first question to the user regarding a device identifier; collecting third voice information from the user, wherein the third voice information is used to represent a first answer matching the first question; sending the third voice information to the robot platform, so that the robot platform can access the third voice information based on a multimodal gateway, and call a target large model from a model library through a voice service central control to perform intent recognition on the third voice information to obtain a second recognition result; when it is determined that the second recognition result indicates that the user has provided a device identifier, determining the device type based on the device identifier; and generating an operation command based on the device identifier, device type, and operation type to complete the operation of the device.
[0006] According to embodiments of this disclosure, an operation instruction is generated based on a device identifier, device type, and operation type to complete the operation of the device, including: determining whether to invoke the device provider's application based on the device type and operation type; if it is determined that no application invocation is required, generating an operation instruction based on the device identifier and operation type to complete the operation of the device; if it is determined that an application invocation is required, invoking the application based on the device identifier so that the application can control the device to turn on Bluetooth; connecting via Bluetooth and sending a connection result to the application so that the application can return operation information matching the device if it determines that the connection result indicates a successful connection; and generating an operation instruction based on the operation information to complete the operation of the device.
[0007] According to embodiments of this disclosure, determining whether to invoke the device provider's application based on the device type and operation type includes: determining that no application needs to be invoked when the device type is determined to be a non-directly connected peripheral device; or determining that an application needs to be invoked when the device type is determined to be a third-party cloud device and the operation type is a network configuration operation; or determining that no application needs to be invoked when the device type is determined to be a third-party cloud device and the operation type is a binding operation or an unbinding operation.
[0008] According to embodiments of this disclosure, when it is determined that no application needs to be invoked, an operation instruction is generated based on the device identifier and operation type to complete the operation on the device, including: when it is determined that no application needs to be invoked and the operation type is a network distribution operation, sending a network distribution operation request matching the device identifier to a robot platform; receiving network distribution operation information matching the network distribution operation request sent by the robot platform; generating an operation instruction for connecting the device based on the network distribution operation information; and sending the operation instruction to the device to complete the network distribution operation.
[0009] According to embodiments of this disclosure, when it is determined that no application needs to be invoked, an operation instruction is generated based on the device identifier and operation type to complete the operation on the device. This includes: when it is determined that no application needs to be invoked and the operation type is a binding operation, for a non-directly connected peripheral device, the following operations are performed: sending a network configuration query request to the robot platform, so that the robot platform responds to the network configuration query request and performs a binding operation based on the authentication result of the non-directly connected peripheral device; and when it is determined that the network configuration success and binding success results sent by the robot platform are received, the network configuration success and binding success are fed back to the user in voice form; for a third-party cloud device, the following operations are performed: sending the result of the completed network configuration operation to the application, so that the application sends a binding instruction to the robot platform; and receiving a binding success result sent by the robot platform after executing the binding instruction, and feeding back the binding success result to the user in voice form.
[0010] According to embodiments of this disclosure, the operation type includes an unbinding operation; wherein, based on the device identifier, device type, and operation type, an operation instruction is generated to complete the operation on the device, including: sending a fourth voice message to the user, wherein the fourth voice message represents a second question to confirm whether the device identifier is correct; collecting the user's fifth voice message, wherein the fifth voice message represents a second answer matching the second question; sending the fifth voice message to a robot platform, so that the robot platform can call a target large model from the model library through the voice service central control to perform intent recognition on the fifth voice message to obtain a third recognition result; if it is determined that the third recognition result represents that the user confirms the device identifier is correct, generating an unbinding operation instruction for the device corresponding to the device identifier, and sending the unbinding operation instruction to the robot platform; and receiving a successful unbinding result sent by the robot platform after performing the unbinding operation, and feeding back the successful unbinding to the user in the form of voice.
[0011] Another aspect of this disclosure provides a device operation method, comprising: accessing first voice information of a user sent by a robot based on a multimodal gateway; performing intent recognition on the first voice information by calling a target large model from a model library based on a voice service central control to obtain a first recognition result; sending the first recognition result to the robot so that the robot can send second voice information to the user, wherein the second voice information is used to represent a first question to the user regarding a device identifier; receiving third voice information sent by the robot, wherein the third voice information is used to represent a first answer matching the first question, and the third voice information is collected by the robot from the user; performing intent recognition on the third voice information by calling a target large model from a model library based on a voice service central control to obtain a second recognition result; and sending the second recognition result to the robot so that, if the robot determines that the second recognition result represents that the user has provided a device identifier, it generates an operation command based on the device identifier, device type, and operation type, wherein the device type is determined according to the device identifier.
[0012] Another aspect of this disclosure provides a device operation apparatus, comprising: a first sending module, configured to send first voice information collected from a user to a robot platform, so that the robot platform can access the first voice information based on a multimodal gateway and call a target large model from a model library through a voice service central control to perform intent recognition on the first voice information to obtain a first recognition result, wherein the robot platform includes a multimodal gateway, a voice service central control, and a model library; a second sending module, configured to send second voice information to the user when it is determined that the first recognition result indicates that the user needs to operate the device, wherein the second voice information is used to indicate a first question to the user regarding the device identifier; and a collection module. The system comprises the following modules: a third voice information module for collecting user's third voice information, which represents a first answer matching the first question; a third sending module for sending the third voice information to the robot platform, enabling the robot platform to access the third voice information via a multimodal gateway and call a target large model from the model library through the voice service central control to perform intent recognition on the third voice information to obtain a second recognition result; a determination module for determining the device type based on the device identifier, provided that the second recognition result represents that the user has provided a device identifier; and a generation module for generating operation instructions based on the device identifier, device type, and operation type to complete the operation of the device.
[0013] Another aspect of this disclosure provides a device operation apparatus, comprising: a first access module for accessing first voice information sent by a robot via a multimodal gateway; a first invocation module for invoking a target large model from a model library to perform intent recognition on the first voice information to obtain a first recognition result based on a voice service central control; a fourth sending module for sending the first recognition result to the robot so that the robot can send second voice information to the user, wherein the second voice information represents a first question to the user regarding a device identifier; a second access module for receiving third voice information sent by the robot, wherein the third voice information represents a first answer matching the first question, and the third voice information is collected by the robot from the user; a second invocation module for invoking a target large model from a model library to perform intent recognition on the third voice information to obtain a second recognition result based on a voice service central control; and a fifth sending module for sending the second recognition result to the robot so that, if the robot determines that the second recognition result represents that the user has provided a device identifier, it can generate an operation command based on the device identifier, device type, and operation type, wherein the device type is determined according to the device identifier.
[0014] Another aspect of this disclosure provides an electronic device comprising: one or more processors; and a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the methods described above.
[0015] Another aspect of this disclosure provides a computer-readable storage medium storing computer-executable instructions that, when executed, are used to implement the methods described above.
[0016] Another aspect of this disclosure provides a computer program product including computer-executable instructions that, when executed, are used to implement the methods described above.
[0017] According to embodiments of this disclosure, a robot-centric approach is used to interact with the user via voice. The robot platform's multimodal gateway, voice service control center, and model library are used to recognize the user's voice intent, accurately determine the device the user needs to operate, and generate operation commands to complete the device operation. Since the operation commands are generated based on the device identifier, device type, and operation type determined through voice interaction, there is no need for the user to install or operate any application corresponding to the device. This simplifies the device operation process and enhances the user experience. Furthermore, since the entire interaction between the user and the robot is conducted via voice, user operation is simplified, especially for users unfamiliar with the operation process. This provides a more intuitive and user-friendly device operation method. It at least partially solves the problem in related technologies where the operation process of IoT devices is mostly based on applications or touch-based interactions, resulting in long and cumbersome processes. Attached Figure Description
[0018] The above and other objects, features and advantages of this disclosure will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:
[0019] Figure 1A The illustrations illustrate application scenarios of the device operation method, apparatus, device, medium, and program product according to embodiments of the present disclosure.
[0020] Figure 1B A schematic diagram illustrating a device operation method according to an embodiment of the present disclosure is shown.
[0021] Figure 2 A flowchart illustrating a device operation method according to an embodiment of the present disclosure is shown schematically.
[0022] Figure 3A The illustration shows an interactive diagram of a network configuration operation for a non-directly connected peripheral device according to an embodiment of the present disclosure;
[0023] Figure 3B The illustration shows an interactive diagram of network configuration operations for third-party cloud devices according to embodiments of the present disclosure;
[0024] Figure 3C The illustration schematically shows an interactive diagram of a binding operation for a non-directly connected peripheral device according to an embodiment of the present disclosure;
[0025] Figure 3D The illustration schematically shows an interaction diagram of a binding operation for a third-party cloud device according to an embodiment of the present disclosure;
[0026] Figure 3EThis illustration schematically shows an interactive diagram of an unbinding operation according to an embodiment of the present disclosure;
[0027] Figure 4 A flowchart illustrating a method of operating a device according to another embodiment of the present disclosure is shown schematically;
[0028] Figure 5 A block diagram of a device operating apparatus according to an embodiment of the present disclosure is shown schematically;
[0029] Figure 6 A block diagram of a device operating apparatus according to another embodiment of the present disclosure is schematically shown; and
[0030] Figure 7 A block diagram schematically illustrates an electronic device suitable for implementing a device operation method according to an embodiment of the present disclosure. Detailed Implementation
[0031] The embodiments of the present disclosure will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the disclosure. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the present disclosure for ease of explanation. However, it will be apparent that one or more embodiments may be practiced without these specific details. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concepts of the present disclosure.
[0032] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. The terms “comprising,” “including,” etc., as used herein indicate the presence of features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0033] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.
[0034] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).
[0035] In the embodiments disclosed herein, the collection, updating, analysis, processing, use, transmission, provision, disclosure, and storage of data (e.g., including but not limited to user personal information) comply with relevant laws and regulations, are used for legitimate purposes, and do not violate public order and good morals. In particular, necessary measures have been taken to prevent unauthorized access to user personal information data and to safeguard user personal information security and network security.
[0036] In the embodiments disclosed herein, user authorization or consent is obtained before acquiring or collecting user personal information.
[0037] Embodiments of this disclosure provide a device operation method, apparatus, device, medium, and program product. The method includes: sending first voice information collected from a user to a robot platform, so that the robot platform accesses the first voice information via a multimodal gateway and performs intent recognition on the first voice information by calling a target large model from a model library through a voice service central control to obtain a first recognition result, wherein the robot platform includes a multimodal gateway, a voice service central control, and a model library; if the first recognition result indicates that the user needs to operate the device, sending second voice information to the user, wherein the second voice information is used to represent a first question to the user regarding a device identifier; collecting third voice information from the user, wherein the third voice information is used to represent a first answer matching the first question; sending the third voice information to the robot platform, so that the robot platform accesses the third voice information via a multimodal gateway and performs intent recognition on the third voice information by calling a target large model from a model library through a voice service central control to obtain a second recognition result; if the second recognition result indicates that the user has provided a device identifier, determining the device type based on the device identifier; and generating an operation command based on the device identifier, device type, and operation type to complete the operation of the device.
[0038] Figure 1A The illustrations illustrate application scenarios of the device operation method, apparatus, device, medium, and program product according to embodiments of the present disclosure.
[0039] like Figure 1A As shown, the application scenarios of this embodiment may include user 101, robot 102, robot platform 103, and Internet of Things device 104.
[0040] User 101 can interact with robot 102 via voice.
[0041] Robot 102 can interact with robot platform 103 via a network. For example, robot 102 can send user voice information to robot platform 103 via the network. Robot platform 103 performs intent recognition on the user's voice information and returns the recognition result to robot 102. Robot 102 may include, for example, a smart electronic device without a display screen. Robot platform 103 may include, for example, a cloud platform. The cloud platform may be a server providing various services, such as a backend management server that can support user voice information (for example only). The backend management server can process the received user voice information by performing intent recognition and other processing, and feed the recognition result back to robot 102. Robot platform 103 can manage various robots 102, etc.
[0042] Robot 102 can also interact with IoT device 104 via a network. The network can include various connection types, such as wired and / or wireless communication links, etc.
[0043] The Internet of Things (IoT) device 104 can include non-directly connected terminal devices and third-party cloud devices. For example, IoT device 104 may include, but is not limited to, sensors 1041, positioning devices 1042, vehicle terminals 1043, and smart devices 1044, etc., without specific limitations. Non-directly connected terminal devices can be understood as those that do not directly connect to a central server or cloud platform, but instead transmit data through intermediary devices (such as gateways or routers). Third-party cloud devices can be understood as those capable of interacting with the systems or platforms of multiple cloud service providers, enabling cross-cloud service data exchange and communication. The connection methods and cloud service interactivity of sensors 1041, positioning devices 1042, vehicle terminals 1043, and smart devices 1044 can be used for classification. For example, industrial sensors (pressure, vibration, flow) can be classified as non-directly connected terminal devices. Smart industrial sensors capable of sending data to multiple cloud platforms for analysis can be classified as third-party cloud devices.
[0044] It should be understood that Figure 1A The number of users, robots, robot platforms, and IoT devices shown is merely illustrative. Depending on implementation needs, there can be any number of users, robots, robot platforms, and IoT devices.
[0045] Figure 1B A schematic diagram illustrating a system architecture of a device operation method according to an embodiment of the present disclosure is provided.
[0046] like Figure 1BAs shown, the system architecture of this embodiment may include a robot terminal 120, a robot platform terminal 130, and an IoT device terminal 140. The robot terminal 120 may include a robot main controller (screen) 1201, etc. The robot terminal 120 and the robot platform terminal 130 can communicate via voice. For example, the user's voice information collected by the robot terminal 120 can be accessed through a multimodal gateway 1301, and then the user's voice information can be processed using Automatic Speech Recognition (ASR), Voice Activity Detection (VAD), and Text-to-Speech (TTS). Then, the voice service central control unit 1302 can call large models in the model library 1303, or domain natural language understanding (domain NLU), or a knowledge base for intent recognition. The voice service central control unit 1302 can also invoke Remote Procedure Call (RPC) to provide skill services 1304 through voice skills, such as IoT, music, radio, weather, etc. IoT control RPC is performed with the IoT platform 1305 through skill communication. The IoT platform 1305 can manage devices, products, object models, rule engines, scene linkage, over-the-air (OTA) technology, alarm management, and log management through IoT interfaces. It can also connect IoT devices on the IoT device side 140. These IoT devices can include the robot itself, robot peripheral sensors (such as wearables and environmental monitoring devices), and third-party cloud devices. Both the robot side 120 and the robot platform side 130 can interact with other services such as audio / video services and communication services.
[0047] It is important to note that Figure 1B The examples shown are merely examples of system architectures that can be applied to the embodiments of this disclosure, in order to help those skilled in the art understand the technical content of this disclosure, but do not mean that the embodiments of this disclosure cannot be used in other devices, systems, environments or scenarios.
[0048] Figure 2 A flowchart illustrating a device operation method according to an embodiment of the present disclosure is shown schematically.
[0049] like Figure 2 As shown, the method includes operations S210~S260, which can be performed by a robot.
[0050] During operation S210, the first voice information of the user collected is sent to the robot platform so that the robot platform can access the first voice information based on the multimodal gateway and call the target large model from the model library through the voice service central control to perform intent recognition on the first voice information to obtain the first recognition result.
[0051] According to embodiments of this disclosure, the robot platform may include a multimodal gateway, a voice service control center, and a model library. The model library may include various large models.
[0052] According to embodiments of this disclosure, when a user needs to operate the device, they can send a first voice message to the robot, such as "I want to operate the device." The robot can collect the first voice message through a microphone or other device. The voice service control center can call any large model from the model library as the target large model, or it can call a large model related to operating the device as the target large model; no specific limitation is made here. The first recognition result can be used to characterize whether the user needs to operate the device. For the first voice message such as "I want to operate the device," the first recognition result can characterize that the user needs to operate the device.
[0053] In operation S220, if it is determined that the first recognition result indicates that the user needs to operate the device, the second voice information is sent to the user.
[0054] According to embodiments of this disclosure, the second voice information is used to represent a first question inquiring of the device identifier from the user. The device identifier may be a device serial number, etc., and the second voice information may be something like, "Do you know the serial number of the device you want to operate?"
[0055] When operating the S230, third-party voice information of the user is collected.
[0056] According to embodiments of this disclosure, after the user receives the second voice information from the robot, a third voice information can be sent to the robot in voice form based on the device identifier of the device the user needs to operate. The robot can collect the user's third voice information through a microphone.
[0057] The third voice information is used to represent the first answer that matches the first question. For example, for the second voice information, such as "Do you know the serial number of the device you want to operate?", the third voice information could be "The serial number is XXXXXX".
[0058] During operation S240, the third voice information is sent to the robot platform so that the robot platform can access the third voice information based on the multimodal gateway, and call the target large model from the model library through the voice service central control to perform intent recognition on the third voice information to obtain the second recognition result.
[0059] According to embodiments of this disclosure, the second identification result is used to characterize whether the user has provided a device identifier. For third voice information such as "serial number XXXXXX", the second identification result can characterize whether the user has provided a device identifier.
[0060] In operation S250, if it is determined that the second identification result indicates that the user has provided a device identifier, the device type of the device is determined based on the device identifier.
[0061] According to embodiments of this disclosure, the robot can determine the device matching the device identifier provided by the user. The device type is then determined based on whether it can interact with cloud services. For example, device types can be divided into two categories: non-directly connected terminal devices and third-party cloud devices.
[0062] In operation S260, operation instructions are generated based on the device identifier, device type, and operation type to complete the operation of the device.
[0063] According to embodiments of this disclosure, the operation type can be determined based on the operation performed on the device during user interaction. Then, based on the operation type, device type, and device identifier, an operation command can be generated and sent to the corresponding device for execution, thus completing the operation on the device.
[0064] According to embodiments of this disclosure, a robot-centric approach is used to interact with the user via voice. The robot platform's multimodal gateway, voice service control center, and model library are used to recognize the user's voice intent, accurately determine the device the user needs to operate, and generate operation commands to complete the device operation. Since the operation commands are generated based on the device identifier, device type, and operation type determined through voice interaction, there is no need for the user to install or operate any application corresponding to the device. This simplifies the device operation process and enhances the user experience. Furthermore, since the entire interaction between the user and the robot is conducted via voice, user operation is simplified, especially for users unfamiliar with the operation process. This provides a more intuitive and user-friendly device operation method. It at least partially solves the problem in related technologies where the operation process of IoT devices is mostly based on applications or touch-based interactions, resulting in long and cumbersome processes.
[0065] The following is for reference. Figures 3A-3E In conjunction with specific embodiments, Figure 2 The method shown will be further explained.
[0066] In realizing this disclosed concept, the inventors also discovered that the method of generating operation instructions can differ depending on the type of device and the type of operation. If operation instructions are generated in a similar manner for all, there will be unnecessary waste of resources and low operational efficiency.
[0067] Based on this, according to embodiments of this disclosure, for example... Figure 2The illustrated operation S260 generates operation instructions based on the device identifier, device type, and operation type to complete the operation on the device. This may include the following operations: determining whether to invoke the device provider's application based on the device type and operation type; if it is determined that invoking the application is unnecessary, generating operation instructions based on the device identifier and operation type to complete the operation on the device; if it is determined that invoking the application is necessary, invoking the application based on the device identifier so that the application can control the device to turn on Bluetooth; connecting via Bluetooth and sending a connection result to the application so that the application, upon determining that the connection result indicates a successful connection, returns operation information matching the device; and generating operation instructions based on the operation information to complete the operation on the device.
[0068] According to embodiments of this disclosure, it can be determined whether a device can interact with cloud services based on the device type. It can also be determined whether operations on the device require the assistance of the device provider's application based on the operation type. If it is determined that the device can interact with cloud services and operations on the device require the assistance of the device provider's application, it is determined that the device provider's application needs to be invoked. If it is determined that the device cannot interact with cloud services, or that the device can interact with cloud services but operations on the device do not require the assistance of the device provider's application, it is determined that the device provider's application does not need to be invoked.
[0069] According to embodiments of this disclosure, operational information may include, but is not limited to, Service Set Identifier (SSID), password, security protocol, and other information.
[0070] According to embodiments of this disclosure, by determining whether to invoke the device provider's application, it is possible to save resources and improve operational efficiency without invoking the device provider's application.
[0071] According to embodiments of this disclosure, determining whether to invoke the device provider's application based on the device type and operation type may include the following operations: if the device type is determined to be a non-directly connected peripheral device, it is determined that no application needs to be invoked; or if the device type is determined to be a third-party cloud device and the operation type is a network configuration operation, it is determined that the application needs to be invoked; or if the device type is determined to be a third-party cloud device and the operation type is a binding operation or an unbinding operation, it is determined that no application needs to be invoked.
[0072] According to embodiments of this disclosure, third-party cloud devices can interact with cloud services. Non-directly connected peripheral devices cannot interact with cloud services. Network configuration operations for third-party cloud devices require calling an application. The application can guide the network configuration process. Binding and unbinding operations for third-party cloud devices do not require calling an application.
[0073] According to embodiments of this disclosure, by determining whether an application needs to be invoked based on different device types and operation types, operation instructions can be generated in a personalized manner for different device types and operation types.
[0074] Figure 3A The illustration shows an interactive diagram of a network configuration operation for a non-directly connected peripheral device according to an embodiment of the present disclosure; Figure 3B The illustration shows an interactive diagram of a network configuration operation for a third-party cloud device according to an embodiment of the present disclosure.
[0075] According to embodiments of this disclosure, when it is determined that no application invocation is required, an operation instruction is generated based on the device identifier and operation type to complete the operation on the device. This operation may include: when it is determined that no application invocation is required and the operation type is a network configuration operation, sending a network configuration operation request matching the device identifier to the robot platform; receiving network configuration operation information matching the network configuration operation request sent by the robot platform; generating an operation instruction for connecting the device based on the network configuration operation information; and sending the operation instruction to the device to complete the network configuration operation.
[0076] According to embodiments of this disclosure, network configuration information may include, but is not limited to, SSID, network security protocol, wireless frequency band, address configuration, subnet mask, default gateway, etc.
[0077] According to embodiments of this disclosure, the robot platform may further include an Internet of Things (IoT) platform, which can store network configuration operation information for various networked devices. Operation instructions may include, but are not limited to, discovering devices and issuing network commands.
[0078] For network configuration operations involving non-directly connected peripheral devices, for example, it can be done as follows: Figure 3A As shown, user 301 sends a first voice message 3011 to robot 302. After receiving the first voice message 3011, robot 302 sends it to robot platform 303. Robot platform 303 accesses the first voice message 3011 via a multimodal gateway and, through the voice service central control, calls the target large model from the model library to perform intent recognition on the first voice message 3011, obtaining a first recognition result 3031, which is then returned to robot 302. Robot 302 then prompts user 301 via voice to enable network configuration for the device. User 301 enables network configuration for the non-directly connected peripheral device 304.
[0079] User 301 sends a voice message 3012 to robot 302 indicating that the device has started network configuration. Robot 302 sends the voice message 3012 to robot platform 303. Robot platform 303 accesses the voice message 3012 based on the multimodal gateway and calls the target large model from the model library through the voice service central control to perform intent recognition on the voice message 3012 indicating that the device has started network configuration, obtaining the network configuration recognition result 3032, and returns the network configuration recognition result 3032 to robot 302.
[0080] Robot 302 sends a second voice message 3021 to user 301, requesting the serial number of the network-configured device. User 301 sends a third voice message 3013 to robot 302, stating the corresponding serial number. Robot 302 sends the third voice message 3013 to robot platform 303. Robot platform 303 accesses the third voice message 3013 via a multimodal gateway and uses the voice service central control to call the target large model from the model library to perform intent recognition on the third voice message 3013, obtaining a second recognition result 3033. Robot 302 sends a network configuration operation request 3022 matching the serial number to robot platform 303. In response to the network configuration operation request 3022, robot platform 303 sends network configuration operation information 3034 matching the request 3022 to robot 302. Based on the network configuration operation information 3034, robot 302 obtains the Software Development Kit (SDK) and generates operation instructions 3023 for connecting the device, such as discovering the device and issuing network permissions. Robot 302 sends operation command 3023 to non-directly connected peripheral device 304 via local area network communication. Non-directly connected peripheral device 304 connects to the network and completes the network distribution operation.
[0081] For example, interactions with third-party cloud devices during network configuration operations can be performed as follows: Figure 3B As shown, user 301 sends a first voice message 3011 to robot 302. After receiving the first voice message 3011, robot 302 sends it to robot platform 303. Robot platform 303 accesses the first voice message 3011 through a multimodal gateway, and calls the target large model from the model library through the voice service central control to perform intent recognition on the first voice message 3011 to obtain a first recognition result 3031, and returns the first recognition result 3031 to robot 302.
[0082] Based on the device identifier of the third-party cloud device 306, robot 302 invokes the device provider's application 305 (such as a brand owner's mini-program) to begin network configuration. Application 305 guides the network configuration process. Application 305 controls the third-party cloud device 306 to turn on Bluetooth. Application 305 calls an interface to scan the Bluetooth list and connect to the Bluetooth of the third-party cloud device 306. Robot 302 scans the Bluetooth list and connects to the Bluetooth of the third-party cloud device 306. The third-party cloud device 306 returns a connection success message to robot 302. Robot 302 sends a connection success result to application 305. Upon confirming a successful connection, application 305 returns operation information 3051 matching the third-party cloud device 306 to robot 302. Robot 302 generates operation instructions 3023 based on operation information 3051 and sends operation instructions 3024 to the third-party cloud device 306 using Bluetooth Low Energy (BLE) technology. The third-party cloud device 306 executes the operation instructions to complete network connection. The third-party cloud device 306 can successfully report the network connection to the robot 302.
[0083] According to embodiments of this disclosure, for non-directly connected peripheral devices, the robot can operate the device through voice interaction with the user, using the robot platform. This eliminates the need to call up a device-compatible application, download or operate the device based on an application, simplifying the process, saving resources, and enhancing the user experience.
[0084] According to embodiments of this disclosure, for third-party cloud devices, the robot invokes the application without requiring the user to download the application, simplifying the process, saving resources, and enhancing the user experience.
[0085] Figure 3C The illustration schematically shows an interactive diagram of a binding operation for a non-directly connected peripheral device according to an embodiment of the present disclosure; Figure 3D The illustration shows an interactive diagram of a binding operation for a third-party cloud device according to an embodiment of the present disclosure.
[0086] According to embodiments of this disclosure, when it is determined that no application needs to be invoked, an operation instruction is generated based on the device identifier and operation type to complete the operation on the device. This operation may include: when it is determined that no application needs to be invoked and the operation type is a binding operation, for a non-directly connected peripheral device, such as... Figure 3CAs shown, robot 302 can send a network configuration query request 3024 to robot platform 303. Robot platform 303 authenticates the non-directly connected peripheral device 304. In response to the network configuration query request 3024, based on the authentication result of the non-directly connected peripheral device 304, a binding operation is performed. Robot platform 303 sends the results of successful network configuration and successful binding to robot 302. Robot 302 then sends the feedback of successful network configuration and successful binding to user 301 via voice.
[0087] The authentication of the non-directly connected peripheral device 304 by the robot platform 303 may include the following operations: the robot platform 303 performs connection authentication on the non-directly connected peripheral device 304, and if the authentication is successful, the non-directly connected peripheral device 304 goes online.
[0088] For third-party cloud devices, such as Figure 3D As shown, robot 302 can successfully report network connection to application 305. Application 305 sends binding command 3052 to robot platform 303. After executing binding command 3052, robot platform 303 sends a successful binding result to robot 302. Robot 302 then feeds back the successful binding result to user 301 via voice.
[0089] Figure 3E The illustration shows an interactive diagram of the unbinding operation according to an embodiment of the present disclosure.
[0090] According to embodiments of this disclosure, the operation type may include an unbinding operation.
[0091] For example Figure 2 Operation S260, as shown, generates an operation command based on the device identifier, device type, and operation type to complete the operation on the device. This may include the following operations: sending a fourth voice message to the user; collecting the user's fifth voice message; sending the fifth voice message to the robot platform, so that the robot platform, through the voice service control center, calls the target large model from the model library to perform intent recognition on the fifth voice message to obtain a third recognition result; if the third recognition result indicates that the user confirms the device identifier is correct, generating an unbinding operation command corresponding to the device identifier and sending the unbinding operation command to the robot platform; receiving a successful unbinding result sent by the robot platform after performing the unbinding operation and providing voice feedback to the user on the successful unbinding.
[0092] According to embodiments of this disclosure, the fourth voice information is used to characterize a second question asking the user whether the device identifier is correct.
[0093] According to embodiments of this disclosure, the fifth voice information is used to characterize a second answer that matches the second question.
[0094] For devices that are not directly connected peripherals or third-party cloud devices, the unbinding process can be as follows: Figure 3E As shown, user 301 sends a first voice message 3011 to robot 302. This first voice message 3011 is related to the unbinding operation. After receiving the first voice message 3011, robot 302 sends it to robot platform 303. Robot platform 303 accesses the first voice message 3011 through a multimodal gateway and calls the target large model from the model library through the voice service central control to perform intent recognition on the first voice message 3011, obtaining a first recognition result 3031, and returning the first recognition result 3031 to robot 302. This first recognition result 3031 indicates that the user needs to unbind the device.
[0095] Robot 302 sends a second voice message 3021 to user 301, asking user 301 about the device that needs to be unbound. User 301 sends a third voice message 3013 to robot 302, stating the corresponding device to be unbound. Robot 302 sends the third voice message 3013 to robot platform 303. Robot platform 303 accesses the third voice message 3013 through a multimodal gateway, and through the voice service central control, calls the target large model from the model library to perform intent recognition on the third voice message 3013 to obtain a second recognition result 3033, which is then returned to robot 302.
[0096] Robot 302 sends a fourth voice message 3025 to the user. The user sends a fifth voice message 3014. The fifth voice message 3014 is sent to the robot platform 303, so that the robot platform 303 can call the target large model from the model library through the voice service central control to perform intent recognition on the fifth voice message 3014 and obtain a third recognition result 3035. If the robot 302 determines that the third recognition result 3035 indicates that the user 301 has confirmed the device unbinding, it generates an unbinding operation command 3026 corresponding to the device identifier and sends the unbinding operation command 3026 to the robot platform 303. The robot platform 303 executes the unbinding operation and sends the unbinding success result to the robot 302. The robot 302 then feeds back the unbinding success to the user 301 in the form of voice.
[0097] According to embodiments of this disclosure, the unbinding operation for third-party cloud devices may further include the following steps: After unbinding the device, the robot platform 303 generates an instruction for unbinding the device in the application. After the application successfully unbinds the device, it returns a successful unbinding result to the robot platform 303. The robot platform 303 then sends the successful unbinding result to the robot 302, and the robot 302 then provides voice feedback to the user 301 regarding the successful unbinding.
[0098] According to embodiments of this disclosure, personalized binding and unbinding operations are provided for different devices, simplifying the process, saving resources, and enhancing the user experience.
[0099] Figure 4 A flowchart illustrating a method of operating a device according to another embodiment of the present disclosure is shown schematically.
[0100] like Figure 4 As shown, the method includes operations S410~S460, which can be executed by the robot platform.
[0101] When operating the S410, based on the multimodal gateway, the first voice information of the user sent by the robot is received.
[0102] When operating the S420, based on the voice service central control, the target large model is called from the model library to perform intent recognition on the first voice information to obtain the first recognition result.
[0103] When operating S430, the first recognition result is sent to the robot so that the robot can send the second voice information to the user.
[0104] According to embodiments of this disclosure, the second voice information is used to characterize a first question asked to the user regarding the device identifier.
[0105] When operating the S440, it receives third-party voice information sent by the robot.
[0106] According to embodiments of this disclosure, third voice information is used to characterize a first answer that matches a first question, and the third voice information is collected by the robot from the user.
[0107] When operating the S450, based on the voice service central control, the target large model is called from the model library to perform intent recognition on the third voice information to obtain the second recognition result.
[0108] In operation S460, the second identification result is sent to the robot so that the robot can generate operation instructions based on the device identifier, device type and operation type, if it determines that the second identification result indicates that the user has provided a device identifier.
[0109] According to embodiments of this disclosure, the device type is determined based on a device identifier.
[0110] It should be noted that the device operation method executed by the robot platform in this embodiment corresponds to the device operation method executed by the robot described above. For a detailed description of the device operation method executed by the robot platform, please refer to the device operation method executed by the robot section, which will not be repeated here.
[0111] Figure 5A block diagram of a device operating apparatus according to an embodiment of the present disclosure is shown schematically.
[0112] like Figure 5 As shown, the device operation device 500 includes a first sending module 510, a second sending module 520, a data acquisition module 530, a third sending module 540, a determination module 550, and a generation module 560.
[0113] The first sending module 510 is used to send the collected first voice information of the user to the robot platform, so that the robot platform can access the first voice information based on the multimodal gateway, and call the target large model from the model library through the voice service central control to perform intent recognition on the first voice information to obtain the first recognition result. The robot platform includes a multimodal gateway, a voice service central control and a model library.
[0114] The second sending module 520 is used to send second voice information to the user when it is determined that the first recognition result indicates that the user needs to operate the device. The second voice information is used to indicate that the user is asking a first question about the device identifier.
[0115] The acquisition module 530 is used to acquire the user's third voice information, wherein the third voice information is used to represent the first answer that matches the first question.
[0116] The third sending module 540 is used to send the third voice information to the robot platform, so that the robot platform can access the third voice information based on the multimodal gateway, and call the target large model from the model library through the voice service central control to perform intent recognition on the third voice information to obtain the second recognition result.
[0117] The determination module 550 is used to determine the device type of the device based on the device identifier when the second identification result indicates that the user has provided a device identifier.
[0118] The generation module 560 is used to generate operation instructions based on the device identifier, device type and operation type in order to complete the operation of the device.
[0119] Figure 6 A block diagram of a device operating apparatus according to another embodiment of the present disclosure is shown schematically.
[0120] like Figure 6 As shown, the device operation device 600 includes a first access module 610, a first invocation module 620, a fourth sending module 630, a second access module 640, a second invocation module 650, and a fifth sending module 660.
[0121] The first access module 610 is used to access the user's first voice information sent by the robot based on a multimodal gateway.
[0122] The first calling module 620 is used for central control based on voice service, calling the target large model from the model library to perform intent recognition on the first voice information to obtain the first recognition result.
[0123] The fourth sending module 630 is used to send the first recognition result to the robot so that the robot can send second voice information to the user, wherein the second voice information is used to represent a first question to the user regarding the device identifier.
[0124] The second access module 640 is used to receive third voice information sent by the robot, wherein the third voice information is used to represent the first answer that matches the first question, and the third voice information is collected by the robot from the user.
[0125] The second calling module 650 is used for central control based on voice service, calling the target large model from the model library to perform intent recognition on the third voice information to obtain the second recognition result.
[0126] The fifth sending module 660 is used to send the second identification result to the robot so that, when the robot determines that the second identification result indicates that the user has provided a device identifier, it can generate an operation instruction based on the device identifier, device type and operation type, wherein the device type is determined according to the device identifier.
[0127] Any one or more of the modules, submodules, units, and subunits according to embodiments of the present disclosure, or at least part of the functions of any one or more of them, can be implemented in one module. Any one or more of the modules, submodules, units, and subunits according to embodiments of the present disclosure can be implemented by dividing them into multiple modules. Any one or more of the modules, submodules, units, and subunits according to embodiments of the present disclosure can be at least partially implemented as hardware circuitry, such as a Field-Programmable Gate Array (FPGA), a Programmable Logic Array (PLA), a System-on-Chip, a System-on-a-Substrate, a System-on-Package, an Application-Specific Integrated Circuit (ASIC), or implemented in hardware or firmware by any other reasonable means of integrating or packaging circuitry, or implemented in software, hardware, or firmware, or in any suitable combination of any of these three implementation methods. Alternatively, one or more of the modules, submodules, units, and subunits according to embodiments of the present disclosure can be at least partially implemented as computer program modules, which, when run, can perform corresponding functions.
[0128] For example, any and more of the first transmitting module 510, the second transmitting module 520, the acquisition module 530, the third transmitting module 540, the determination module 550, and the generation module 560 can be combined into one module / unit / subunit, or any one of these modules / units / subunits can be split into multiple modules / units / subunits. Alternatively, at least some of the functions of one or more of these modules / units / subunits can be combined with at least some of the functions of other modules / units / subunits and implemented in one module / unit / subunit. According to embodiments of this disclosure, at least one of the first transmitting module 510, the second transmitting module 520, the acquisition module 530, the third transmitting module 540, the determination module 550, and the generation module 560 can be at least partially implemented as hardware circuits, such as field-programmable gate arrays (FPGAs), programmable logic arrays (PLAs), systems-on-a-chip, systems-on-a-substrate, systems-on-package, application-specific integrated circuits (ASICs), or any other reasonable means of integrating or packaging circuits, or implemented in software, hardware, or firmware, or in any suitable combination of any of these three implementation methods. Alternatively, at least one of the first transmitting module 510, the second transmitting module 520, the acquisition module 530, the third transmitting module 540, the determination module 550, and the generation module 560 can be at least partially implemented as computer program modules, which can perform corresponding functions when the computer program module is run.
[0129] For example, any and more of the first access module 610, first invocation module 620, fourth sending module 630, second access module 640, second invocation module 650, and fifth sending module 660 can be combined into one module / unit / subunit, or any one of these modules / units / subunits can be split into multiple modules / units / subunits. Alternatively, at least some of the functions of one or more of these modules / units / subunits can be combined with at least some of the functions of other modules / units / subunits and implemented in one module / unit / subunit. According to embodiments of this disclosure, at least one of the first access module 610, the first calling module 620, the fourth sending module 630, the second access module 640, the second calling module 650, and the fifth sending module 660 can be at least partially implemented as hardware circuits, such as field-programmable gate arrays (FPGAs), programmable logic arrays (PLAs), systems-on-a-chip, systems-on-a-substrate, systems-on-package, application-specific integrated circuits (ASICs), or any other reasonable means of integrating or packaging circuits, or implemented in hardware or firmware, or in any one of software, hardware, and firmware implementations, or in a suitable combination of any of these. Alternatively, at least one of the first access module 610, the first calling module 620, the fourth sending module 630, the second access module 640, the second calling module 650, and the fifth sending module 660 can be at least partially implemented as a computer program module, which, when run, can perform corresponding functions.
[0130] It should be noted that the device operation device part in the embodiments of this disclosure corresponds to the device operation method part in the embodiments of this disclosure. For a detailed description of the device operation device part, please refer to the device operation method part, which will not be repeated here.
[0131] Figure 7 A block diagram schematically illustrates an electronic device suitable for implementing a device operation method according to an embodiment of the present disclosure. Figure 7 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0132] like Figure 7As shown, an electronic device 700 according to an embodiment of the present disclosure includes a processor 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage portion 708 into a random access memory (RAM) 703. The processor 701 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 701 may also include onboard memory for caching purposes. The processor 701 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.
[0133] RAM 703 stores various programs and data required for the operation of electronic device 700. Processor 701, ROM 702, and RAM 703 are interconnected via bus 704. Processor 701 performs various operations of the method flow according to embodiments of the present disclosure by executing programs in ROM 702 and / or RAM 703. It should be noted that programs may also be stored in one or more memories other than ROM 702 and RAM 703. Processor 701 may also perform various operations of the method flow according to embodiments of the present disclosure by executing programs stored in one or more memories.
[0134] According to embodiments of this disclosure, the electronic device 700 may further include an input / output (I / O) interface 705, which is also connected to a bus 704. The electronic device 700 may also include one or more of the following components connected to the input / output (I / O) interface 705: an input section 706 including a keyboard, mouse, etc.; an output section 707 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a LAN card, modem, etc. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the input / output (I / O) interface 705 as needed. A removable medium 711, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 710 as needed so that computer programs read from it can be installed into the storage section 708 as needed.
[0135] According to embodiments of this disclosure, the method flow according to embodiments of this disclosure can be implemented as a computer software program. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable storage medium, the computer program containing program code for performing the methods shown in the flowchart. In such embodiments, the computer program can be downloaded and installed from a network via communication section 709, and / or installed from removable medium 711. When the computer program is executed by processor 701, it performs the functions defined in the system of embodiments of this disclosure. According to embodiments of this disclosure, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.
[0136] This disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs that, when executed, implement the method according to the embodiments of this disclosure.
[0137] According to embodiments of this disclosure, the computer-readable storage medium can be a non-volatile computer-readable storage medium. Examples include, but are not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0138] For example, according to embodiments of this disclosure, a computer-readable storage medium may include the ROM 702 and / or RAM 703 described above and / or one or more memories other than ROM 702 and RAM 703.
[0139] Embodiments of this disclosure also include a computer program product comprising a computer program containing program code for performing the methods provided in the embodiments of this disclosure. When the computer program product is run on an electronic device, the program code is used to enable the electronic device to implement the methods provided in the embodiments of this disclosure.
[0140] When the computer program is executed by the processor 701, it performs the functions defined in the system / apparatus of this disclosure embodiments. According to embodiments of this disclosure, the systems, apparatuses, modules, units, etc., described above can be implemented by computer program modules.
[0141] In one embodiment, the computer program may rely on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of signals over a network medium, and may be downloaded and installed via the communication section 709, and / or installed from a removable medium 711. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.
[0142] According to embodiments of this disclosure, program code for executing the computer programs provided in embodiments of this disclosure can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, languages such as Java, C++, Python, "C", or similar programming languages. The program code can execute entirely on a user's computing device, partially on a user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0143] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions. Those skilled in the art will understand that the features described in the various embodiments of the present disclosure can be combined and / or combined in various ways, even if such combinations are not explicitly described in the present disclosure. In particular, the features described in the various embodiments of this disclosure may be combined and / or combined in various ways without departing from the spirit and teachings of this disclosure. All such combinations and / or combinations fall within the scope of this disclosure.
[0144] The embodiments of this disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of this disclosure. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. Various substitutions and modifications can be made by those skilled in the art without departing from the scope of this disclosure, and all such substitutions and modifications should fall within the scope of this disclosure.
Claims
1. A method for operating a device, comprising: The first voice information collected from the user is sent to the robot platform, so that the robot platform can access the first voice information based on the multimodal gateway, and call the target large model from the model library through the voice service central control to perform intent recognition on the first voice information to obtain the first recognition result. The robot platform includes the multimodal gateway, the voice service central control and the model library. If the first identification result indicates that the user needs to operate the device, a second voice message is sent to the user, wherein the second voice message is used to indicate that a first question is asked to the user regarding the device identifier; Collect the user's third voice information, wherein the third voice information is used to represent the first answer that matches the first question; The third voice information is sent to the robot platform so that the robot platform can access the third voice information based on the multimodal gateway, and call the target large model from the model library through the voice service central control to perform intent recognition on the third voice information to obtain a second recognition result; If the second identification result indicates that the user has provided the device identifier, then based on the device identifier, the device type of the device is determined; and Based on the device identifier, the device type, and the operation type, an operation instruction is generated to complete the operation of the device.
2. The method according to claim 1, wherein, The step of generating operation instructions based on the device identifier, the device type, and the operation type to complete the operation of the device includes: Based on the device type and the operation type, determine whether to invoke the device provider's application; If it is determined that the application does not need to be invoked, an operation instruction is generated based on the device identifier and the operation type to complete the operation on the device; If it is determined that the application needs to be invoked, the application is invoked based on the device identifier so that the application controls the device to turn on Bluetooth; Connect via Bluetooth and send a connection result to the application, so that the application, upon determining that the connection result indicates a successful connection, returns operation information matching the device; and Based on the operation information, operation instructions are generated to complete the operation of the device.
3. The method according to claim 2, wherein, The step of determining whether to invoke the device provider's application based on the device type and the operation type includes: If it is determined that the device type is a non-directly connected peripheral device, it is determined that the application does not need to be invoked; or If the device type is determined to be a third-party cloud device and the operation type is a network configuration operation, then it is determined that the application needs to be invoked; or If the device type is determined to be a third-party cloud device and the operation type is a binding operation or an unbinding operation, it is determined that there is no need to call the application.
4. The method according to claim 2, wherein, The step of generating operation instructions based on the device identifier and the operation type to complete the operation on the device when it is determined that the application does not need to be invoked includes: If it is determined that there is no need to invoke the application and the operation type is a network configuration operation, a network configuration operation request matching the device identifier is sent to the robot platform; Receive network distribution operation information sent by the robot platform that matches the network distribution operation request; Based on the distribution network operation information, generate the operation instructions for connecting the device; and The operation command is sent to the device to complete the distribution network operation.
5. The method according to claim 3, wherein, The step of generating operation instructions based on the device identifier and the operation type to complete the operation on the device when it is determined that the application does not need to be invoked includes: If it is determined that no application needs to be invoked, and the operation type is a binding operation, For the device type described as a non-directly connected peripheral device, perform the following operations: Send a network configuration query request to the robot platform, so that the robot platform responds to the network configuration query request and performs the binding operation based on the authentication result of the non-directly connected peripheral device; and Upon confirming that the network configuration and binding success results have been received from the robot platform, the network configuration and binding success results are fed back to the user in the form of voice. For the device type described above, which is a third-party cloud device, perform the following operations: The result of the completed network configuration operation is sent to the application, so that the application can send a binding instruction to the robot platform; and The system receives a successful binding result sent by the robot platform after executing the binding instruction, and then feeds back the successful binding result to the user in the form of voice.
6. The method according to claim 1, wherein the operation type includes an unbinding operation; in, The step of generating operation instructions based on the device identifier, device type, and operation type to complete the operation of the device includes: Send a fourth voice message to the user, wherein the fourth voice message is used to represent a second question to the user to confirm whether the device identifier is correct; Collect the user's fifth voice information, wherein the fifth voice information is used to represent the second answer that matches the second question; The fifth voice information is sent to the robot platform, so that the robot platform can call the target large model from the model library through the voice service center to perform intent recognition on the fifth voice information and obtain the third recognition result; If the third identification result indicates that the user has confirmed the device identifier is correct, an unbinding operation command corresponding to the device identifier is generated, and the unbinding operation command is sent to the robot platform; and The system receives a successful unbinding result sent by the robot platform after performing the unbinding operation, and then sends the successful unbinding result back to the user via voice.
7. A method for operating a device, comprising: Based on a multimodal gateway, the system accesses the user's first voice information sent by the robot; Based on the voice service central control, the target large model is called from the model library to perform intent recognition on the first voice information to obtain the first recognition result. The first recognition result is sent to the robot so that the robot can send second voice information to the user, wherein the second voice information is used to represent a first question to the user regarding the device identifier; Receive third voice information sent by the robot, wherein the third voice information is used to represent a first answer that matches the first question, and the third voice information is collected by the robot from the user; Based on the aforementioned voice service central control, the target large model is called from the model library to perform intent recognition on the third voice information to obtain a second recognition result; and The second identification result is sent to the robot so that, if the robot determines that the second identification result indicates that the user has provided the device identifier, it generates an operation instruction based on the device identifier, device type, and operation type, wherein the device type is determined according to the device identifier.
8. A device operating apparatus, comprising: The first sending module is used to send the first voice information collected from the user to the robot platform, so that the robot platform can access the first voice information based on the multimodal gateway, and call the target large model from the model library through the voice service central control to perform intent recognition on the first voice information to obtain the first recognition result. The robot platform includes the multimodal gateway, the voice service central control and the model library. The second sending module is used to send second voice information to the user when it is determined that the first identification result indicates that the user needs to operate the device, wherein the second voice information is used to indicate that the user is asking the user a first question about the device identifier; The acquisition module is used to acquire the user's third voice information, wherein the third voice information is used to represent the first answer that matches the first question; The third sending module is used to send the third voice information to the robot platform, so that the robot platform can access the third voice information based on the multimodal gateway, and call the target large model from the model library through the voice service central control to perform intent recognition on the third voice information to obtain a second recognition result; The determining module is configured to, when determining that the second identification result indicates that the user has provided the device identifier, determine the device type of the device based on the device identifier; and The generation module is used to generate operation instructions based on the device identifier, the device type, and the operation type, so as to complete the operation of the device.
9. A device operating apparatus, comprising: The first access module is used to access the user's first voice information sent by the robot based on a multimodal gateway; The first calling module is used to call the target large model from the model library based on the voice service central control to perform intent recognition on the first voice information and obtain the first recognition result; The fourth sending module is used to send the first recognition result to the robot so that the robot can send second voice information to the user, wherein the second voice information is used to represent a first question to the user regarding the device identifier; The second access module is used to receive third voice information sent by the robot, wherein the third voice information is used to represent a first answer that matches the first question, and the third voice information is collected by the robot from the user; The second invocation module is used to, based on the voice service central control, invoke the target large model from the model library to perform intent recognition on the third voice information to obtain a second recognition result; and The fifth sending module is used to send the second identification result to the robot, so that when the robot determines that the second identification result indicates that the user has provided the device identifier, it generates an operation instruction based on the device identifier, device type and operation type, wherein the device type is determined according to the device identifier.
10. An electronic device, comprising: One or more processors; Memory, used to store one or more programs. Wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the method of any one of claims 1 to 7.
11. A computer-readable storage medium having stored thereon executable instructions that, when executed by a processor, cause the processor to perform the method of any one of claims 1 to 7.
12. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-7.