Interaction method, device, terminal, system and storage medium

By obtaining user instructions and information to be identified by the terminal, and using a predetermined model to determine and share target operation instructions, the complex user operations in information sharing are solved, and the user experience of smart driving and smart home scenarios is improved.

CN120343039APending Publication Date: 2025-07-18XIAOMI EV TECH CO LTD +2
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510474886.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the prior art, user operations are complex during information sharing, resulting in a decrease in user experience.

Method used

By acquiring the first user instructions and the information to be identified in the first terminal, the target operation instructions are determined using a predetermined model, and a network connection is established through the device identification to share with the second terminal for execution, simplifying user input and improving operational convenience.

Benefits of technology

The instruction sharing between the first terminal and the second terminal is realized, which simplifies user operations and improves user experience, and is particularly suitable for smart driving and smart home scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120343039A_ABST
    Figure CN120343039A_ABST
Patent Text Reader

Abstract

The invention provides an interaction method and device, a terminal, a system and a storage medium, and relates to the technical field of information processing.The method comprises the steps that a first user instruction is obtained, and the first user instruction comprises first operation information; acquiring to-be-identified information of a first application in a first terminal to determine a target operation instruction associated with the to-be-identified information and the first operation information; the target operation instruction is used for triggering a second application of the second terminal to execute a target task, and the first terminal and the second terminal establish network connection through a device identifier. The to-be-identified information is collected through the first terminal, the target operation instruction is confirmed, and the target operation instruction is executed at the second terminal, so that instruction sharing between the first terminal and the second terminal is realized, user experience in an instruction sharing scene is optimized, and the method is suitable for instruction sharing scenes such as an intelligent driving scene and an intelligent home scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of information processing technologies, and in particular, to an interaction method, apparatus, terminal, system, and storage medium. Background Art

[0002] Currently, information sharing during interaction is widely used in fields such as driving and smart home. However, in related technologies, information sharing through interaction has the problem of relatively complex user operations, which reduces the user experience. Summary of the Invention

[0003] This application aims to solve at least one of the technical problems in related technologies to some extent.

[0004] To this end, this application proposes a method, apparatus, terminal, system, and storage medium.

[0005] The first aspect embodiment of this application proposes an interaction method, including:

[0006] Obtain a first user instruction, where the first user instruction includes first operation information;

[0007] Obtain the information to be recognized of a first application in a first terminal to determine a target operation instruction associated with the information to be recognized and the first operation information; the target operation instruction is used to trigger a second application in a second terminal to execute a target task, where the first terminal and the second terminal establish a network connection through a device identifier.

[0008] Optionally, it further includes:

[0009] Based on the information to be recognized and the first operation information, determine a target operation instruction associated with the information to be recognized and the first operation information.

[0010] Optionally, based on the information to be recognized and the first operation information, determining a target operation instruction associated with the information to be recognized and the first operation information includes:

[0011] Use the first operation information and the information to be recognized as input parameters and input them into a predetermined model to obtain a matching result; where the matching result is related to the target operation instruction.

[0012] Optionally, it further includes:

[0013] Receive at least one candidate address information, where the candidate address information is related to the information to be recognized and the first operation information;

[0014] Receive a second user instruction, and obtain the target operation instruction based on at least one of the candidate address information.

[0015] Optionally, receive a second user instruction, and obtain a target operation instruction based on at least one of the candidate address information, including at least one of the following methods:

[0016] Receive the user's voice instruction, select a target address from at least one of the candidate addresses, and obtain the target operation instruction, where the target operation instruction includes the target address;

[0017] Receive the operation instruction input by the user through a click operation or a selection operation, determine the target address from at least one of the candidate addresses, and obtain the target operation instruction, where the target operation instruction includes the target address.

[0018] Optionally, it further includes:

[0019] Receive at least one piece of candidate call object information, where the candidate call object information is related to the information to be recognized and the first operation information;

[0020] Receive a second user instruction, and obtain a target operation instruction based on at least one of the candidate call object information.

[0021] Optionally, the receiving the second user instruction and obtaining the target operation instruction based on at least one of the candidate call object information includes at least one of the following methods:

[0022] Receive the user's voice instruction, select a target call object from at least one of the candidate call objects, and obtain the target operation instruction, where the target operation instruction includes the target call object;

[0023] Receive the operation instruction input by the user through a click operation or a selection operation, determine the target call object from at least one of the candidate call objects, and obtain the target operation instruction, where the target operation instruction includes the target call object.

[0024] Optionally, the predetermined model is a multi-modal large language model trained with training data including the information to be recognized and operation information.

[0025] Optionally, the method further includes:

[0026] Obtain a first wake-up instruction, where the first wake-up instruction is triggered by an operation on the first terminal;

[0027] In response to the first wake-up instruction, collect the user's input information and obtain the first user instruction.

[0028] Optionally, the obtaining the information to be recognized of the first application in the first terminal includes:

[0029] Obtain the image data or text data in the first terminal as the information to be recognized.

[0030] Optionally, obtaining the image data or text data in the first terminal as the information to be recognized includes at least one of the following methods:

[0031] Capturing an image displayed on the first terminal to obtain the information to be recognized;

[0032] Obtaining an image of a specified area designated by the user in the first terminal as the information to be recognized;

[0033] Obtaining an image captured by the first terminal as the information to be recognized.

[0034] Optionally, the method further includes:

[0035] Receiving result information sent by a second terminal, where the result information is related to the execution progress or execution result of the target task.

[0036] An embodiment of the second aspect of the present application provides an interaction method, including:

[0037] Receiving a target operation instruction, where the target operation instruction is used to instruct the second terminal to execute a target task; the target operation instruction is related to the first operation information and the information to be recognized of the first terminal, and the first terminal and the second terminal establish a network connection through a device identifier;

[0038] Triggering the second application to execute the target task according to the target operation instruction.

[0039] Optionally, the method further includes:

[0040] Receiving at least one candidate address information, where the candidate address information is related to the information to be recognized and the first operation information;

[0041] Receiving a second user instruction, and confirming at least one of the candidate address information based on the second user instruction to obtain the target operation instruction.

[0042] Optionally, receiving a second user instruction, and confirming at least one of the candidate address information based on the second user instruction to obtain the target operation instruction includes at least one of the following methods:

[0043] Receiving a voice instruction of the user, selecting a target address from at least one of the candidate addresses to obtain the target operation instruction, where the target operation instruction includes the target address;

[0044] Receiving an operation instruction input by the user through a click operation or a selection operation, determining a target address from at least one of the candidate addresses to obtain the target operation instruction, where the target operation instruction includes the target address.

[0045] Optionally, the method further includes:

[0046] Receiving a target address sent by a first terminal to obtain the target operation instruction, where the target operation instruction includes the target address.

[0047] Optionally, the method further includes:

[0048] Receiving at least one piece of candidate call object information, where the candidate call object information is related to the information to be recognized and the first operation information;

[0049] Receiving a second user instruction, and based on the second user instruction, confirming at least one piece of the candidate call object information to obtain a target operation instruction.

[0050] Optionally, the receiving the second user instruction and, based on the second user instruction, confirming at least one piece of the candidate call object information to obtain a target operation instruction includes at least one of the following methods:

[0051] Receiving a voice instruction of a user, selecting a target call object from at least one of the candidate call objects to obtain the target operation instruction, where the target operation instruction includes the target call object;

[0052] Receiving an operation instruction input by the user through a click operation or a selection operation, determining a target call object from at least one of the candidate call objects to obtain a target operation instruction, where the target operation instruction includes the target call object.

[0053] Optionally, the method further includes:

[0054] Receiving a target call object sent by a first terminal to obtain the target operation instruction, where the target operation instruction includes the target call object.

[0055] Optionally, the method further includes:

[0056] Generating result information, where the result information is related to the execution progress or execution result of the target task; or,

[0057] Generating result information and sending the result information to the first terminal, where the result information is related to the execution progress or execution result of the target task.

[0058] An embodiment of the third aspect of the present application provides an interaction method, including:

[0059] Obtaining a first user instruction, where the first user instruction includes first operation information;

[0060] Obtaining information to be recognized of a first application in a first terminal;

[0061] Determine a target operation instruction based on the information to be recognized and the first operation information;

[0062] Trigger the second application of the second terminal to execute a target task based on the target operation instruction, where the first terminal and the second terminal establish a network connection through a device identifier.

[0063] Optionally, determining a target operation instruction based on the information to be recognized and the first operation information includes:

[0064] Use the first operation information and the information to be recognized as input parameters and input them into a predetermined model to obtain a matching result; wherein the matching result is related to the target operation instruction.

[0065] Optionally, it further includes:

[0066] Output at least one candidate address information through the predetermined model, where the candidate address information is related to the information to be recognized and the first operation information;

[0067] Receive a second user instruction, and confirm at least one of the candidate address information based on the second user instruction to obtain a target operation instruction.

[0068] Optionally, receiving a second user instruction and confirming at least one of the candidate address information based on the second user instruction to obtain a target operation instruction includes at least one of the following methods:

[0069] Receive a voice instruction from the user, select a target address from at least one of the candidate addresses, and obtain the target operation instruction, where the target operation instruction includes the target address;

[0070] Receive an operation instruction input by the user through a click operation or a selection operation, determine a target address from at least one of the candidate addresses, and obtain the target operation instruction, where the target operation instruction includes the target address.

[0071] Optionally, it further includes:

[0072] Output at least one candidate call object information through the predetermined model, where the candidate call object information is related to the information to be recognized and the first operation information;

[0073] Receive a second user instruction, and confirm at least one of the candidate call object information based on the second user instruction to obtain a target operation instruction.

[0074] Optionally, receiving the second user instruction and confirming at least one of the candidate call object information based on the second user instruction to obtain a target operation instruction includes at least one of the following methods:

[0075] Receiving a voice instruction from the user, selecting a target call object from at least one of the candidate call objects, and obtaining the target operation instruction, where the target operation instruction includes the target call object;

[0076] Receiving an operation instruction input by the user through a click operation or a selection operation, determining a target call object from at least one of the candidate call objects, and obtaining the target operation instruction, where the target operation instruction includes the target call object.

[0077] Optionally, the predetermined model is a multimodal large language model trained with training data including information to be recognized and operation information.

[0078] Optionally, the method further includes:

[0079] Obtaining a first wake-up instruction, where the first wake-up instruction is triggered by an operation on the first terminal;

[0080] In response to the first wake-up instruction, collecting input information of the user and obtaining the first user instruction.

[0081] Optionally, obtaining the information to be recognized of the first application in the first terminal includes:

[0082] Obtaining image data or text data in the first terminal as the information to be recognized.

[0083] Optionally, obtaining the image data or text data in the first terminal as the information to be recognized includes at least one of the following methods:

[0084] Capturing an image displayed on the first terminal to obtain the information to be recognized;

[0085] Obtaining an area image specified by the user in the first terminal as the information to be recognized;

[0086] Obtaining an image captured by the first terminal as the information to be recognized.

[0087] An embodiment of one aspect of the present application provides an interaction device for executing the method in the foregoing first aspect, including:

[0088] A transceiver module for obtaining a first user instruction, where the first user instruction includes first operation information;

[0089] A processing module, configured to obtain the information to be recognized of a first application in a first terminal, so as to determine a target operation instruction associated with the information to be recognized and the first operation information; the target operation instruction is used to trigger a second application in the second terminal to execute a target task, where the first terminal and the second terminal establish a network connection through a device identifier.

[0090] Optionally, it further includes:

[0091] An instruction determination module, configured to determine a target operation instruction associated with the information to be recognized and the first operation information based on the information to be recognized and the first operation information.

[0092] Optionally, the instruction determination module includes:

[0093] A matching module, configured to use the first operation information and the information to be recognized as input parameters and input them into a predetermined model to obtain a matching result; the matching result is related to the target operation instruction.

[0094] An interaction device is proposed in an embodiment of one aspect of the present application. The device is used to execute the method in the second aspect described above, and includes:

[0095] A transceiver module, configured to receive a target operation instruction, where the target operation instruction is used to instruct the second terminal to execute a target task; the target operation instruction is related to the first operation information and the information to be recognized of the first terminal, and the first terminal and the second terminal establish a network connection through a device identifier;

[0096] An execution module, configured to trigger the second application to execute a target task according to the target operation instruction.

[0097] An interaction device is proposed in an embodiment of one aspect of the present application. The device is used to execute the method in the third aspect described above, and includes:

[0098] A transceiver module, configured to obtain a first user instruction, where the first user instruction includes first operation information; obtain the information to be recognized of a first application in the first terminal;

[0099] A processing module, configured to determine a target operation instruction based on the information to be recognized and the first operation information;

[0100] An execution module, configured to trigger the second application in the second terminal to execute a target task based on the target operation instruction, where the first terminal and the second terminal establish a network connection through a device identifier.

[0101] Optionally, the processing module includes:

[0102] A matching module, configured to use the first operation information and the information to be recognized as input parameters, and input them into a predetermined model to obtain a matching result; wherein, the matching result is related to the target operation instruction.

[0103] Another embodiment of this application provides a first terminal, including at least one processor, and the at least one processor is coupled to a memory; the at least one processor is configured to execute the method described in the foregoing first aspect.

[0104] Another embodiment of this application provides a second terminal, including at least one processor, and the at least one processor is coupled to a memory; the at least one processor is configured to execute the method described in the foregoing second aspect.

[0105] Another embodiment of this application provides a system, which is configured to implement the method described in the foregoing first aspect, or the second aspect, or the third aspect.

[0106] Another embodiment of this application provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the method described in the foregoing first aspect, second aspect or third aspect.

[0107] Another embodiment of this application provides a computer program product, and when the program is executed by a processor, it implements the method described in the foregoing first aspect, second aspect or third aspect.

[0108] The interaction method, device, electronic device, chip and storage medium provided by this application collect the information to be recognized and the first user instruction by the first terminal, confirm the target operation instruction, and share the target operation instruction to the second terminal through a network connection, and execute the target operation instruction on the second terminal, realizing the instruction sharing between the first terminal and the second terminal, and optimizing the user experience in the instruction sharing scenario. Compared with the information sharing scheme in the related art, in this application, there is no need for the user to input complex instructions, only basic operation information needs to be input, and the information to be recognized is extracted by a predetermined model to supplement and improve the first operation information in the first user instruction, and an accurate target operation instruction is obtained. The target operation instruction can accurately represent the user's intention, simplify the user's operation, improve the convenience of the user issuing instructions, and is applicable to instruction sharing scenarios such as intelligent driving scenarios and smart home scenarios.

[0109] The additional aspects and advantages of this application will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of this application. Description of the Drawings

[0110] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the following description of embodiments in conjunction with the accompanying drawings, in which:

[0111] Figure 1 It is a schematic flowchart of an interaction method provided by an embodiment of the present application;

[0112] Figure 2 It is a schematic flowchart of an interaction method provided by an embodiment of the present application;

[0113] Figure 3 It is a schematic flowchart of an interaction method provided by an embodiment of the present application;

[0114] Figure 4 It is a schematic structural diagram of an interaction device provided by an embodiment of the present application;

[0115] Figure 5 It is a schematic structural diagram of an interaction device provided by an embodiment of the present application;

[0116] Figure 6 It is a schematic structural diagram of an interaction device provided by an embodiment of the present application;

[0117] Figure 7 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed Embodiments

[0118] The embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present application, and should not be construed as limiting the present application.

[0119] The interaction method, device, electronic device, chip, and storage medium of the embodiments of the present application will be described below with reference to the accompanying drawings.

[0120] Figure 1 It is a schematic flowchart of an interaction provided by an embodiment of the present application.

[0121] As an implementation, the interaction method of the embodiments of the present application can be configured in an interaction device, and the interaction device can be applied to any electronic device so that the electronic device can perform an interaction function.

[0122] Among them, the electronic device can be any device with computing capabilities. For example, it can be a mobile terminal, and the mobile terminal can be a hardware device such as a mobile phone, a tablet computer, a personal digital assistant, a wearable device, a vehicle, etc., having various operating systems, touch screens, and / or display screens.

[0123] Such asFigure 1 As shown, the method may include the following steps:

[0124] Step 101, obtain a first user instruction, where the first user instruction includes first operation information;

[0125] Step 102, obtain the information to be recognized of a first application in a first terminal to determine a target operation instruction associated with the information to be recognized and the first operation information; the target operation instruction is used to trigger a second application in the second terminal to execute a target task, where the first terminal and the second terminal establish a network connection through a device identifier.

[0126] In this embodiment, the method is executed by the first terminal, and the first terminal is the terminal that provides the information to be recognized. In the scenario of instruction intercommunication, the first terminal is used to determine the target operation instruction and share it with the second terminal, and the second terminal executes the target operation instruction.

[0127] Before the first terminal determines the target operation instruction, it is necessary to obtain the first user instruction input by the user, and obtain the information to be recognized of the first application in the first terminal, and determine the target operation instruction through the information to be recognized and the first user instruction.

[0128] In one embodiment, the first operation information included in the first user instruction represents the operation that the user wants to perform. The user indicates the determined operation action to the first terminal through the first user instruction. From the perspective of the terminal device, the exact operation object is not known. In order to obtain the exact operation object, in this embodiment, the information to be recognized is obtained from the first application running in the first terminal. The information to be recognized includes information related to the operation object. The target operation instruction is determined by combining the information to be recognized and the first operation information. The target operation instruction includes the exact operation action and the operation object.

[0129] In one possible embodiment, the first terminal is a mobile terminal, such as a mobile phone, a tablet computer, a laptop computer, an augmented reality (AR) device, a virtual reality (VR) device, and so on.

[0130] In one possible embodiment, the second terminal includes a mobile phone, a smart watch, a vehicle, and the like.

[0131] In a possible embodiment, in a navigation scenario, when a user sees a certain restaurant in a first application (such as a social application) on a first terminal and wants to navigate to that restaurant, the user issues a voice command, such as "I want to go to this restaurant". The first user command obtained by the first terminal contains first operation information that indicates the user's intention to navigate, but it is not yet determined which specific restaurant to navigate to. Then, by identifying the information to be recognized in the first application of the first terminal, such as the picture information of the current application interface, the text information of the current application interface, or the positioning data of the current application interface, etc., the specific address of the restaurant (for example, No. X, XX Road) is obtained. By combining the above voice command and the information to be recognized, it can be obtained that the user needs to navigate to this restaurant at this time, and the navigation information is shared with the second terminal, realizing the information interaction in the intelligent cockpit scenario quickly and conveniently, without the need to manually enter the address on the second terminal (such as the in-vehicle infotainment system central control screen) again, greatly improving the efficiency of multi-terminal interaction and enhancing the user experience.

[0132] In a possible embodiment, in a call scenario, when a user sees a certain phone number in a first application (such as a contact list, a recruitment application) on a first terminal and wants to call this phone number, the user issues a voice command such as "I want to contact the person in the picture". The first user command obtained by the first terminal contains first operation information that indicates the user's intention to make a call, but it is not yet determined which specific phone number to call. By identifying the information to be recognized (such as a screenshot) in the first application of the first terminal, such as the picture information of the current interface and the text information of the current interface, the specific contact information of this person (such as phone number 139XXXXXXXX) is determined. By combining the above voice command and the information to be recognized, it can be obtained that the user needs to call and contact this person at this time, and the phone number is shared with the second terminal, realizing the information interaction in the intelligent cockpit scenario quickly and conveniently, without the need to manually enter the phone number on the second terminal (such as the in-vehicle infotainment system central control screen) again, greatly improving the efficiency of multi-terminal interaction and enhancing the user experience.

[0133] In a possible embodiment, there is no strict order between step 101 and step 102. Step 101 can be executed first, step 102 can be executed first, or step 101 and step 102 can be executed simultaneously.

[0134] Optionally, the first terminal and the second terminal are connected by means of communication.

[0135] Optionally, the communication means includes short - range communication and / or long - range communication. Short - range communication such as Bluetooth, ZigBee, etc., and long - range communication includes cellular networks, WIFI, etc. Optionally, the first terminal and the second terminal can directly establish a connection through short - range communication, or can establish a connection through short - range communication and / or long - range communication under the same user account, so as to quickly share instructions.

[0136] The first terminal collects the information to be recognized and the first user instruction and confirms the target operation instruction, and shares the target operation instruction to the second terminal through the network connection. The second terminal executes the target operation instruction, realizing the instruction sharing between the first terminal and the second terminal, and optimizing the user experience in the instruction sharing scenario. Compared with the information sharing scheme in the related technology, in this application, the user does not need to input complex instructions, only needs to input basic operation information. The information to be recognized is extracted by a predetermined model to supplement and improve the first operation information in the first user instruction, and an accurate target operation instruction is obtained. The target operation instruction can accurately represent the user's intention, simplifies the user's operation, improves the convenience of the user issuing instructions, and is applicable to instruction sharing scenarios such as intelligent driving scenarios and smart home scenarios.

[0137] Optionally, it further includes:

[0138] Based on the information to be recognized and the first operation information, determine the target operation instruction associated with the information to be recognized and the first operation information.

[0139] In this embodiment, the first operation information contains exact operation actions and inexact operation objects, and the information to be recognized contains exact operation objects. In order to obtain an instruction that the second terminal can execute, the information to be recognized and the first operation information are recognized to extract the exact operation object from the information to be recognized, and a target operation instruction is generated according to the exact operation object and the exact operation action. Such a target operation instruction is an instruction executable by the terminal.

[0140] In a possible embodiment, in a navigation scenario, when a user sees a certain restaurant in a first application (such as a social application) on a first terminal and wants to navigate to that restaurant, the user issues a voice command, such as "I want to go to this restaurant". The first terminal obtains a first user command, and the first operation information included in the first user command indicates the user's intention to navigate, but it is not yet determined which restaurant to navigate to. Then, by identifying the information to be recognized in the first application of the first terminal, such as the picture information of the current application interface, the text information of the current application interface, or the positioning data of the current application interface, etc., the specific address of the restaurant (such as, No. X, XX Road) is obtained. By combining the above first operation information and the information to be recognized, it can be obtained that the user needs to navigate to this restaurant at this time, and further determine the target operation command such as "Navigate to No. X, XX Road". By combining the information to be recognized and the first operation information to generate the target operation command, the terminal can obtain the command that the user wants to execute but is not clearly stated in the first operation information. The user does not need to further input more complex commands, and the terminal can accurately understand the user's intention, improving the user experience.

[0141] Optionally, the first user command is "I want to ride a bike to the park in the picture", and the information to be recognized is a screenshot in the first application of the first terminal. It is recognized that the address of the park is "the intersection of Road A and Road B". From the information to be recognized, it can be obtained that the user's intention is to ride to a certain destination. According to the information to be recognized, it can be determined that the user's destination is the intersection of Road A and Road B. Then, the target operation command is "Generate a bike ride to the intersection of Road A and Road B". The navigation information generated by the second application in the second terminal based on the target operation command is the navigation to ride to the intersection of Road A and Road B, which better meets the user's needs, and the generated navigation path will be shared between the first terminal and the second terminal.

[0142] In a possible embodiment, in a call scenario, when a user sees a certain phone number in a first application (such as an address book or a recruitment application) on a first terminal and wants to make a call, the user issues a voice command such as "I want to contact the person in the picture". The first operation information included in the first user command indicates the user's intention to make a call, but the specific phone number to be called has not been determined yet. By identifying the information to be recognized (such as a screenshot) in the first application on the first terminal, such as the picture information and text information of the current interface, the specific contact information of this person (such as phone number 139XXXXXXXX) is determined. By combining the above voice command and the information to be recognized, and by combining the above first operation information and the information to be recognized, it can be obtained that the user needs to contact this person's phone number at this time. Further, a target operation command such as "call 139XXXXXXXX" is determined. By generating a target operation command by combining the information to be recognized and the first operation information, the terminal can obtain the command that the user wants to execute but is not clearly stated in the first operation information. The user does not need to input more complex commands, and the terminal can accurately understand the user's intention, improving the user experience.

[0143] Optionally, determining a target operation command associated with the information to be recognized and the first operation information based on the information to be recognized and the first operation information includes:

[0144] Taking the first operation information and the information to be recognized as input parameters and inputting them into a predetermined model to obtain a matching result; wherein, the matching result is related to the target operation command.

[0145] In this embodiment, a pre-trained model is used to process the first operation information and the information to be recognized, extract the features therein, and process the feature data to obtain a matching result.

[0146] In a possible embodiment, the steps executed by the model are as follows: extracting feature vectors of the information to be recognized to obtain a first feature vector; extracting feature vectors of the first operation information in the first user command to obtain a second feature vector; determining the matching result according to the first feature vector and the second feature vector. By feature extraction, high-dimensional features are extracted from the first operation information and the information to be recognized, and processing is performed based on the features. The data that matches the features of the first operation information is extracted from the features of the information to be recognized as the matching result. By extracting features, the model can more accurately understand the user's intention in the first operation information and the exact operation object in the information to be recognized. Combining these features can accurately understand the user's intention and output the target operation command that the user wants to execute but is not clearly stated.

[0147] In a possible embodiment, the information to be recognized is a screenshot of a first application in a first terminal. Since the screenshot directly obtained by taking a screenshot has a high resolution and occupies a large storage space, directly inputting the screenshot into the model for analysis takes a long time. Moreover, analyzing the screenshot is often for the text therein, so appropriately reducing the resolution of the screenshot has little impact on the analysis effect of the model. Therefore, the screenshot is compressed, and then the compressed screenshot is divided into multiple small pictures. Feature extraction is performed on each small picture respectively to obtain multiple third feature vectors, and these third feature vectors are concatenated to obtain a first feature vector. By performing compression processing on the information to be recognized, the storage space occupied by the information to be recognized is reduced, the computing power required by the model to process the information to be recognized is reduced, the efficiency of the model to process information is improved, the efficiency of generating the target operation instruction is higher, and the user experience is enhanced.

[0148] In a possible embodiment, the first instruction data is voice data. The first instruction data is converted into an instruction text through speech-to-text technology, the instruction text is split into multiple words, the features corresponding to each word are extracted respectively and concatenated to obtain a second feature vector.

[0149] In a possible embodiment, the model is set in a server on the edge side or in the cloud. The first user instruction and the information to be recognized collected in the first terminal are sent to the model for processing, and a matching result is output, and the matching result is shared to the first terminal or the second terminal. The server is set on the edge side. Due to local processing, the processing and response rates of the model will be greatly improved; the server is set in the cloud. Due to sufficient computing power in the cloud, data synchronization and update can be achieved, and the accuracy is significantly improved. In this embodiment, it is not limited to whether the model is set in the cloud or on the edge side, and it can be set according to the current user scenario.

[0150] Optionally, it further includes:

[0151] Receiving at least one candidate address information, where the candidate address information is related to the information to be recognized and the first operation information;

[0152] Receiving a second user instruction, and obtaining the target operation instruction based on at least one of the candidate address information.

[0153] In this embodiment, in a navigation scenario, the information to be recognized and the first operation information in the first terminal are obtained, and the information to be recognized and the first operation information are processed to extract the corresponding address from the information to be recognized according to the intention in the first operation information. Since there are multiple addresses in the information to be recognized, multiple candidate address information is generated after processing, and multiple candidate address information is returned to the first terminal.

[0154] Since multiple candidate address information is received, the first terminal still cannot confirm an exact address. To generate an exact target operation instruction, the first terminal presents these candidate address information to the user for selection. The user inputs a second user instruction, and the first terminal receives the second user instruction to select one or more from the candidate address information to generate a target operation instruction. By selecting the candidate address information through the second user instruction and determining the target operation instruction, the user's intention can be accurately known, the user's instruction can be accurately executed, and the user experience is improved.

[0155] Optionally, the address information may be an entity name or detailed address information. Entity names such as "The Forbidden City", "National Museum", etc., and detailed address information such as "XXX, Haidian District, Beijing", "No. X, Road X", "Intersection of Road X and Road X", etc.

[0156] In a possible embodiment, the information to be recognized is a screenshot of a first application, and the screenshot contains an address. After processing the information to be recognized and the first operation information, only one candidate address information is generated. At this time, the user's destination is determined. Then, after receiving it, the first terminal can directly generate a target operation instruction based on this candidate address information and the first operation information.

[0157] In a possible embodiment, the information to be recognized is a screenshot of a first application, and the screenshot contains multiple addresses. After processing the information to be recognized and the first operation information, multiple candidate address information (Address A, Address B, Address C) is generated. After receiving these candidate address information, the first terminal needs to further confirm with the user the address the user wants to go to. The first terminal will display these candidate address information through a display, or, announce the candidate address information through voice, to prompt the user to select from these addresses. The user inputs a second user instruction "Go to Address A", then the first terminal can determine that the target operation instruction is "Navigate to Address A". Optionally, the user inputs a second user instruction "Go to Address A first, then go to Address B", then the first terminal can determine that the target operation instruction is "Navigate from the current location to Address A, and then navigate from A to Address B".

[0158] Optionally, receiving the second user instruction and obtaining the target operation instruction based on at least one of the candidate address information includes at least one of the following methods:

[0159] Receiving the user's voice instruction, selecting a target address from at least one of the candidate addresses, and obtaining the target operation instruction, where the target operation instruction includes the target address;

[0160] Receiving the operation instruction input by the user through a click operation or a selection operation, determining a target address from at least one of the candidate addresses, and obtaining the target operation instruction, where the target operation instruction includes the target address.

[0161] In this embodiment, there are various ways for the user to input the second user instruction. The user can input a voice instruction by directly speaking to the first terminal, or can manually operate on the first terminal. The input voice instruction includes a selection decision for the candidate address. After receiving multiple candidate address information, the first terminal displays the multiple candidate address information on the display screen. The user can click on the screen to select the candidate address information, or circle the candidate address information on the screen, or select the candidate address information through key operations. By obtaining the second user instruction in various ways, the user can input the instruction conveniently, and the interaction efficiency between the user and the terminal is higher, improving the user experience.

[0162] In a possible embodiment, after receiving multiple candidate address information (Address A, Address B, Address C), the first terminal displays these candidate address information on the screen. When the user says "Navigate to the first address" or "Navigate to Address A", after receiving the voice instruction, the first terminal converts the voice instruction into text and matches it with the candidate address information, and then can determine that the address the user wants to go to is Address A, and Address A is the target address. Then, the target operation instruction "Navigate to Address A" can be determined according to Address A.

[0163] In a possible embodiment, after receiving multiple candidate address information (Address A, Address B, Address C), the first terminal displays these candidate address information on the screen. When the user clicks on "Address A" or circles "Address A" on the touch screen, after receiving the operation instruction, the first terminal can determine that the address the user wants to go to is Address A, and Address A is the target address. Then, the target operation instruction "Navigate to Address A" can be determined according to Address A.

[0164] Optionally, it further includes:

[0165] Receiving at least one candidate call object information, where the candidate call object information is related to the to-be-identified information and the first operation information;

[0166] Receiving a second user instruction, and obtaining a target operation instruction based on at least one of the candidate call object information.

[0167] In this embodiment, in a call scenario, the to-be-identified information and the first operation information in the first terminal are obtained, and the to-be-identified information and the first operation information are processed to extract the corresponding call object from the to-be-identified information according to the intention in the first operation information. Since there are multiple call objects in the to-be-identified information, multiple candidate call object information is generated after processing, and multiple candidate call object information is returned to the first terminal.

[0168] Since multiple candidate call object information is received, the first terminal still cannot confirm a specific call object. To generate a specific target operation instruction, the first terminal displays the candidate call object information to the user for selection. The user inputs a second user instruction, and the first terminal receives the second user instruction to select one or more from the candidate call object information to generate a target operation instruction. By selecting the candidate call object information through the second user instruction and determining the target operation instruction, the user's intention can be accurately known, the user's instruction can be accurately executed, and the user experience is improved.

[0169] Optionally, the call object information may be an object name or a phone number, and the object name may be "Teacher A", "Teacher B", etc. in the address book.

[0170] In a possible embodiment, the information to be recognized is a screenshot of a first application, and the screenshot contains a call object. After processing the information to be recognized and the first operation information, only one candidate call object information "Teacher A" is generated. At this time, the user's destination is determined, so the first terminal can directly generate a target operation instruction "Call Teacher A" according to this candidate call object information and the first operation information after receiving it.

[0171] In a possible embodiment, the information to be recognized is a screenshot of a first application, and the screenshot contains multiple call objects. After processing the information to be recognized and the first operation information, multiple candidate call object information (Call Object A, Call Object B, Call Object C) is generated. After receiving these candidate call object information, the first terminal needs to further confirm with the user the call object the user wants to call. The first terminal will display these candidate call object information through a display, or, announce the candidate call object information through voice to prompt the user to select from these call objects. The user inputs a second user instruction "Call Call Object A", then the first terminal can determine that the target operation instruction is "Contact Call Object A".

[0172] Optionally, the receiving the second user instruction and obtaining the target operation instruction based on at least one of the candidate call object information includes at least one of the following methods:

[0173] Receiving the user's voice instruction, selecting a target call object from at least one of the candidate call objects, and obtaining the target operation instruction, where the target operation instruction includes the target call object;

[0174] Receiving the operation instruction input by the user through a click operation or a selection operation, determining a target call object from at least one of the candidate call objects, and obtaining the target operation instruction, where the target operation instruction includes the target call object.

[0175] In this embodiment, there are various ways for the user to input the second user instruction. The user can input a voice instruction by directly speaking to the first terminal, or can manually operate on the first terminal. The input voice instruction includes a selection decision for the candidate call object. After receiving multiple candidate call object information, the first terminal displays the multiple candidate call object information on the display screen. The user can click on the screen to select the candidate call object information, or circle the candidate call object information on the screen, or select the candidate call object information through key operations. By obtaining the second user instruction in multiple ways, the user can input the instruction conveniently, and the interaction efficiency between the user and the terminal is higher, improving the user experience.

[0176] In a possible embodiment, after receiving multiple candidate call object information (call object A, call object B, call object C), the first terminal displays these candidate call object information on the screen. When the user says "contact the first call object" or "call the phone of call object A", after receiving the voice instruction, the first terminal converts the voice instruction into text and matches it with the candidate call object information, and then it can be determined that the user wants to contact call object A, and call object A is the target call object. Then the target operation instruction "contact call object A" can be determined according to call object A.

[0177] In a possible embodiment, after receiving multiple candidate call object information (call object A, call object B, call object C), the first terminal displays these candidate call object information on the screen. When the user clicks on "call object A" on the touch screen, or circles "call object A", after the first terminal receives the operation instruction, it can be determined that the user wants to contact call object A, and call object A is the target call object. Then the target operation instruction "contact call object A" can be determined according to call object A.

[0178] Optionally, the predetermined model is a multimodal large language model trained with training data including information to be recognized and operation information.

[0179] In this embodiment, the multimodal large language model (MLLMs) is a deep learning model that combines a large language model (LLM) and a large vision model (LVM), and is capable of processing and understanding various types of data, such as text, images, and audio.

[0180] Multimodal large language models can process various types of input data, such as text, images, audio, video, etc., and can generate outputs in multiple modalities, such as generating images based on text, or generating descriptions based on images. During the training process, multimodal large language models are trained on large-scale datasets containing text, visual, auditory, and sometimes even sensor data, enabling them to establish connections between different modalities, thus supporting tasks that require understanding and generating content across multiple data types.

[0181] In one possible embodiment, by collecting historical data of user interactions with the terminal, the information to be recognized and operation information therein are obtained to generate training data. Among them, the operation information is text or voice data, and the information to be recognized is data in modalities such as text, images, audio, video, etc. One sample in the training data contains the information to be recognized, operation information, and label data. The label data is the matching result corresponding to the information to be recognized and operation information in the sample, such as candidate address information or candidate target object information.

[0182] Each sample in the training data is input into the multimodal large language model for processing and then the matching result is output. The matching result is compared with the label data corresponding to the sample, and a loss function representing the size of the difference is calculated. With the goal of reducing the loss function, the parameters in the multimodal large language model are optimized, and after multiple rounds of optimization iterations of the parameters, a trained predetermined model can be obtained.

[0183] Processing the multimodal large language model through the information to be recognized and operation information can improve the multimodal large language model's ability to understand various modal data, thereby better processing information in cross-modal tasks, improving the efficiency and accuracy of matching, and enabling users to obtain more accurate matching results, enhancing the user experience.

[0184] Optionally, the method further includes:

[0185] Obtain a first wake-up instruction, where the first wake-up instruction is triggered by an operation on the first terminal;

[0186] In response to the first wake-up instruction, collect the user's input information to obtain the first user instruction.

[0187] In this embodiment, before the user inputs the first user instruction, the module or application responsible for voice recognition in the first terminal can be awakened by the first wake-up instruction, so that the first terminal starts to receive voice or text instructions. After receiving the first wake-up instruction, the first terminal learns that the user intends to input an instruction, and then starts to call the modules or applications therein, collect the user's input information, and obtain the first user instruction.

[0188] In a possible embodiment, the voice assistant in the first terminal collects and analyzes the user's instructions. After the user utters the first wake-up instruction "Little X classmate", the first terminal wakes up the voice assistant application and is ready to receive the user's voice instructions.

[0189] Optionally, obtaining the information to be recognized of the first application in the first terminal includes:

[0190] Obtaining the image data or text data in the first terminal as the information to be recognized.

[0191] In this embodiment, the instruction that the user wants to give is related to the information displayed by the first application in the first terminal. The first terminal can extract a screenshot of the application or obtain the text in the application from the background as the information to be recognized.

[0192] In a possible embodiment, in the content displayed in the first application (a certain blog APP) in the first terminal (such as a mobile phone), the user sees a restaurant and wants to go to this restaurant for dinner. Then the user can input the first user instruction "I want to go to the restaurant in the picture" by voice. The first terminal intercepts the image of the first application or obtains the text in the first application as the data to be recognized.

[0193] Optionally, obtaining the image data or text data in the first terminal as the information to be recognized includes at least one of the following methods:

[0194] Intercepting the image displayed on the first terminal to obtain the information to be recognized;

[0195] Obtaining the image of the area specified by the user in the first terminal as the information to be recognized;

[0196] Obtaining the image collected by the first terminal as the information to be recognized.

[0197] In this embodiment, there are various ways for the first terminal to obtain the image data of the first application. It can take a screenshot of the entire screen, or obtain the image of the area selected by the user, or schedule the sensor in the first terminal to collect an image.

[0198] In a possible embodiment, in the content displayed in the first application (a certain blog APP) in the first terminal (such as a mobile phone), the user sees a restaurant and wants to go to this restaurant for dinner. Then the user can input the first user instruction "I want to go to the restaurant in the picture" by voice. The first terminal intercepts the image of the first application or obtains the text in the first application as the data to be recognized, or determines the area circled by the user and intercepts the image of the area circled by the user. By specifying the information to be recognized in the first terminal in various ways, the user can input instructions conveniently, and the interaction efficiency between the user and the terminal is higher, improving the user experience.

[0199] Optionally, the method further includes:

[0200] Receiving result information sent by a second terminal, where the result information is related to the execution progress or execution result of the target task.

[0201] In this embodiment, after the second terminal executes the target operation instruction, result information will be generated, and the result information will be transmitted to the first terminal through a data transmission means, so that the user can understand the execution situation of the target operation instruction. Through the sharing of the result information between terminals, the user can intuitively understand whether their intention has been executed and whether it has been executed accurately, improving the user's interaction experience.

[0202] In a possible embodiment, in a navigation scenario, after the second terminal schedules the second application therein to execute the target operation instruction, result information of "navigation started" will be generated and synchronized to the first terminal, and the user can then know that navigation has started.

[0203] In a possible embodiment, in a navigation scenario, after the second terminal schedules the second application therein to execute the target operation instruction, result information of "navigation route" will be generated and synchronized to the first terminal, and the user can view the navigation route on the first terminal.

[0204] In a possible embodiment, in a call scenario, after the second terminal schedules the second application therein to execute the target operation instruction, result information of "call made" will be generated and synchronized to the first terminal.

[0205] In a possible embodiment, in a call scenario, after the second terminal schedules the second application therein to execute the target operation instruction, result information of "call voice" will be generated and synchronized to the first terminal, and the user can view the call being answered simultaneously on the first terminal.

[0206] Figure 2 It is a schematic diagram of an interaction process provided by an embodiment of the present application. As Figure 2 shown, the method includes:

[0207] Step 201, receiving a target operation instruction, where the target operation instruction is used to instruct the second terminal to execute a target task; the target operation instruction is related to the first operation information and the information to be recognized of the first terminal, and the first terminal and the second terminal establish a network connection through a device identifier;

[0208] Step 202, triggering the second application to execute the target task according to the target operation instruction.

[0209] In this embodiment, the target operation instruction determined based on the first operation information of the first terminal and the information to be recognized is received by the second terminal, and the second terminal is the execution end of the target operation instruction. The user indicates the determined operation action to the first terminal through the first user instruction. From the perspective of the terminal device, the exact operation object is not known. To obtain the exact operation object, in this embodiment, the information to be recognized is obtained from the first application running on the first terminal. The information to be recognized contains information related to the operation object. The information to be recognized and the first operation information are combined for recognition to determine the target operation instruction. The target operation instruction contains the exact operation action and operation object.

[0210] In a possible embodiment, the first terminal is a mobile terminal, such as a mobile phone, a tablet computer, a laptop computer, an augmented reality (AR) device, a virtual reality (VR) device, and so on.

[0211] In a possible embodiment, the second terminal includes a mobile phone, a smart watch, a vehicle, and the like.

[0212] In a possible embodiment, in a navigation scenario, the user sees a certain restaurant in the first application (such as a social application) on the first terminal and wants to navigate to that restaurant. The user issues a voice instruction, such as "I want to go to this restaurant". The first terminal obtains the first user instruction, and the first operation information contained in the first user instruction indicates the user's intention to navigate, but which restaurant to navigate to has not been determined yet. Then, by recognizing the information to be recognized in the first application of the first terminal, such as the picture information of the current application interface, the text information of the current application interface, or the positioning data of the current application interface, etc., the specific address of the restaurant (such as, No. X, XX Road) is obtained. By combining the above voice instruction and the information to be recognized, it can be obtained that the user needs to navigate to this restaurant at this time, and the navigation information is shared with the second terminal, realizing the information interaction in the intelligent cockpit scenario quickly and conveniently, without the need to manually enter the address on the second terminal (such as the in-vehicle infotainment system screen) again, greatly improving the efficiency of multi-terminal interaction and enhancing the user experience.

[0213] In a possible embodiment, in a call scenario, when seeing a certain phone number in a first application (such as an address book, a recruitment application) in a first terminal and wanting to make a call to this number, the user issues a voice command such as "I want to contact the person in the picture". The first operation information included in the first user command indicates the intention of the user to make a call, but the specific phone number to be called has not been determined yet. By identifying the information to be recognized (such as a screenshot) in the first application of the first terminal, such as the picture information of the current interface and the text information of the current interface, the specific contact information of this person (such as phone number 139XXXXXXXX) is determined. By combining the above voice command and the information to be recognized, it can be obtained that the user needs to make a call to contact this person at this time and share the phone number with a second terminal, which quickly and conveniently realizes information interaction in the intelligent cockpit scenario without having to manually enter the phone number on the second terminal (such as the in-vehicle infotainment system central control screen), greatly improving the efficiency of multi-terminal interaction and enhancing the user experience.

[0214] In a possible embodiment, a pre-trained model is used to process the first operation information and the information to be recognized, extract the features therein, and process the feature data to obtain a matching result.

[0215] In a possible embodiment, the steps performed by the model are as follows: extracting features from the information to be recognized to obtain a first feature vector; extracting features from the first operation information in the first user command to obtain a second feature vector; determining the matching result according to the first feature vector and the second feature vector. Through feature extraction, high-dimensional features are extracted from the first operation information and the information to be recognized, and based on these features, data that matches the features of the first operation information is extracted from the features of the information to be recognized as the matching result. By extracting features, the model can more accurately understand the intention of the user in the first operation information and the exact operation object in the information to be recognized. Combining these features can accurately understand the intention of the user and output the target operation command that the user wants to execute but has not made clear.

[0216] In a possible embodiment, the information to be recognized is a screenshot of a first application in a first terminal. Since the screenshot directly obtained by taking a screenshot has a high resolution and occupies a large storage space, directly inputting the screenshot into the model for analysis takes a long time. Moreover, the analysis of the screenshot is often carried out for the text therein, so appropriately reducing the resolution of the screenshot has little impact on the analysis effect of the model. Therefore, the screenshot is compressed, and then the compressed screenshot is divided into multiple small pictures. Feature extraction is performed on each small picture to obtain multiple third feature vectors, and these third feature vectors are spliced to obtain a first feature vector. By performing compression processing on the information to be recognized, the storage space occupied by the information to be recognized is reduced, the computing power required by the model to process the information to be recognized is reduced, the efficiency of the model to process information is improved, the efficiency of generating the target operation instruction is higher, and the user experience is enhanced.

[0217] In a possible embodiment, the first instruction data is voice data. The first instruction data is converted into an instruction text through speech-to-text technology, the instruction text is split into multiple words, the features corresponding to each word are extracted respectively and spliced to obtain a second feature vector.

[0218] In a possible embodiment, the model is set in a server on the edge side or in the cloud. The first user instruction and the information to be recognized collected in the first terminal are sent to the model for processing, and a matching result is output, and the matching result is shared to the first terminal or the second terminal. The server is set on the edge side. Due to local processing, the processing and response rate of the model will be greatly improved; the server is set in the cloud. Due to sufficient computing power in the cloud, data synchronization and update can be achieved, and the accuracy is significantly improved. In this embodiment, it is not limited to whether the model is set in the cloud or on the edge side, and it can be set according to the current user scenario.

[0219] Optionally, the method further includes:

[0220] Receiving at least one candidate address information, where the candidate address information is related to the information to be recognized and the first operation information;

[0221] Receiving a second user instruction, and confirming at least one of the candidate address information based on the second user instruction to obtain the target operation instruction.

[0222] In this embodiment, in a navigation scenario, the information to be recognized and the first operation information in the first terminal are obtained, and the information to be recognized and the first operation information are processed to extract the corresponding address from the information to be recognized according to the intention in the first operation information. Since there are multiple addresses in the information to be recognized, multiple candidate address information is generated after processing, and multiple candidate address information is returned to the second terminal.

[0223] Since multiple candidate address information is received, the second terminal still cannot confirm an exact address. To generate an exact target operation instruction, the second terminal displays the candidate address information to the user for selection. The user inputs a second user instruction, and the second terminal receives the second user instruction to select one or more from the candidate address information to generate a target operation instruction. By selecting the candidate address information through the second user instruction and determining the target operation instruction, the user's intention can be accurately known, the user's instruction can be accurately executed, and the user experience is improved.

[0224] Optionally, the address information may be an entity name or detailed address information. Entity names such as "The Forbidden City", "National Museum of China", etc., and detailed address information such as "XXX, Haidian District, Beijing", "No. X, XX Road", "Intersection of XX Road and XX Road", etc.

[0225] In a possible embodiment, the information to be recognized is a screenshot of a first application, and the screenshot contains an address. After processing the information to be recognized and the first operation information, only one candidate address information is generated. At this time, the user's destination is determined, so the second terminal can directly generate a target operation instruction based on this candidate address information and the first operation information after receiving it.

[0226] In a possible embodiment, the information to be recognized is a screenshot of a first application, and the screenshot contains multiple addresses. After processing the information to be recognized and the first operation information, multiple candidate address information (Address A, Address B, Address C) is generated. After receiving these candidate address information, the second terminal needs to further confirm with the user the address the user wants to go to. The second terminal will display these candidate address information through a display, or, announce the candidate address information by voice to prompt the user to select from these addresses. The user inputs a second user instruction "Go to Address A", then the second terminal can determine that the target operation instruction is "Navigate to Address A". Optionally, the user inputs a second user instruction "Go to Address A first, then go to Address B", then the second terminal can determine that the target operation instruction is "Navigate from the current location to Address A, and then navigate from A to Address B".

[0227] Optionally, receiving the second user instruction and confirming at least one of the candidate address information based on the second user instruction to obtain the target operation instruction includes at least one of the following methods:

[0228] Receiving the user's voice instruction, selecting a target address from at least one of the candidate addresses to obtain the target operation instruction, and the target operation instruction includes the target address;

[0229] Receive the operation instruction input by the user through a click operation or a selection operation, determine the target address from at least one of the candidate addresses, and obtain the target operation instruction, where the target operation instruction includes the target address.

[0230] In this embodiment, there are various ways for the user to input the second user instruction. The user can input a voice instruction by directly speaking to the first terminal, or can manually operate on the second terminal. The input voice instruction includes a selection decision for the candidate address. After receiving multiple candidate address information, the second terminal displays the multiple candidate address information on the display screen. The user can click on the screen to select the candidate address information, or circle the candidate address information on the screen, or select the candidate address information through a key operation. By obtaining the second user instruction in various ways, the user can input the instruction conveniently, and the interaction efficiency between the user and the terminal is higher, improving the user experience.

[0231] In a possible embodiment, after receiving multiple candidate address information (Address A, Address B, Address C), the second terminal displays these candidate address information on the screen. The user says "Navigate to the first address" or "Navigate to Address A". After receiving the voice instruction, the second terminal converts the voice instruction into text and matches it with the candidate address information, and then can determine that the user wants to go to Address A, and Address A is the target address. Then, the target operation instruction "Navigate to Address A" can be determined according to Address A.

[0232] In a possible embodiment, after receiving multiple candidate address information (Address A, Address B, Address C), the second terminal displays these candidate address information on the screen. The user clicks on "Address A" or circles "Address A" on the touch screen. After receiving the operation instruction, the second terminal can determine that the user wants to go to Address A, and Address A is the target address. Then, the target operation instruction "Navigate to Address A" can be determined according to Address A.

[0233] Optionally, the method further includes:

[0234] Receive the target address sent by the first terminal to obtain the target operation instruction, where the target operation instruction includes the target address.

[0235] In this embodiment, the first terminal receives at least one candidate address information, where the candidate address information is related to the information to be recognized and the first operation information; receives the second user instruction, and obtains the target operation instruction based on at least one of the candidate address information. The first terminal directly sends the selected address to the second terminal, and the second terminal can perform navigation according to the address. By interacting with the first terminal, the second terminal can efficiently obtain the target address selected by the user and execute the target operation instruction, improving the interaction efficiency between the terminals, and can execute the user's instruction more efficiently, improving the user experience.

[0236] Optionally, the method further includes:

[0237] Receiving at least one piece of candidate call object information, where the candidate call object information is related to the information to be recognized and the first operation information;

[0238] Receiving a second user instruction, and confirming at least one piece of the candidate call object information based on the second user instruction to obtain a target operation instruction.

[0239] In this embodiment, in a call scenario, the information to be recognized and the first operation information in the first terminal are obtained, and the information to be recognized and the first operation information are processed to extract a corresponding call object from the information to be recognized according to the intention in the first operation information. Since there are multiple call objects in the information to be recognized, multiple pieces of candidate call object information are generated after processing, and the multiple pieces of candidate call object information are returned to the second terminal.

[0240] Since multiple pieces of candidate call object information are received, the second terminal still cannot confirm an exact call object. To generate an exact target operation instruction, the second terminal displays these candidate call object information to the user for selection. The user inputs a second user instruction, and the second terminal receives the second user instruction to select one or more from the candidate call object information to generate a target operation instruction. By selecting the candidate call object information through the second user instruction and determining the target operation instruction, the user's intention can be accurately known, the user's instruction can be accurately executed, and the user experience is improved.

[0241] Optionally, the call object information may be an object name or a telephone number, and the object name may be, for example, "Teacher A", "Teacher B", etc. in the address book.

[0242] In a possible embodiment, the information to be recognized is a screenshot of a first application, and the screenshot contains a call object. After processing the information to be recognized and the first operation information, only one piece of candidate call object information "Teacher A" is generated. At this time, the user's destination is determined, so the second terminal can directly generate a target operation instruction "Call Teacher A" based on this candidate call object information and the first operation information after receiving it.

[0243] In a possible embodiment, the information to be recognized is a screenshot of a first application, and the screenshot contains multiple call objects. After processing the information to be recognized and the first operation information, multiple candidate call object information (Call Object A, Call Object B, Call Object C) is generated. After receiving these candidate call object information, the second terminal needs to further confirm with the user which call object the user wants to call. The second terminal will display these candidate call object information on the display, or announce the candidate call object information through voice, so as to prompt the user to select from these call objects. The user inputs the second user instruction "Call Call Object A", then the second terminal can determine that the target operation instruction is "Contact Call Object A".

[0244] Optionally, receiving the second user instruction and confirming at least one of the candidate call object information based on the second user instruction to obtain a target operation instruction includes at least one of the following methods:

[0245] Receiving a voice instruction from the user, selecting a target call object from at least one of the candidate call objects, and obtaining the target operation instruction, where the target operation instruction includes the target call object;

[0246] Receiving an operation instruction input by the user through a click operation or a selection operation, determining a target call object from at least one of the candidate call objects, and obtaining the target operation instruction, where the target operation instruction includes the target call object.

[0247] In this embodiment, there are various ways for the user to input the second user instruction. The user can input a voice instruction by directly speaking to the second terminal, or can manually operate on the second terminal. The input voice instruction contains a selection decision for the candidate call object. After receiving multiple candidate call object information, the second terminal displays the multiple candidate call object information on the display screen. The user can click on the screen to select the candidate call object information, or circle the candidate call object information on the screen, or select the candidate call object information through a key operation. By obtaining the second user instruction in multiple ways, the user can input the instruction conveniently, and the interaction efficiency between the user and the terminal is higher, improving the user experience.

[0248] In a possible embodiment, after receiving multiple candidate call object information (Call Object A, Call Object B, Call Object C), the second terminal displays these candidate call object information on the screen. The user says "Contact the first call object" or "Call the phone number of Call Object A". After receiving the voice instruction, the second terminal converts the voice instruction into text and matches it with the candidate call object information, and then it can be determined that the user wants to contact Call Object A, and Call Object A is the target call object. Then the target operation instruction "Contact Call Object A" can be determined according to Call Object A.

[0249] In a possible embodiment, after receiving multiple candidate call object information (call object A, call object B, call object C), the second terminal displays the candidate call object information on the screen. When the user clicks on "call object A" or circles "call object A" on the touch screen, and the second terminal receives the operation instruction, it can determine that the user wants to contact call object A, and call object A is the target call object. Then, the target operation instruction "contact call object A" can be determined according to call object A.

[0250] Optionally, the method further includes:

[0251] Receiving the target call object sent by the first terminal to obtain the target operation instruction, where the target operation instruction includes the target call object.

[0252] In this embodiment, the first terminal receives at least one call object information, and the candidate call object information is related to the information to be recognized and the first operation information; receives the second user instruction, and obtains the target operation instruction based on at least one candidate call object information. The first terminal directly sends the selected call object to the second terminal, and the second terminal can make a call according to the call object. By interacting with the first terminal, the second terminal can efficiently obtain the target call object selected by the user and execute the target operation instruction, improving the interaction efficiency between terminals, more efficiently executing the user's instructions, and enhancing the user experience.

[0253] Optionally, the method further includes:

[0254] Generating result information, where the result information is related to the execution progress or execution result of the target task; or,

[0255] Generating result information and sending the result information to the first terminal, where the result information is related to the execution progress or execution result of the target task.

[0256] In this embodiment, after the second terminal executes the target operation instruction, result information will be generated and transmitted to the first terminal through data transmission means so that the user can understand the execution situation of the target operation instruction.

[0257] In a possible embodiment, in a navigation scenario, after the second terminal schedules the second application therein to execute the target operation instruction, a result information of "navigation has started" will be generated and synchronized to the first terminal, and the user can then know that the navigation has started.

[0258] In a possible embodiment, in a navigation scenario, after the second terminal schedules the second application therein to execute the target operation instruction, a result information of "navigation route" will be generated and synchronized to the first terminal, and the user can view the navigation route on the first terminal.

[0259] In a possible embodiment, in a call scenario, after the second terminal schedules the second application therein to execute the target operation instruction, it generates result information of "call made" and synchronizes this information to the first terminal.

[0260] In a possible embodiment, in a call scenario, after the second terminal schedules the second application therein to execute the target operation instruction, it generates result information of "call voice" and synchronizes this information to the first terminal, and the user can view and answer the call on the first terminal.

[0261] Figure 3 It is a schematic flowchart of another interaction method provided by an embodiment of the present application, as Figure 3 shown, and this method includes the following steps:

[0262] Step 301, obtain a first user instruction, where the first user instruction includes first operation information;

[0263] Step 302, obtain information to be recognized of a first application in the first terminal;

[0264] Step 303, determine a target operation instruction based on the information to be recognized and the first operation information;

[0265] Step 304, trigger the second application of the second terminal to execute a target task based on the target operation instruction, where the first terminal and the second terminal establish a network connection through a device identifier.

[0266] In this embodiment, the method is implemented by a system composed of a first terminal and a second terminal, and the first terminal is the terminal that provides the information to be recognized. In the scenario of instruction intercommunication, the first terminal is used to determine the target operation instruction and share it with the second terminal, and the second terminal executes the target operation instruction.

[0267] Before the first terminal determines the target operation instruction, it is necessary to obtain the first user instruction input by the user and obtain the information to be recognized of the first application in the first terminal. The target operation instruction is determined through the information to be recognized and the first user instruction.

[0268] In an embodiment, the first operation information included in the first user instruction represents the operation that the user wants to perform. The user indicates the determined operation action to the first terminal through the first user instruction. From the perspective of the terminal device, the exact operation object is not known. To obtain the exact operation object, in this embodiment, the information to be recognized is obtained from the first application running in the first terminal. The information to be recognized includes information related to the operation object. The target operation instruction is determined by combining the information to be recognized and the first operation information. The target operation instruction includes the exact operation action and operation object.

[0269] In a possible embodiment, the first terminal is a mobile terminal, such as a mobile phone, a tablet computer, a laptop computer, an augmented reality (AR) device, a virtual reality (VR) device, and so on.

[0270] In a possible embodiment, the second terminal includes a mobile phone, a smart watch, a vehicle, etc.

[0271] In a possible embodiment, in a navigation scenario, when a user sees a certain restaurant in a first application (such as a social application) on the first terminal and wants to navigate to that restaurant, the user issues a voice command, such as "I want to go to this restaurant". The first user instruction obtained by the first terminal contains first operation information indicating the user's intention to navigate, but the specific restaurant to which the user wants to navigate has not been determined yet. Then, by identifying the information to be recognized in the first application of the first terminal, such as the picture information of the current application interface, the text information of the current application interface, or the positioning data of the current application interface, etc., the specific address of the restaurant (such as, No. X, XX Road) is obtained. By combining the above voice command and the information to be recognized, it can be obtained that the user needs to navigate to this restaurant at this time, and the navigation information is shared with the second terminal, quickly and conveniently realizing information interaction in the intelligent cockpit scenario, without the need to manually enter the address on the second terminal (such as the in-vehicle infotainment system central control screen) again, greatly improving the efficiency of multi-terminal interaction and enhancing the user experience.

[0272] In a possible embodiment, in a call scenario, when a user sees a certain phone number in a first application (such as a contact book, a recruitment application) on the first terminal and wants to call this phone number, the user issues a voice command such as "I want to contact the person in the picture". The first user instruction obtained by the first terminal contains first operation information indicating the user's intention to make a call, but the specific phone number to be called has not been determined yet. By identifying the information to be recognized (such as a screenshot) in the first application of the first terminal, such as the picture information of the current interface and the text information of the current interface, the specific contact information of this person (such as phone number 139XXXXXXXX) is determined. By combining the above voice command and the information to be recognized, it can be obtained that the user needs to call and contact this person at this time, and the phone number is shared with the second terminal, quickly and conveniently realizing information interaction in the intelligent cockpit scenario, without the need to manually enter the phone number on the second terminal (such as the in-vehicle infotainment system central control screen) again, greatly improving the efficiency of multi-terminal interaction and enhancing the user experience.

[0273] In a possible embodiment, there is no limitation on the execution order of steps 301, 302, 303, and 304. These steps can be executed sequentially or synchronously.

[0274] Optionally, the first terminal and the second terminal are connected by means of communication.

[0275] Optionally, the communication means includes short-distance communication and / or long-distance communication. Short-distance communication includes Bluetooth, ZigBee, etc., and long-distance communication includes cellular networks, Wi-Fi, etc. Optionally, the first terminal and the second terminal can directly establish a connection through short-distance communication, or can establish a connection through short-distance communication and / or long-distance communication under the same user account, so as to quickly share instructions.

[0276] Optionally, determining the target operation instruction based on the information to be recognized and the first operation information includes:

[0277] Taking the first operation information and the information to be recognized as input parameters, and inputting them into a predetermined model to obtain a matching result; wherein, the matching result is related to the target operation instruction.

[0278] In this embodiment, the first operation information and the information to be recognized are processed by a predetermined model, features are extracted therefrom, and the feature data is processed to obtain a matching result.

[0279] In a possible embodiment, the steps executed by the model are as follows: extracting features from the information to be recognized to obtain a first feature vector; extracting features from the first operation information in the first user instruction to obtain a second feature vector; determining the matching result according to the first feature vector and the second feature vector. Through feature extraction, high-dimensional features are extracted from the first operation information and the information to be recognized, and processing is performed based on the features, and data that matches the features of the first operation information is extracted from the features of the information to be recognized as the matching result. By extracting features, the model can more accurately understand the user's intention in the first operation information and the exact operation object in the information to be recognized, and combine these features to accurately understand the user's intention and output the target operation instruction that the user wants to execute but has not clearly defined.

[0280] In a possible embodiment, the information to be recognized is a screenshot of a first application in a first terminal. Since the resolution of the screenshot directly obtained by taking a screenshot is relatively high, it occupies a large storage space, and it takes a long time to directly input the screenshot into the model for analysis. And analyzing the screenshot is often to analyze the text therein, so appropriately reducing the resolution of the screenshot has little impact on the effect of model analysis. Therefore, the screenshot is compressed, and then the compressed screenshot is divided into multiple small pictures, and feature extraction is performed on each small picture respectively to obtain multiple third feature vectors, and these third feature vectors are spliced to obtain a first feature vector. By performing compression processing on the information to be recognized, the storage space occupied by the information to be recognized is reduced, the computing power required for the model to process the information to be recognized is reduced, the efficiency of the model to process the information is improved, the efficiency of generating the target operation instruction is higher, and the user experience is improved.

[0281] In a possible embodiment, the first instruction data is voice data. The first instruction data is converted into an instruction text through speech-to-text technology. The instruction text is split into multiple words, and the features corresponding to each word are extracted and spliced to obtain a second feature vector.

[0282] In a possible embodiment, the model is set in a server on the edge side or in the cloud. The first user instruction and the information to be recognized collected in the first terminal are sent to the model for processing, and a matching result is output. The matching result is shared to the first terminal or the second terminal. When the server is set on the edge side, due to local processing, the processing and response rate of the model will be greatly improved; when the server is set in the cloud, due to sufficient computing power in the cloud, data synchronization and update can be achieved, and the accuracy is significantly improved. In this embodiment, it is not limited to whether the model is set in the cloud or on the edge side, and it can be set according to the current user scenario.

[0283] Optionally, it further includes:

[0284] Output at least one candidate address information through the predetermined model, and the candidate address information is related to the information to be recognized and the first operation information;

[0285] Receive a second user instruction, and confirm at least one of the candidate address information based on the second user instruction to obtain a target operation instruction.

[0286] In this embodiment, the execution subject of the step can be the first terminal or the second terminal. In a navigation scenario, the information to be recognized and the first operation information in the first terminal are obtained, and the information to be recognized and the first operation information are processed to extract the corresponding address from the information to be recognized according to the intention in the first operation information. Since there are multiple addresses in the information to be recognized, multiple candidate address information are generated after processing, and multiple candidate address information are returned to the first terminal.

[0287] Since multiple candidate address information are received, the terminal still cannot confirm an exact address. To generate an exact target operation instruction, the terminal displays these candidate address information to the user for selection. The user inputs a second user instruction, and the terminal receives the second user instruction to select one or more from the candidate address information to generate a target operation instruction. By selecting the candidate address information through the second user instruction and determining the target operation instruction, the user's intention can be accurately known, the user's instruction can be accurately executed, and the user experience is improved.

[0288] Optionally, the address information can be an entity name or a detailed address information. Entity nouns such as "The Forbidden City", "National Museum", etc., and detailed address information such as "XXX, Haidian District, Beijing", "No. X, XX Road", "Intersection of XX Road and XX Road", etc.

[0289] In a possible embodiment, the information to be recognized is a screenshot of a first application, and the screenshot contains an address. After processing the information to be recognized and the first operation information, only one candidate address information is generated. At this time, the user's destination is determined. Then, after receiving it, the terminal can directly generate a target operation instruction according to this candidate address information and the first operation information.

[0290] In a possible embodiment, the information to be recognized is a screenshot of a first application, and the screenshot contains multiple addresses. After processing the information to be recognized and the first operation information, multiple candidate address information (Address A, Address B, Address C) are generated. After receiving these candidate address information, the terminal needs to further confirm with the user the address the user wants to go to. The terminal will display these candidate address information through a display, or announce the candidate address information through voice, to prompt the user to select from these addresses. The user inputs a second user instruction "Go to Address A", then the terminal can determine that the target operation instruction is "Navigate to Address A". Optionally, the user inputs a second user instruction "Go to Address A first, then go to Address B", then the terminal can determine that the target operation instruction is "Navigate from the current location to Address A, and then navigate from A to Address B".

[0291] Optionally, receiving a second user instruction, and based on the second user instruction, confirming at least one of the candidate address information to obtain a target operation instruction, includes at least one of the following methods:

[0292] Receiving a voice instruction from the user, selecting a target address from at least one of the candidate addresses, to obtain the target operation instruction, where the target operation instruction includes the target address;

[0293] Receiving an operation instruction input by the user through a click operation or a selection operation, determining a target address from at least one of the candidate addresses, to obtain the target operation instruction, where the target operation instruction includes the target address.

[0294] In this embodiment, the execution subject of the step may be a first terminal or a second terminal. The way for the user to input the second user instruction includes multiple types. The user can input a voice instruction by directly speaking to the terminal, or can manually operate on the terminal. The input voice instruction contains a selection decision for the candidate address. After the terminal receives multiple candidate address information, it displays the multiple candidate address information on the display screen. The user can click on the screen to select the candidate address information, or circle the candidate address information on the screen, or select the candidate address information through a key operation. By obtaining the second user instruction in multiple ways, the user can input the instruction conveniently, and the interaction efficiency between the user and the terminal is higher, improving the user experience.

[0295] In a possible embodiment, after receiving multiple candidate address information (Address A, Address B, Address C), the terminal displays these candidate address information on the screen. When the user says "Navigate to the first address" or "Navigate to Address A", after receiving the voice command, the terminal converts the voice command into text and matches it with the candidate address information, and then it can be determined that the address the user wants to go to is Address A, and Address A is the target address. Then the target operation instruction "Navigate to Address A" can be determined according to Address A.

[0296] In a possible embodiment, after receiving multiple candidate address information (Address A, Address B, Address C), the terminal displays these candidate address information on the screen. When the user clicks on "Address A" or circles "Address A" on the touch screen, after receiving the operation instruction, the terminal can determine that the address the user wants to go to is Address A, and Address A is the target address. Then the target operation instruction "Navigate to Address A" can be determined according to Address A.

[0297] Optionally, it further includes:

[0298] Output at least one candidate call object information through the predetermined model, and the candidate call object information is related to the information to be recognized and the first operation information;

[0299] Receive a second user instruction, and confirm at least one of the candidate call object information based on the second user instruction to obtain a target operation instruction.

[0300] In this embodiment, the execution subject of the step can be the first terminal or the second terminal. In a call scenario, obtain the information to be recognized and the first operation information in the first terminal, and process the information to be recognized and the first operation information to extract the corresponding call object from the information to be recognized according to the intention in the first operation information. Since there are multiple call objects in the information to be recognized, multiple candidate call object information is generated after processing, and multiple candidate call object information is returned to the first terminal.

[0301] Since multiple candidate call object information is received, the terminal still cannot confirm an exact call object. To generate an exact target operation instruction, the terminal displays these candidate call object information to the user for selection. The user inputs a second user instruction, and the terminal receives the second user instruction to select one or more from the candidate call object information to generate a target operation instruction. By selecting the candidate call object information through the second user instruction and determining the target operation instruction, the user's intention can be accurately known, the user's instruction can be accurately executed, and the user experience is improved.

[0302] Optionally, the call object information can be an object name or a phone number, and the object name can be "Teacher A", "Teacher B", etc. in the address book.

[0303] In a possible embodiment, the information to be recognized is a screenshot of a first application, and the screenshot contains a call object. After processing the information to be recognized and the first operation information, only one candidate call object information "Teacher A" is generated. At this time, the user's destination is determined. Then, after receiving it, the terminal can directly generate a target operation instruction "Call Teacher A" based on this candidate call object information and the first operation information.

[0304] In a possible embodiment, the information to be recognized is a screenshot of a first application, and the screenshot contains multiple call objects. After processing the information to be recognized and the first operation information, multiple candidate call object information (Call Object A, Call Object B, Call Object C) is generated. After receiving these candidate call object information, the terminal needs to further confirm with the user which call object the user wants to call. The terminal will display these candidate call object information on the display, or, announce the candidate call object information by voice, to prompt the user to select from these call objects. The user inputs a second user instruction "Call Call Object A", then the terminal can determine that the target operation instruction is "Contact Call Object A".

[0305] Optionally, the receiving the second user instruction and confirming at least one of the candidate call object information based on the second user instruction to obtain a target operation instruction includes at least one of the following methods:

[0306] Receiving a voice instruction from the user, selecting a target call object from at least one of the candidate call objects, and obtaining the target operation instruction, where the target operation instruction includes the target call object;

[0307] Receiving an operation instruction input by the user through a click operation or a selection operation, determining a target call object from at least one of the candidate call objects, and obtaining the target operation instruction, where the target operation instruction includes the target call object.

[0308] In this embodiment, the execution subject of the step may be a first terminal or a second terminal. There are various ways for the user to input the second user instruction. The user can input a voice instruction by directly speaking to the terminal, or can manually operate on the terminal. The input voice instruction contains a selection decision for the candidate call object. After receiving multiple candidate call object information, the terminal displays multiple candidate call object information on the display screen. The user can click on the screen to select the candidate call object information, or, circle the candidate call object information on the screen, or, select the candidate call object information through a key operation. By obtaining the second user instruction in multiple ways, the user can input the instruction conveniently, and the interaction efficiency between the user and the terminal is higher, improving the user experience.

[0309] In a possible embodiment, after receiving multiple candidate call object information (call object A, call object B, call object C), the terminal displays this candidate call object information on the screen. When the user says "Contact the first call object" or "Call call object A", after receiving the voice command, the terminal converts the voice command into text and matches it with the candidate call object information, and then it can be determined that the user wants to contact call object A, and call object A is the target call object. Then the target operation command "Contact call object A" can be determined according to call object A.

[0310] In a possible embodiment, after receiving multiple candidate call object information (call object A, call object B, call object C), the terminal displays this candidate call object information on the screen. When the user clicks on "Call object A" or circles "Call object A" on the touch screen, after receiving the operation command, the terminal can determine that the user wants to contact call object A, and call object A is the target call object. Then the target operation command "Contact call object A" can be determined according to call object A.

[0311] Optionally, the predetermined model is a multimodal large language model trained with training data containing information to be recognized and operation information.

[0312] In this embodiment, the multimodal large language model (MLLMs) is a deep learning model that combines a large language model (LLM) and a large vision model (LVM), and can process and understand various types of data, such as text, images, and audio.

[0313] The multimodal large language model can process various types of input data, such as text, images, audio, video, etc., and can generate outputs in multiple modalities, such as generating images according to text, or generating descriptions according to images. During the training process of the multimodal large language model, it is trained on a large-scale dataset containing text, vision, audition, and sometimes even sensing data, and can establish connections between different modalities, thus supporting tasks that require understanding and generating content across multiple data types.

[0314] In a possible embodiment, by collecting the historical data of the interaction between the user and the terminal, the information to be recognized and the operation information therein are obtained to generate training data. Among them, the operation information is text or voice data, and the information to be recognized is data in modalities such as text, images, audio, and video. A sample in the training data contains information to be recognized, operation information, and label data, and the label data is the matching result corresponding to the information to be recognized and the operation information in the sample, such as candidate address information or candidate target object information.

[0315] Each sample in the training data is input into the multi-modal large language model for processing, and then a matching result is output. The matching result is compared with the label data corresponding to the sample, and a loss function representing the difference size is calculated. Taking reducing the loss function as the goal, the parameters in the multi-modal large language model are optimized. After multiple rounds of optimization iterations on the parameters, a trained predetermined model can be obtained.

[0316] Processing the multi-modal large language model with the information to be recognized and the operation information can improve the multi-modal large language model's understanding ability of various modal data, so as to better process information in cross-modal tasks, improve the efficiency and accuracy of matching, and the matching results obtained by users are more accurate, enhancing the user experience.

[0317] Optionally, the method further includes:

[0318] Obtain a first wake-up instruction, where the first wake-up instruction is triggered by an operation on the first terminal;

[0319] In response to the first wake-up instruction, collect the user's input information and obtain the first user instruction.

[0320] In this embodiment, before the user inputs the first user instruction, the module or application responsible for recognizing speech in the first terminal can be woken up by the first wake-up instruction, so that the first terminal starts to receive voice or text instructions. After receiving the first wake-up instruction, the first terminal learns that the user intends to input an instruction, and then starts to call the modules or applications therein, collect the user's input information, and obtain the first user instruction.

[0321] In a possible embodiment, the voice assistant in the first terminal collects and analyzes the user's instructions. After the user says the first wake-up instruction of "Little X classmate", the first terminal starts to wake up the voice assistant application and is ready to receive the user's voice instructions.

[0322] Optionally, the obtaining the information to be recognized of the first application in the first terminal includes:

[0323] Obtain the image data or text data in the first terminal as the information to be recognized.

[0324] In this embodiment, the instruction that the user wants to give is related to the information displayed by the first application in the first terminal. The first terminal can extract the screenshot of the application or obtain the text in the application from the background as the information to be recognized.

[0325] In a possible embodiment, in the content displayed in a first application (a certain blog APP) on a first terminal (such as a mobile phone), the user sees a restaurant and wants to go to this restaurant for dinner. Then, the user can input a first user instruction "I want to go to the restaurant in the picture" by voice. The first terminal captures an image of the first application or obtains the text in the first application as data to be recognized.

[0326] Optionally, the obtaining of the image data or text data in the first terminal as the information to be recognized includes at least one of the following methods:

[0327] Capturing the image displayed on the first terminal to obtain the information to be recognized;

[0328] Obtaining the image of the area specified by the user in the first terminal as the information to be recognized;

[0329] Obtaining the image collected by the first terminal as the information to be recognized.

[0330] In this embodiment, the method for the first terminal to obtain the image data of the first application includes various ways. It can capture the entire screen, or obtain the image of the area selected by the user, or schedule the sensor in the first terminal to collect the image.

[0331] In a possible embodiment, in the content displayed in a first application (a certain blog APP) on a first terminal (such as a mobile phone), the user sees a restaurant and wants to go to this restaurant for dinner. Then, the user can input a first user instruction "I want to go to the restaurant in the picture" by voice. The first terminal captures an image of the first application or obtains the text in the first application as data to be recognized, or determines the area circled by the user and captures the image of the area circled by the user. By specifying the information to be recognized in the first terminal in various ways, the user can input instructions conveniently, and the interaction efficiency between the user and the terminal is higher, improving the user experience.

[0332] To implement the above embodiment, an interaction device is further proposed in an embodiment of the present application.

[0333] Figure 4 It is a schematic structural diagram of an interaction device provided in an embodiment of the present application.

[0334] As Figure 4 shown, the device may include:

[0335] A transceiver module 410, configured to obtain a first user instruction, where the first user instruction includes first operation information;

[0336] A processing module 420, configured to obtain the information to be recognized of a first application in a first terminal, so as to determine a target operation instruction associated with the information to be recognized and the first operation information; the target operation instruction is used to trigger a second application in the second terminal to execute a target task, where the first terminal and the second terminal establish a network connection through a device identifier.

[0337] Optionally, it further includes:

[0338] An instruction determination module, configured to determine a target operation instruction associated with the information to be recognized and the first operation information based on the information to be recognized and the first operation information.

[0339] Optionally, the instruction determination module includes:

[0340] A matching module, configured to use the first operation information and the information to be recognized as input parameters and input them into a predetermined model to obtain a matching result; where the matching result is related to the target operation instruction.

[0341] Optionally, the device further includes:

[0342] An address receiving module, configured to receive at least one candidate address information, where the candidate address information is related to the information to be recognized and the first operation information;

[0343] A second instruction receiving module, configured to receive a second user instruction and obtain the target operation instruction based on at least one of the candidate address information.

[0344] Optionally, the second instruction receiving module includes:

[0345] A voice receiving module, configured to receive a user's voice instruction, select a target address from at least one of the candidate addresses, and obtain the target operation instruction, where the target operation instruction includes the target address;

[0346] An operation receiving module, configured to receive an operation instruction input by a user through a click operation or a selection operation, determine a target address from at least one of the candidate addresses, and obtain the target operation instruction, where the target operation instruction includes the target address.

[0347] Optionally, it further includes:

[0348] An object receiving module, configured to receive at least one candidate call object information, where the candidate call object information is related to the information to be recognized and the first operation information;

[0349] A second instruction receiving module, configured to receive a second user instruction and obtain a target operation instruction based on at least one of the candidate call object information.

[0350] Optionally, the second instruction receiving module includes:

[0351] A voice receiving module, configured to receive a voice instruction from a user, select a target call object from at least one of the candidate call objects, and obtain the target operation instruction, where the target operation instruction includes the target call object;

[0352] An operation receiving module, configured to receive an operation instruction input by the user through a click operation or a selection operation, determine a target call object from at least one of the candidate call objects, and obtain the target operation instruction, where the target operation instruction includes the target call object.

[0353] Optionally, the predetermined model is a multimodal large language model trained with training data including information to be recognized and operation information.

[0354] Optionally, the device further includes:

[0355] A wake-up module, configured to obtain a first wake-up instruction, where the first wake-up instruction is triggered by an operation on the first terminal;

[0356] An acquisition module, configured to collect input information of the user and obtain the first user instruction in response to the first wake-up instruction.

[0357] Optionally, the processing module 420 includes:

[0358] An information acquisition module, configured to obtain image data or text data in the first terminal as the information to be recognized.

[0359] Optionally, the information acquisition module includes:

[0360] A first acquisition sub-module, configured to intercept an image displayed on the first terminal to obtain the information to be recognized;

[0361] A second acquisition sub-module, configured to obtain an image of a specified area in the first terminal by the user as the information to be recognized;

[0362] A third acquisition sub-module, configured to obtain an image collected by the first terminal as the information to be recognized.

[0363] Optionally, the device further includes:

[0364] A result acquisition module, configured to receive result information sent by a second terminal, where the result information is related to the execution progress or execution result of the target task.

[0365] It should be noted that the foregoing explanations of the method embodiments also apply to the device of this embodiment, and will not be elaborated here.

[0366] To implement the above embodiments, an embodiment of the present application further provides an interaction device.

[0367] Figure 5 FIG. is a schematic structural diagram of an interaction device provided by an embodiment of the present application.

[0368] As Figure 5 shown, the device may include:

[0369] A transceiver module 510, configured to receive a target operation instruction, where the target operation instruction is used to instruct a second terminal to execute a target task; the target operation instruction is related to first operation information and information to be recognized of a first terminal, and the first terminal and the second terminal establish a network connection through a device identifier;

[0370] An execution module 520, configured to trigger a second application to execute a target task according to the target operation instruction.

[0371] Optionally, the device further includes:

[0372] An address receiving module, configured to receive at least one candidate address information, where the candidate address information is related to the information to be recognized and the first operation information;

[0373] A second instruction receiving module, configured to receive a second user instruction, and confirm at least one of the candidate address information based on the second user instruction to obtain the target operation instruction.

[0374] Optionally, the second instruction receiving module includes:

[0375] A voice receiving module, configured to receive a user's voice instruction, select a target address from at least one of the candidate addresses, and obtain the target operation instruction, where the target operation instruction includes the target address;

[0376] An operation receiving module, configured to receive an operation instruction input by a user through a click operation or a selection operation, determine a target address from at least one of the candidate addresses, and obtain the target operation instruction, where the target operation instruction includes the target address.

[0377] Optionally, the device further includes:

[0378] A target address receiving module, configured to receive a target address sent by a first terminal to obtain the target operation instruction, where the target operation instruction includes the target address.

[0379] Optionally, the device further includes:

[0380] An object receiving module, configured to receive at least one candidate call object information, where the candidate call object information is related to the information to be recognized and the first operation information;

[0381] A second instruction receiving module, configured to receive a second user instruction, and confirm at least one piece of the candidate call object information based on the second user instruction to obtain a target operation instruction.

[0382] Optionally, the second instruction receiving module includes:

[0383] A voice receiving module, configured to receive a user's voice instruction, select a target call object from at least one of the candidate call objects, and obtain the target operation instruction, where the target operation instruction includes the target call object;

[0384] An operation receiving module, configured to receive an operation instruction input by a user through a click operation or a selection operation, determine a target call object from at least one of the candidate call objects, and obtain the target operation instruction, where the target operation instruction includes the target call object.

[0385] Optionally, the apparatus further includes:

[0386] A target object receiving module, configured to receive a target call object sent by a first terminal to obtain the target operation instruction, where the target operation instruction includes the target call object.

[0387] Optionally, the apparatus further includes:

[0388] A first result generating module, configured to generate result information, where the result information is related to the execution progress or execution result of the target task; or,

[0389] A second result generating module, configured to generate result information and send the result information to the first terminal, where the result information is related to the execution progress or execution result of the target task.

[0390] It should be noted that the foregoing explanation of the method embodiment is also applicable to the apparatus of this embodiment, and will not be elaborated herein.

[0391] To implement the foregoing embodiments, an interaction apparatus is further proposed in an embodiment of the present application.

[0392] Figure 6 It is a schematic structural diagram of an interaction apparatus provided in an embodiment of the present application.

[0393] As Figure 6 shown, the apparatus may include:

[0394] A transceiver module 610, configured to obtain a first user instruction, where the first user instruction includes first operation information; obtain information to be recognized of a first application in a first terminal;

[0395] A processing module 620, configured to determine a target operation instruction based on the information to be recognized and the first operation information;

[0396] An execution module 630, configured to trigger a second application of the second terminal to execute a target task based on the target operation instruction, where the first terminal and the second terminal establish a network connection through a device identifier.

[0397] Optionally, the processing module includes:

[0398] A matching module, configured to use the first operation information and the information to be recognized as input parameters, and input them into a predetermined model to obtain a matching result; where the matching result is related to the target operation instruction.

[0399] Optionally, it further includes:

[0400] An address acquisition module, configured to output at least one candidate address information through the predetermined model, where the candidate address information is related to the information to be recognized and the first operation information;

[0401] A second instruction acquisition module, configured to receive a second user instruction, and confirm at least one of the candidate address information based on the second user instruction to obtain a target operation instruction.

[0402] Optionally, the second instruction acquisition module includes:

[0403] A voice acquisition module, configured to receive a user's voice instruction, select a target address from at least one of the candidate addresses to obtain the target operation instruction, where the target operation instruction includes the target address;

[0404] An operation acquisition module, configured to receive an operation instruction input by a user through a click operation or a selection operation, determine a target address from at least one of the candidate addresses to obtain the target operation instruction, where the target operation instruction includes the target address.

[0405] Optionally, it further includes:

[0406] An object acquisition module, configured to output at least one candidate call object information through the predetermined model, where the candidate call object information is related to the information to be recognized and the first operation information;

[0407] A second instruction acquisition module, configured to receive a second user instruction, and confirm at least one of the candidate call object information based on the second user instruction to obtain a target operation instruction.

[0408] Optionally, the second instruction acquisition module includes:

[0409] A voice receiving module, configured to receive a user's voice instruction, select a target call object from at least one of the candidate call objects, and obtain the target operation instruction, where the target operation instruction includes the target call object;

[0410] An operation receiving module, configured to receive an operation instruction input by the user through a click operation or a selection operation, determine a target call object from at least one of the candidate call objects, and obtain a target operation instruction, where the target operation instruction includes the target call object.

[0411] Optionally, the predetermined model is a multimodal large language model trained with training data including information to be recognized and operation information.

[0412] Optionally, the device further includes:

[0413] A wake-up module, configured to obtain a first wake-up instruction, where the first wake-up instruction is triggered by an operation on the first terminal;

[0414] An instruction collection module, configured to collect the user's input information in response to the first wake-up instruction and obtain the first user instruction.

[0415] Optionally, the transceiver module 610 includes:

[0416] An information-to-be-recognized acquisition module, configured to obtain image data or text data in the first terminal as the information to be recognized.

[0417] Optionally, the information-to-be-recognized acquisition module includes:

[0418] A first acquisition sub-module, configured to intercept an image displayed on the first terminal to obtain the information to be recognized;

[0419] A second acquisition sub-module, configured to obtain an image of a specified area in the first terminal specified by the user as the information to be recognized;

[0420] A third acquisition sub-module, configured to obtain an image collected by the first terminal as the information to be recognized.

[0421] It should be noted that the foregoing explanation of the method embodiment also applies to the device of this embodiment, and details are not described herein again.

[0422] To implement the above embodiments, the present application also proposes a non-transitory computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the method described in the foregoing method embodiment is implemented.

[0423] To implement the above embodiments, the present application also provides a computer program product, having a computer program stored thereon, where the computer program, when executed by a processor, implements the method described in the foregoing method embodiments.

[0424] To implement the above embodiments, the present application also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where, when the processor executes the program, the method described in the foregoing method embodiments is implemented.

[0425] Figure 7 FIG. is a schematic structural diagram of an electronic device provided by an embodiment of the present application. For example, the electronic device 800 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0426] Referring to Figure 7 , the electronic device 800 may include one or more of the following components: a processing component 802, a memory 804, a power component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 812, a sensor component 814, and a communication component 816.

[0427] The processing component 802 generally controls the overall operation of the electronic device 800, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the above method. In addition, the processing component 802 may include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.

[0428] The memory 804 is configured to store various types of data to support the operation of the electronic device 800. Examples of such data include instructions for any application or method operating on the electronic device 800, contact data, phone book data, messages, pictures, videos, etc. The memory 804 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc.

[0429] The power component 806 provides power for various components of the electronic device 800. The power component 806 may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power for the electronic device 800.

[0430] The multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of the touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operations. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the electronic device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.

[0431] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC) that is configured to receive external audio signals when the electronic device 800 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 further includes a speaker for outputting audio signals.

[0432] The I / O interface 812 provides an interface between the processing component 802 and a peripheral interface module, which can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a power button, and a lock button.

[0433] The sensor assembly 814 includes one or more sensors for providing a status assessment of various aspects for the electronic device 800. For example, the sensor assembly 814 can detect the on / off state of the electronic device 800, the relative positioning of components, such as the display and keypad of the electronic device 800. The sensor assembly 814 can also detect a change in the position of the electronic device 800 or a component of the electronic device 800, the presence or absence of user contact with the electronic device 800, the orientation or acceleration / deceleration of the electronic device 800, and the temperature change of the electronic device 800. The sensor assembly 814 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 814 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 814 can also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0434] The communication component 816 is configured to facilitate communication between the electronic device 800 and other devices in a wired or wireless manner. The electronic device 800 can access a wireless network based on communication standards, such as WiFi, 4G, or 5G, or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0435] In an exemplary embodiment, the electronic device 800 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above method.

[0436] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions, and the above instructions can be executed by a processor 820 of the electronic device 800 to complete the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0437] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this application. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0438] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of this application, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically defined.

[0439] Any process or method description shown in the flowchart or described in other ways herein can be understood as representing a module, segment, or part of code including one or more executable instructions for implementing a customized logic function or process, and the scope of the preferred embodiments of this application includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in a reverse order according to the involved functions, rather than in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of this application belong.

[0440] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definitional sequence list of executable instructions for implementing logical functions, which can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in conjunction with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: electrical connection parts with one or more wirings (electronic devices), portable computer disk cartridges (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber devices, and portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other suitable processing as necessary, and then stored in a computer memory.

[0441] It should be understood that various parts of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits with suitable combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0442] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the method of the above embodiments can be completed by a program instructing relevant hardware, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.

[0443] In addition, each functional unit in various embodiments of the present application may be integrated into one processing module, may exist physically alone for each unit, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.

[0444] The above-mentioned storage medium may be a read-only memory, a magnetic disk, an optical disc, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.

Claims

1. An interaction method, characterized in that, Including: Obtain a first user instruction, where the first user instruction includes first operation information; Obtain the information to be recognized of a first application in the first terminal to determine a target operation instruction associated with the information to be recognized and the first operation information; the target operation instruction is used to trigger a second application in the second terminal to execute a target task, where the first terminal and the second terminal establish a network connection through device identifiers.

2. The method according to claim 1, wherein It further includes: Based on the information to be recognized and the first operation information, determine a target operation instruction associated with the information to be recognized and the first operation information.

3. The method according to claim 2, wherein Based on the information to be recognized and the first operation information, determining a target operation instruction associated with the information to be recognized and the first operation information includes: Taking the first operation information and the information to be recognized as input parameters and inputting them into a predetermined model to obtain a matching result; where the matching result is related to the target operation instruction.

4. The method according to claim 2 or 3, characterized in that, It further includes: Receive at least one candidate address information, where the candidate address information is related to the information to be recognized and the first operation information; Receive a second user instruction and obtain the target operation instruction based on at least one of the candidate address information.

5. The method according to claim 4, characterized in that, Receiving a second user instruction and obtaining a target operation instruction based on at least one of the candidate address information includes at least one of the following methods: Receive the user's voice instruction, select a target address from at least one of the candidate addresses to obtain the target operation instruction, and the target operation instruction includes the target address; Receive the operation instruction input by the user through a click operation or a selection operation, determine the target address from at least one of the candidate addresses to obtain the target operation instruction, and the target operation instruction includes the target address.

6. The method according to claim 2 or 3, characterized in that, It further includes: Receive at least one candidate call object information, where the candidate call object information is related to the information to be recognized and the first operation information; Receive a second user instruction and obtain a target operation instruction based on at least one of the candidate call object information.

7. The method according to claim 6, wherein Receiving a second user instruction and obtaining a target operation instruction based on at least one of the candidate call object information includes at least one of the following methods: Receive the user's voice instruction, select a target call object from at least one of the candidate call objects to obtain the target operation instruction, and the target operation instruction includes the target call object; Receive the operation instruction input by the user through a click operation or a selection operation, determine the target call object from at least one of the candidate call objects to obtain the target operation instruction, and the target operation instruction includes the target call object.

8. The method according to claim 3, wherein The predetermined model is a multi-modal large language model trained with training data including information to be recognized and operation information.

9. The method according to claim 1, wherein The method further includes: Obtain a first wake-up instruction, where the first wake-up instruction is triggered by an operation on the first terminal; In response to the first wake-up instruction, collect the user's input information and obtain the first user instruction.

10. The method according to claim 1, wherein The obtaining the information to be recognized of a first application in the first terminal includes: Obtain the image data or text data in the first terminal as the information to be recognized.

11. The method according to claim 10, wherein Obtaining the image data or text data in the first terminal as the information to be recognized includes at least one of the following methods: Capturing the image displayed on the first terminal to obtain the information to be recognized; Obtaining the image of the area specified by the user in the first terminal as the information to be recognized; Obtaining the image collected by the first terminal as the information to be recognized.

12. The method according to claim 1, wherein The method further includes: Receiving the result information sent by the second terminal, where the result information is related to the execution progress or execution result of the target task.

13. An interaction method, characterized in that, Including: Receiving a target operation instruction, where the target operation instruction is used to instruct the second terminal to execute a target task; the target operation instruction is related to the first operation information and the information to be recognized of the first terminal, and the first terminal and the second terminal establish a network connection through the device identifier; Triggering the second application to execute the target task according to the target operation instruction.

14. The method according to claim 13, wherein The method further includes: Receiving at least one candidate address information, where the candidate address information is related to the information to be recognized and the first operation information; Receiving a second user instruction, and based on the second user instruction, confirming at least one of the candidate address information to obtain the target operation instruction.

15. The method according to claim 14, wherein Receiving a second user instruction, and based on the second user instruction, confirming at least one of the candidate address information to obtain the target operation instruction includes at least one of the following methods: Receiving the voice instruction of the user, selecting a target address from at least one of the candidate addresses to obtain the target operation instruction, and the target operation instruction includes the target address; Receiving the operation instruction input by the user through a click operation or a selection operation, determining the target address from at least one of the candidate addresses to obtain the target operation instruction, and the target operation instruction includes the target address.

16. The method according to claim 13, wherein The method further includes: Receiving the target address sent by the first terminal to obtain the target operation instruction, and the target operation instruction includes the target address.

17. The method according to claim 13, wherein The method further includes: Receiving at least one candidate call object information, where the candidate call object information is related to the information to be recognized and the first operation information; Receiving a second user instruction, and based on the second user instruction, confirming at least one of the candidate call object information to obtain the target operation instruction.

18. The method according to claim 17, wherein Receiving a second user instruction, and based on the second user instruction, confirming at least one of the candidate call object information to obtain the target operation instruction includes at least one of the following methods: Receiving the voice instruction of the user, selecting a target call object from at least one of the candidate call objects to obtain the target operation instruction, and the target operation instruction includes the target call object; Receiving the operation instruction input by the user through a click operation or a selection operation, determining the target call object from at least one of the candidate call objects to obtain the target operation instruction, and the target operation instruction includes the target call object.

19. The method according to claim 13, wherein The method further includes: Receiving the target call object sent by the first terminal to obtain the target operation instruction, and the target operation instruction includes the target call object.

20. The method according to claim 13, wherein The method further includes: Generate result information, where the result information is related to the execution progress or result of the target task; or, Generate result information and send the result information to the first terminal, where the result information is related to the execution progress or result of the target task.

21. An interaction method, characterized in that Including: Obtain a first user instruction, where the first user instruction includes first operation information; Obtain the information to be recognized of the first application in the first terminal; Based on the information to be recognized and the first operation information, determine a target operation instruction; Based on the target operation instruction, trigger the second application in the second terminal to execute the target task, where the first terminal and the second terminal establish a network connection through device identifiers.

22. The method according to claim 21, wherein Based on the information to be recognized and the first operation information, determining a target operation instruction includes: Use the first operation information and the information to be recognized as input parameters and input them into a predetermined model to obtain a matching result; where the matching result is related to the target operation instruction.

23. The method according to claim 22, wherein Further including: Output at least one candidate address information through the predetermined model, where the candidate address information is related to the information to be recognized and the first operation information; Receive a second user instruction, and confirm at least one of the candidate address information based on the second user instruction to obtain a target operation instruction.

24. The method according to claim 23, wherein Receiving a second user instruction and confirming at least one of the candidate address information based on the second user instruction to obtain a target operation instruction includes at least one of the following methods: Receive the user's voice instruction, select a target address from at least one of the candidate addresses to obtain the target operation instruction, and the target operation instruction includes the target address; Receive the operation instruction input by the user through a click operation or a selection operation, determine the target address from at least one of the candidate addresses to obtain the target operation instruction, and the target operation instruction includes the target address.

25. The method according to claim 22, wherein Further including: Output at least one candidate call object information through the predetermined model, where the candidate call object information is related to the information to be recognized and the first operation information; Receive a second user instruction and confirm at least one of the candidate call object information based on the second user instruction to obtain a target operation instruction.

26. The method according to claim 25, wherein Receiving a second user instruction and confirming at least one of the candidate call object information based on the second user instruction to obtain a target operation instruction includes at least one of the following methods: Receive the user's voice instruction, select a target call object from at least one of the candidate call objects to obtain the target operation instruction, and the target operation instruction includes the target call object; Receive the operation instruction input by the user through a click operation or a selection operation, determine the target call object from at least one of the candidate call objects to obtain the target operation instruction, and the target operation instruction includes the target call object.

27. The method according to claim 22, wherein The predetermined model is a multimodal large language model trained with training data including information to be recognized and operation information.

28. The method according to claim 21, wherein The method further includes: Obtain a first wake-up instruction, where the first wake-up instruction is triggered by an operation on the first terminal; In response to the first wake-up instruction, collect the input information of the user and obtain the first user instruction.

29. The method according to claim 21, wherein The obtaining of the information to be recognized of the first application in the first terminal includes: Obtain the image data or text data in the first terminal as the information to be recognized.

30. The method according to claim 29, wherein The obtaining of the image data or text data in the first terminal as the information to be recognized includes at least one of the following methods: Capture the image displayed on the first terminal to obtain the information to be recognized; Obtain the area image specified by the user in the first terminal as the information to be recognized; Obtain the image collected by the first terminal as the information to be recognized.

31. An interactive device, characterized in that, The device is used to execute the method according to any one of claims 1-12, and includes: A transceiver module, configured to obtain a first user instruction, wherein the first user instruction includes first operation information; A processing module, configured to obtain the information to be recognized of the first application in the first terminal to determine a target operation instruction associated with the information to be recognized and the first operation information; the target operation instruction is used to trigger the second application of the second terminal to execute a target task, wherein the first terminal and the second terminal establish a network connection through a device identifier.

32. An interactive device, characterized in that, The device is used to execute the method according to any one of claims 13-20, and includes: A transceiver module, configured to receive a target operation instruction, wherein the target operation instruction is used to instruct the second terminal to execute a target task; the target operation instruction is related to the first operation information and the information to be recognized of the first terminal, and the first terminal and the second terminal establish a network connection through a device identifier; An execution module, configured to trigger the second application to execute the target task according to the target operation instruction.

33. An interactive device, characterized in that, The device is used to execute the method according to any one of claims 21-30, and includes: A transceiver module, configured to obtain a first user instruction, wherein the first user instruction includes first operation information; obtain the information to be recognized of the first application in the first terminal; A processing module, configured to determine a target operation instruction based on the information to be recognized and the first operation information; An execution module, configured to trigger the second application of the second terminal to execute a target task based on the target operation instruction, wherein the first terminal and the second terminal establish a network connection through a device identifier.

34. A first terminal, characterized in that It includes at least one processor, and the at least one processor is coupled to a memory; the at least one processor is used to execute the method according to any one of claims 1-12.

35. A second terminal, characterized in that It includes at least one processor, and the at least one processor is coupled to a memory; the at least one processor is used to execute the method according to any one of claims 13-20.

36. A system, characterized in that, The system includes a first terminal and a second terminal, The system is used to execute the method according to any one of claims 21-30; or, The first terminal is used to execute the method according to any one of claims 1-12; or, The second terminal is used to execute the method according to any one of claims 13-20.

37. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method according to any one of claims 1-12, 13-20 or 21-30.

38. A computer program product, characterized in that, When the program is executed by a processor, it implements the method according to claims 1-12, 13-20 or 21-30.

Citation Information

Cited By

  • Interaction method, system and device, vehicle and storage medium

    CN121262179A