Method and device for interacting with user

By using a decision model in a computing device to obtain user input and call specific functions, the problem of poor interaction experience in the prior art is solved, and more accurate user intention recognition and better interactive experience are achieved.

CN120066274APending Publication Date: 2025-05-30ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510222721.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In the prior art, when computing devices interact with users, they directly input and output dialogue content based on the user's voice, which often fails to meet user needs, resulting in poor interaction experience.

Method used

The user's input data is obtained through the decision model, and the decision results are output to indicate the call to specific functions in the computing device, such as hardware control, API calls or interactive language model calls, and then function calls are made to better identify user intent.

Benefits of technology

Through decision-making and function calls of the decision-making model, the output dialogue can better meet user needs and improve the interactive experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066274A_ABST
    Figure CN120066274A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a method and device for interacting with a user. The method comprises the steps that the computing device obtains first input data of a user; outputting a first decision result based on the first input data through a decision model, the first decision result being used for indicating to call a first function provided in the computing device; and calling the first function according to the first decision result. The interaction experience can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One or more embodiments of this specification relate to the field of computers, and in particular, to methods and devices for interacting with users. Background Art

[0002] Currently, computing devices interact with users to better provide services to users, and the interaction is often reflected through conversations. For example, when the user's voice input is "What kind of flower is this?", the computing device not only needs to understand the user's intention, but also obtain an image of the flower, and then identify the name of the flower in the image before it can output a conversation that meets the user's intention.

[0003] In the prior art, during the interaction between a computing device and a user, directly outputting conversation content based on the user's voice input often results in the output conversation not meeting the user's needs, thus resulting in a poor interaction experience. Summary of the Invention

[0004] One or more embodiments of this specification describe a method and device for interacting with users, which can improve the interaction experience.

[0005] In a first aspect, a method for interacting with a user is provided. The method is executed by a computing device and includes:

[0006] Obtain first input data of the user;

[0007] Based on the first input data, output a first decision result through a decision model, where the first decision result is used to indicate calling a first function provided in the computing device;

[0008] According to the first decision result, call the first function.

[0009] In a possible implementation manner, the first input data includes at least one of the following data: voice, image, video for indicating a gesture.

[0010] In a possible implementation manner, the first function includes any one of the following functions:

[0011] Function for controlling hardware, function for calling an application programming interface (API), function for calling an interaction language model.

[0012] In a possible implementation manner, the outputting a first decision result through a decision model based on the first input data includes:

[0013] Obtain a first feature block based on the first input data, and add the first feature block to a first feature block sequence;

[0014] Based on the first feature block and a second feature block in the first feature block sequence that is before the first feature block, a decision model outputs a first decision result. The second feature block is obtained based on second input data, and the second input data is data input by the user before the first input data.

[0015] Further, obtaining the first feature block based on the first input data includes: obtaining a plurality of first feature blocks arranged in order based on the first input data;

[0016] The decision model outputs the first decision result based on the first feature block and a second feature block in the first feature block sequence that is before the first feature block, including:

[0017] Performing attention calculation on each of the plurality of first feature blocks based on the plurality of first feature blocks and a plurality of second feature blocks arranged in order before the plurality of first feature blocks in the first feature block sequence, obtaining a plurality of third feature blocks arranged in order corresponding to the plurality of first feature blocks respectively, and sequentially adding the plurality of third feature blocks to the second feature block sequence;

[0018] For each third feature block, the decision model outputs a first decision result corresponding to the third feature block based on the third feature block and a plurality of fourth feature blocks arranged in order before the third feature block in the second feature block sequence.

[0019] In a possible implementation manner, the method further includes:

[0020] Based on third input data, a decision model outputs a second decision result, and the second decision result is used to indicate that based on the result data of calling a first function, a second function provided in the computing device is called. The second function is a function of calling an interactive language model, and the third input data is data input by the user after the first input data.

[0021] Further, the third input data is initial input information of the user obtained in a streaming manner at the current moment; the result data of the first function is first supplementary input information obtained by executing a first hardware control instruction and / or a first application programming interface (API) call instruction;

[0022] Calling the second function provided in the computing device includes: using the initial input information and the first supplementary input information as input data of the interactive language model.

[0023] Further, the initial input information includes active input information and passive input information.

[0024] In a possible implementation, the computing device includes an interactive wearable device.

[0025] In a possible implementation, the decision model is trained in the following manner:

[0026] Obtain training samples, where the training samples include: model input data and label instruction sequences;

[0027] Input the model input data into the decision model, and output prediction instruction sequences corresponding to each moment respectively. The prediction instruction sequences include predicted hardware control instructions, application programming interface (API) call instructions, and / or model call instructions;

[0028] Train the decision model according to the differences between the prediction instruction sequences and the label instruction sequences.

[0029] Further, the hardware control instructions are used to control the state of the hardware. Controlling the state of the hardware includes:

[0030] Controlling a music player to play music; or,

[0031] Taking a photo control; or,

[0032] Volume control; or,

[0033] Display control.

[0034] Further, the API call instructions are used to call the APIs of target applications. The target applications include:

[0035] Retrieval-augmented generation (RAG) applications, object recognition applications by scanning, or weather information acquisition applications.

[0036] In a possible implementation, the first input data includes active input information and environmental perception information. The environmental perception information includes background noise extracted from the active input information in the speech modality, or background images extracted from the active input information in the picture modality.

[0037] Further, the first input data further includes human body perception information, and the human body perception information includes:

[0038] Eye movement coordinates or emotion labels.

[0039] In a second aspect, a device for interacting with a user is provided. The device is disposed in a computing device and includes:

[0040] An acquisition unit for acquiring the first input data of the user;

[0041] A decision-making unit, configured to output a first decision result based on the first input data obtained by the obtaining unit through a decision-making model, where the first decision result is used to indicate to invoke a first function provided in the computing device;

[0042] An invocation unit, configured to invoke the first function according to the first decision result obtained by the decision-making unit.

[0043] In a third aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed on a computer, the computer is made to execute the method of the first aspect.

[0044] In a fourth aspect, a computing device is provided, including a memory and a processor. An executable code is stored in the memory. When the processor executes the executable code, the method of the first aspect is implemented.

[0045] Through the method and device provided in the embodiments of this specification, the computing device first obtains the first input data of the user; then, through a decision-making model, based on the first input data, outputs a first decision result, where the first decision result is used to indicate to invoke a first function provided in the computing device; and then, according to the first decision result, invokes the first function. As can be seen from the above, in the embodiments of this specification, during the interaction process between the computing device and the user, instead of directly outputting the conversation content based on the user's input data, it first makes a decision through a decision-making model to invoke a specific function provided in the computing device. On the one hand, it makes real-time control decisions by perceiving the input. On the other hand, through function invocation, it can undertake a wider range of applications, and the function invocation result helps to better identify the user's intention. In summary, the output conversation can better meet the user's needs, thereby improving the interaction experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0047] Figure 1 It is a schematic diagram of an implementation scenario of an embodiment disclosed in this specification;

[0048] Figure 2 It shows a flowchart of a method for interacting with a user according to an embodiment;

[0049] Figure 3 It shows a schematic diagram of the streaming process of a decision-making model according to an embodiment;

[0050] Figure 4 Schematic diagram of inference acceleration of a decision model according to an embodiment;

[0051] Figure 5 Schematic diagram of construction of training samples according to an embodiment;

[0052] Figure 6 Schematic block diagram of a device for interacting with a user according to an embodiment;

[0053] Figure 7 Schematic diagram of the structure of a computing device according to an embodiment. Detailed implementation manners

[0054] The solutions provided in this specification will be described below with reference to the accompanying drawings.

[0055] Figure 1 It is a schematic diagram of an implementation scenario of an embodiment disclosed in this specification. This implementation scenario involves interaction with a user. It can be understood that the way for the computing device to interact with the user is usually a dialogue. Referring to Figure 1 , in an embodiment of this specification, after obtaining the input data of the user, instead of directly calling the interactive language model to output the dialogue content, a decision model is first used to determine the specific function in the computing device to be called. This function includes the control function for hardware, the call function for the application programming interface (API), and the call function for the interactive language model. Among them, both the control function for hardware and the call function for the API can obtain additional information, which is used as supplementary information for the input data and input into the interactive language model, so that the dialogue output by the interactive language model can better meet the user's needs, thereby improving the interaction experience.

[0056] Figure 2 Schematic flowchart of a method for interacting with a user according to an embodiment. This method can be executed by a computing device based on Figure 1 the implementation scenario shown. As Figure 2 shown, the method for interacting with the user in this embodiment includes the following steps:

[0057] Step 21, obtain the first input data of the user; Step 22, through the decision model, based on the first input data, output a first decision result, where the first decision result is used to indicate the call of the first function provided in the computing device; Step 23, according to the first decision result, perform the call of the first function. The specific execution manners of the above steps will be described below.

[0058] First, in step 21, obtain the user's first input data. It can be understood that the above first input data can be unimodal data, for example, speech; the above first input data can also be multimodal data, for example, including both speech and images.

[0059] In one example, the first input data includes at least one of the following data: speech, images, video for indicating gestures.

[0060] In this example, speech can be converted into text through real-time speech recognition, or the user's input can also be text.

[0061] In the embodiments of this specification, the input data in text form can include eye movement coordinates, real-time speech recognition results, emotion tags, etc.

[0062] In one example, the computing device includes an interactive wearable device. For example, smart glasses.

[0063] In one example, the first input data includes active input information and environmental perception information, and the environmental perception information includes background noise extracted from the active input information in the speech modality, or background images extracted from the active input information in the picture modality.

[0064] Further, the first input data further includes human body perception information, and the human body perception information includes:

[0065] Eye movement coordinates or emotion tags.

[0066] Among them, the above environmental perception information and human body perception information can be extracted after the user's authorization, and can be called passive input information, so as to make decisions and initiate controls according to the user's perception, which can bring a good interaction experience to the user.

[0067] Then, in step 22, through a decision model, based on the first input data, output a first decision result, and the first decision result is used to indicate the invocation of the first function provided in the computing device. It can be understood that the computing device can provide multiple functions, and the decision model selects one or more functions from multiple functions as the first function according to the first input data.

[0068] In one example, the first function includes any one of the following functions:

[0069] The control function for hardware, the invocation function for the application programming interface API, the invocation function for the interactive language model.

[0070] In this example, the control functions of the hardware may include multiple controls for different hardware, such as music players, camera controls, volume controls, display controls, etc.; the functions of calling APIs may include calls to different APIs, such as retrieval-augmented generation (RAG) applications, object recognition applications by scanning, or weather information acquisition applications; and the functions of calling interactive language models may include calls to different interactive language models.

[0071] Among them, the interactive language model can directly output answers to user questions, or output some statements for interacting with the user before specifically answering the user's questions. For example, output the tone word "querying, please wait a moment" for high-latency tasks, or output operation guidance statements such as "The photo is a bit blurry. Can you take another one?", or output consumer sentiment statements such as "Tomorrow is expected to be cloudy, but don't let the weather affect your current good mood~".

[0072] In one example, the output of the first decision result based on the first input data through the decision model includes:

[0073] Obtain a first feature block based on the first input data, and add the first feature block to the first feature block sequence;

[0074] Output the first decision result through the decision model based on the first feature block and a second feature block in the first feature block sequence that is before the first feature block, where the second feature block is obtained based on second input data, and the second input data is the data input by the user before the first input data.

[0075] In this example, an end-to-end streaming processing method is adopted, and real-time streaming responses can be achieved at the input end, inside the model for decision-making, and the output end. Without waiting for the user to express a complete conversation instruction, it will think and respond based on the content that has been expressed at present. This greatly reduces the execution time inside the model and provides users with a high-quality usage experience.

[0076] Figure 3 Show a schematic diagram of the streaming processing of the decision model according to an embodiment. Refer to Figure 3, the user inputs "What kind of flower is this?", and the above input is a streaming input. In the embodiments of this specification, when the user inputs "This is", the decision-making model has already made a decision; when the user inputs "What is this", the decision-making model makes a decision again; when the user inputs "What kind of flower is this", the decision-making model makes a decision again, and the input at this time can also include the execution result of the function called after the previous decision, for example, the taken photo. The encoder in the figure is used to encode the input in the image modality, and the tokenizer is used to encode the input in the text modality. After the encodings of the two are concatenated, they are input into the decision-making model, and the decision-making model outputs an instruction sequence, which determines which functions to call. The call result can further be used as the input of the decision-making model for subsequent decisions. The input "None" in the figure represents no input, the output "None" in the figure represents no output, the output "C: Photo" in the figure represents photo shooting control, the output "A: Flower" in the figure represents calling the application for identifying objects by scanning, and the output "A: llm" in the figure represents calling the large language model, which is a kind of interactive language model.

[0077] Among them, the above encoder realizes preprocessing acceleration, and is mainly used to convert the input data in the image modality into a representation form suitable for processing by the decision-making model, and can solve the asymmetry problem between high-dimensional visual input and low-dimensional language representation.

[0078] Further, obtaining the first feature block based on the first input data includes: obtaining a plurality of first feature blocks arranged in sequence based on the first input data;

[0079] The outputting, by the decision-making model, the first decision result based on the first feature block and the second feature block in the first feature block sequence that is before the first feature block includes:

[0080] Performing attention calculation on each of the plurality of first feature blocks based on the plurality of first feature blocks and the plurality of second feature blocks arranged in sequence before the plurality of first feature blocks in the first feature block sequence to obtain a plurality of third feature blocks arranged in sequence corresponding to the plurality of first feature blocks respectively, and sequentially adding the plurality of third feature blocks to the second feature block sequence;

[0081] For each third feature block, the decision-making model outputs a first decision result corresponding to the third feature block based on the third feature block and the plurality of fourth feature blocks arranged in sequence before the third feature block in the second feature block sequence.

[0082] In this example, attention calculation is adopted to realize the inference acceleration of the decision-making model.

[0083] Figure 4 Shows a schematic diagram of the inference acceleration of the decision-making model according to an embodiment. Refer toFigure 4 , h k-3 , h k-2 , h k-1 form a feature block, which is denoted as Chunki-1; h k , h k+1 , h k+2 form a feature block, which is denoted as Chunki; h k+3 , h k+4 , h k+5 form a feature block, which is denoted as Chunki+1. In chronological order, Chunki-1 is the feature block before Chunki, and Chunki is the feature block before Chunki+1. If the currently input feature block is Chunki, then attention calculation is performed on Chunki based on Chunki-1 and Chunki, where h k , h k+1 , h k+2 form the query vector Queries, h k-3 , h k-2 , h k-1 , h k , h k+1 , h k+2 form the key vector Keys, and after attention calculation, g k , g k+1 , g k+2 form a feature block, and then it is convolved with the feature block formed by g k-3 , g k-2 , g k-1 to obtain the output corresponding to the current input as f k , f k+1 , f k+2 . If the currently input feature block is Chunki+1, then attention calculation is performed on Chunki+1 based on Chunki and Chunki+1, where h k+3 , h k+4 , h k+5 form the query vector Queries, h k , h k+1 , h k+2 , h k+3 , h k+4 , h k+5 form the key vector Keys, and after attention calculation, g k+3 , g k+4 , g k+5 form a feature block, and then it is convolved with the feature block formed by g k , g k+1 , g k+2 to obtain the output corresponding to the current input as fk+3 , f k+4 , f k+5 .

[0084] In one example, the decision model is trained in the following manner:

[0085] Obtain training samples, where the training samples include: model input data and label instruction sequences;

[0086] Input the model input data into the decision model, and output prediction instruction sequences corresponding to each moment respectively. The prediction instruction sequences include predicted hardware control instructions, application programming interface (API) call instructions, and / or model call instructions;

[0087] Train the decision model according to the difference between the prediction instruction sequence and the label instruction sequence.

[0088] In this example, the model input data can be multi-modal data, and can be the complete user input under streaming input, or can be a partial user input under streaming input.

[0089] Figure 5 Shows a schematic diagram of constructing a training sample according to an embodiment. Refer to Figure 5 , first generate user Q&A data based on an image of a car. For example, the user asks questions such as "What is the fuel consumption of this car?" or "What's the weather like today?"; then generate multi-round conversation data. For example, the user asks "What brand is this car?", the smart glasses answer "×××", the user asks "What's the fuel consumption of this car?", the smart glasses answer "Fuel consumption per 100 kilometers is ××"; finally generate training samples for training the decision model, and this training sample corresponds to streaming input and streaming output. When the user inputs "This", the decision model outputs "None", where None means nothing; when the user inputs "This car", the decision model outputs "None C-take picture", where C-takepicture represents taking a picture control; when the user inputs "This car is", the decision model outputs "None C-take picture API-Scan", where API-Scan represents calling the API of the scan application, and at this time the input includes the result of taking a picture; when the user inputs "This car is...", the decision model outputs "None C-take picture API-Scan None"; when the user inputs "This car is... brand", the decision model outputs "None C-take picture API-Scan None None chat", where chat represents an interactive language model.

[0090] Further, the hardware control instruction is used to control the state of the hardware, and the control of the state of the hardware includes:

[0091] Controlling the music player to play music; or,

[0092] Taking photo control; or,

[0093] Volume control; or,

[0094] Display control.

[0095] Further, the API call instruction is used to call the API of the target application, and the target application includes:

[0096] Retrieval-Augmented Generation (RAG) application, object recognition by scanning application, or weather information acquisition application.

[0097] Finally, in step 23, according to the first decision result, the first function is called. It can be understood that the first function may include one or more functions. When multiple functions are included, the above multiple functions can be executed in sequence, and the call result of the previous function can be used as the input of the subsequent call function.

[0098] For example, after taking photo control, if the object recognition by scanning application is called, the photo obtained by the taking photo control can be used as the input of the object recognition by scanning application.

[0099] In one example, the method further includes:

[0100] Based on the third input data, a second decision result is output through a decision model, and the second decision result is used to indicate that based on the result data of calling the first function, a second function provided in the computing device is called. The second function is a function for calling an interactive language model, and the third input data is the data input by the user after the first input data.

[0101] Further, the third input data is the initial input information of the user obtained in a streaming manner at the current moment; the result data of the first function is the first supplementary input information obtained by executing the first hardware control instruction and / or the first Application Programming Interface (API) call instruction;

[0102] The calling of the second function provided in the computing device includes: using the initial input information and the first supplementary input information as the input data of the interactive language model.

[0103] Further, the initial input information includes active input information and passive input information.

[0104] Among them, the actively input information may include voice, image, video, etc. actively input by the user; the passively input information may be human perception information, environmental perception information, etc. extracted with the authorization of the user.

[0105] Through the method provided in the embodiments of this specification, the computing device first obtains the first input data of the user; then, through a decision model, based on the first input data, outputs a first decision result, where the first decision result is used to indicate the invocation of a first function provided in the computing device; and then, according to the first decision result, invokes the first function. As can be seen from the above, in the embodiments of this specification, during the interaction process between the computing device and the user, instead of directly outputting the conversation content based on the input data of the user, it first makes a decision through the decision model to invoke the specific function provided in the computing device. On the one hand, it makes real-time control decisions through sensing input, and on the other hand, it can undertake a wider range of applications through function invocation, and the function invocation result helps to better identify the user's intention. In summary, it makes the output conversation better meet the user's needs, thereby improving the interaction experience.

[0106] According to an embodiment on the other hand, there is also provided a device for interacting with a user, and this device is used to execute the method provided in the embodiments of this specification. Figure 6 The schematic block diagram of a device for interacting with a user according to an embodiment is shown. As Figure 6 shown, the device 600 includes:

[0107] An obtaining unit 61, configured to obtain the first input data of the user;

[0108] A decision-making unit 62, configured to output a first decision result through a decision model based on the first input data obtained by the obtaining unit 61, where the first decision result is used to indicate the invocation of a first function provided in the computing device;

[0109] An invocation unit 63, configured to invoke the first function according to the first decision result obtained by the decision-making unit 62.

[0110] Optionally, as an embodiment, the first input data includes at least one of the following data: voice, image, video for indicating gestures.

[0111] Optionally, as an embodiment, the first function includes any one of the following functions:

[0112] The control function for hardware, the invocation function for the application programming interface API, the invocation function for the interaction language model.

[0113] Optionally, as an embodiment, the decision-making unit 62 includes:

[0114] An adding subunit is configured to obtain a first feature block based on the first input data and add the first feature block to a first feature block sequence;

[0115] A processing subunit is configured to output a first decision result through a decision model based on the first feature block and a second feature block before the first feature block in the first feature block sequence obtained by the adding subunit, where the second feature block is obtained based on second input data, and the second input data is data input by the user before the first input data.

[0116] Further, the adding subunit is specifically configured to obtain a plurality of first feature blocks arranged in order based on the first input data;

[0117] The processing subunit is specifically configured to:

[0118] Perform attention calculation on each first feature block based on the plurality of first feature blocks and a plurality of second feature blocks arranged in order before the plurality of first feature blocks in the first feature block sequence, to obtain a plurality of third feature blocks arranged in order corresponding to the plurality of first feature blocks respectively, and sequentially add the plurality of third feature blocks to a second feature block sequence;

[0119] For each third feature block, the decision model outputs a first decision result corresponding to the third feature block based on the third feature block and a plurality of fourth feature blocks arranged in order before the third feature block in the second feature block sequence.

[0120] Optionally, as an embodiment, the decision unit 62 is further configured to output a second decision result through a decision model based on third input data, where the second decision result is used to indicate to call a second function provided in the computing device based on result data of calling a first function, the second function is a function of calling an interactive language model, and the third input data is data input by the user after the first input data.

[0121] Further, the third input data is initial input information of the user obtained in a streaming manner at the current moment; the result data of the first function is first supplementary input information obtained by executing a first hardware control instruction and / or a first application programming interface (API) call instruction;

[0122] The calling the second function provided in the computing device includes: using the initial input information and the first supplementary input information as input data of the interactive language model.

[0123] Further, the initial input information includes active input information and passive input information.

[0124] Optionally, as an embodiment, the computing device includes an interactive wearable device.

[0125] Optionally, as an embodiment, the decision-making model is trained in the following manner:

[0126] Obtain training samples, where the training samples include: model input data and label instruction sequences;

[0127] Input the model input data into the decision-making model, and output prediction instruction sequences corresponding to each moment respectively. The prediction instruction sequences include predicted hardware control instructions, application programming interface (API) call instructions, and / or model call instructions;

[0128] Train the decision-making model according to the difference between the prediction instruction sequence and the label instruction sequence.

[0129] Furthermore, the hardware control instructions are used to control the state of the hardware. Controlling the state of the hardware includes:

[0130] Controlling a music player to play music; or,

[0131] Taking a photo control; or,

[0132] Volume control; or,

[0133] Display control.

[0134] Furthermore, the API call instructions are used to call the API of the target application. The target applications include:

[0135] Retrieval augmented generation (RAG) applications, object recognition applications by scanning, or weather information acquisition applications.

[0136] Optionally, as an embodiment, the first input data includes active input information and environmental perception information. The environmental perception information includes background noise extracted from the active input information in the speech modality, or background images extracted from the active input information in the picture modality.

[0137] Furthermore, the first input data further includes human body perception information. The human body perception information includes:

[0138] Eye movement coordinates or emotion labels.

[0139] Through the device provided by the embodiments of this specification, the computing device first obtains the user's first input data by the obtaining unit 61; then the decision-making unit 62 outputs a first decision result based on the first input data through the decision-making model, and the first decision result is used to indicate to call the first function provided in the computing device; then the calling unit 63 calls the first function according to the first decision result. As can be seen from the above, in the embodiments of this specification, during the interaction process between the computing device and the user, instead of directly outputting the conversation content based on the user's input data, it first decides to call the specific functions provided in the computing device through the decision-making model. On the one hand, it makes real-time control decisions by perceiving the input. On the other hand, through function calls, it can undertake a wider range of applications, and the function call results help to better identify the user's intention. In summary, the output conversation can better meet the user's needs, thereby improving the interaction experience.

[0140] Figure 7 FIG. shows a schematic structural diagram of a computing device according to an embodiment. Referring to Figure 7 , in addition to the decision-making model on the computing device, some software and hardware need to be equipped so that the computing device can implement various functions. For example, the computing device supports cloud platform function 71, camera function 72, positioning function 73, scan and identify function 74, etc. The embodiments of this specification can cover various common functions and will not be listed one by one here.

[0141] According to an embodiment of another aspect, there is also provided a computer-readable storage medium, on which a computer program is stored. When the computer program is executed in a computer, it causes the computer to execute the method described in conjunction with Figure 2 .

[0142] According to an embodiment of still another aspect, there is also provided a computing device, including a memory and a processor. An executable code is stored in the memory. When the processor executes the executable code, it implements the method described in conjunction with Figure 2 .

[0143] Those skilled in the art should be able to realize that in the above one or more examples, the functions described in the present invention can be implemented by hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.

[0144] The specific embodiments described above further elaborate on the object, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only the specific embodiments of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solution of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for interacting with a user, the method being performed by a computing device, comprising: Obtaining the user's first input data; Outputting a first decision result based on the first input data through a decision model, wherein the first decision result is used to indicate calling a first function provided in the computing device; The first function is called according to the first decision result.

2. The method of claim 1, wherein: The first input data includes at least one of the following data: voice, image, and video for indicating gestures.

3. The method of claim 1, wherein: The first function includes any of the following functions: The functions of controlling hardware, calling application programming interface API, and calling interactive language model.

4. The method of claim 1, wherein: The step of outputting a first decision result based on the first input data through the decision model includes: Acquire a first feature block based on the first input data, and add the first feature block to a first feature block sequence; A first decision result is outputted through a decision model based on the first feature block and a second feature block preceding the first feature block in the first feature block sequence, wherein the second feature block is acquired based on second input data, and the second input data is data input by the user before the first input data.

5. The method according to claim 4, wherein obtaining a first feature block based on the first input data comprises: Acquire a plurality of first feature blocks arranged in sequence based on the first input data; The step of outputting a first decision result based on the first feature block and a second feature block preceding the first feature block in the first feature block sequence through a decision model comprises: Performing attention calculation on each first feature block based on the multiple first feature blocks and multiple second feature blocks in the first feature block sequence that are sequentially arranged before the multiple first feature blocks, obtaining multiple third feature blocks that are sequentially arranged and correspond to the multiple first feature blocks respectively, and sequentially adding the multiple third feature blocks to the second feature block sequence; For each third feature block, the decision model outputs a first decision result corresponding to the third feature block based on the third feature block and a plurality of fourth feature blocks in the second feature block sequence that are sequentially arranged before the third feature block.

6. The method of claim 1, wherein: The method further comprises: Through the decision model, based on the third input data, a second decision result is output, the second decision result is used to indicate calling a second function provided in the computing device based on the result data of calling the first function, the second function is a calling function of the interactive language model, and the third input data is the data input by the user after the first input data.

7. The method of claim 6, wherein: The third input data is initial input information of the user at the current moment obtained in a streaming manner; the result data of the first function is first supplementary input information obtained by executing the first hardware control instruction and / or the first application programming interface API call instruction; The calling of the second function provided in the computing device includes: using the initial input information and the first supplementary input information as input data of the interactive language model.

8. The method of claim 6, wherein: The initial input information includes active input information and passive input information.

9. The method of claim 1, wherein: The computing device includes an interactive wearable device.

10. The method of claim 1, wherein: The decision model is trained in the following way: Acquire a training sample, wherein the training sample includes: model input data and a label instruction sequence; Inputting the model input data into the decision model, and outputting the prediction instruction sequence corresponding to each moment, wherein the prediction instruction sequence includes the predicted hardware control instruction, application programming interface API call instruction and / or model call instruction; The decision model is trained according to the difference between the predicted instruction sequence and the label instruction sequence.

11. The method of claim 10, wherein: The hardware control instruction is used to control the state of the hardware, and the state of the hardware control includes: Control the music player to play music; or, Photo control; or, volume control; or, Display controls.

12. The method of claim 10, wherein: The API call instruction is used to call the API of the target application, and the target application includes: Retrieve enhanced generation of RAG applications, scan-and-recognize object applications, or weather information acquisition applications.

13. The method of claim 1, wherein: The first input data includes active input information and environmental perception information, and the environmental perception information includes background noise extracted from active input information in a speech modality, or background images extracted from active input information in a picture modality.

14. The method of claim 13, wherein: The first input data also includes human perception information, and the human perception information includes: Eye movement coordinates or emotion labels.

15. An apparatus for interacting with a user, the apparatus being disposed in a computing device, comprising: An acquisition unit, used for acquiring first input data of a user; a decision unit, configured to output a first decision result based on the first input data acquired by the acquisition unit through a decision model, wherein the first decision result is used to indicate calling a first function provided in the computing device; A calling unit is used to call the first function according to the first decision result obtained by the decision unit.

16. A computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to execute the method according to any one of claims 1 to 14.

17. A computing device comprising a memory and a processor, wherein the memory stores executable codes, and when the processor executes the executable codes, the method according to any one of claims 1 to 14 is implemented.