Human-computer interaction method and device and electronic equipment

By acquiring user input and images in electronic devices, determining areas of interest and using their processing results, the problem of insufficient accuracy of processing results is solved, and more efficient computing and privacy protection is achieved.

CN120010653APending Publication Date: 2025-05-16BEIJING ZHICUN (WITIN) TECH CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411794689.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-07
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

When existing electronic devices process user input, it is difficult to ensure the accuracy of processing results, which affects the user's human-computer interaction experience.

Method used

By obtaining user input and images, identifying the region of interest (ROI), and using only the ROI processing results, reducing computational overhead and protecting privacy.

Benefits of technology

Improves the accuracy of processing results, improves the user's human-computer interaction experience, while reducing computing overhead and protecting privacy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120010653A_ABST
    Figure CN120010653A_ABST
Patent Text Reader

Abstract

The invention provides a man-machine interaction method, a man-machine interaction device and electronic equipment, which can be applied to the field of man-machine interaction. The method comprises the following steps: acquiring an image and input of a user; determining a region of interest (ROI) from the image according to the input; obtaining a processing result of the ROI; and controlling the prompt device to prompt the processing result to the user. According to the method and the device, the accuracy of the processing result finally presented to the user can be improved, so that the human-computer interaction experience of the user can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of human-computer interaction, and more specifically, to a human-computer interaction method, device and electronic device. Background Art

[0002] Currently, artificial intelligence (AI) is increasingly used in electronic devices. Users can input text or voice on electronic devices to obtain relevant processing results. For example, if the electronic device is a vehicle, when the vehicle receives the voice input from the user "help me generate a paragraph of text related to the scenery outside the vehicle", the vehicle can output relevant text content. The accuracy of the processing results presented by the electronic device to the user directly affects the user's human-computer interaction experience.

[0003] Therefore, how to improve the accuracy of processing results presented by electronic devices has become an issue that needs to be urgently addressed. Summary of the invention

[0004] The present application provides a human-computer interaction method, device and electronic device, which help to improve the accuracy of processing results presented to users, thereby helping to improve the user's human-computer interaction experience.

[0005] In a first aspect, the present application provides a human-computer interaction method, the method comprising: acquiring a first image and a first input from a user; determining a first region of interest (ROI) from the first image based on the first input; acquiring a processing result of the first ROI; and controlling a prompt device to prompt the user with the processing result.

[0006] Based on the above technical solution, the ROI is determined from the image according to the user's input and the processing result of the ROI is obtained, so that the processing result can be more in line with the user's expectations, thereby helping to improve the user's human-computer interaction experience. At the same time, since only the ROI in the image is used when obtaining the processing result instead of the entire image, the computational overhead required to obtain the processing result can be reduced.

[0007] In addition, since only the ROI is used instead of the entire image, it helps to avoid unnecessary privacy leakage in the process of obtaining processing results, and can protect the privacy security of oneself and / or others.

[0008] In some possible implementations, acquiring a first image and a first input from a user includes: in response to acquiring the first input, acquiring the first image.

[0009] In some possible implementations, the first image may be an image taken some time or at a certain moment before the first input is obtained; or, the first image may be an image taken when the first input is obtained; or, the first image may be an image taken some time or at a certain moment after the first input is obtained.

[0010] In combination with the first aspect, in some possible implementations of the first aspect, the first input is a first voice input or a first text input, and based on the first input, a first ROI is determined from a first image, including: inputting the first input and the first image into a target detection model to obtain a first ROI, and an output of the target detection model changes based on a change in the semantics of the first target in the first input.

[0011] Based on the above technical solution, the first ROI can be obtained by inferring the target detection model. Since the input of the target detection model changes based on the change of the target semantics, the accuracy of the model inference result can be improved, and thus the accuracy of the final processing result can be improved.

[0012] In some possible implementations, the object detection model may also be referred to as an open vocabulary object detection model.

[0013] In combination with the first aspect, in some possible implementations of the first aspect, the first input is a first voice input, and based on the first input, a first ROI is determined from the first image, including: determining first text content based on the first voice input; and determining the first ROI from the first image based on the first text content.

[0014] Based on the above technical solution, by analyzing the text content corresponding to the voice input, the ROI in the image can be extracted, which helps to improve the accuracy of the extracted ROI, thereby helping to improve the accuracy of the final processing result.

[0015] In combination with the first aspect, in some possible implementations of the first aspect, determining a first ROI from a first image based on the first text content includes: when the first text content includes second text content related to direction, determining the first ROI based on the second text content; or, when the first text content includes third text content related to attributes of the first target, determining the first ROI based on the third text content, the first ROI including the first target.

[0016] Based on the above technical solution, ROI can be extracted from the image through the target text content related to the direction or the target text content related to the attribute of the target in the text content. Since the target text content can indicate the area of ​​interest to the user in the image, it helps to improve the accuracy of the ROI extracted from the image, thereby helping to improve the accuracy of the final processing result.

[0017] In some possible implementations, the attribute of the first object includes one or more of the color, size, brand, and type of the first object.

[0018] In combination with the first aspect, in some possible implementations of the first aspect, determining the first ROI based on the third text content includes: determining multiple ROIs from the first image based on the third text content; and determining the first ROI from the multiple ROIs based on the user's line of sight when the first voice input is obtained.

[0019] Based on the above technical solution, when multiple ROIs are determined by the attributes of the target, the first ROI can be determined from the multiple ROIs in combination with the user's sight direction. In this way, by combining the user's sight direction as the basis for screening ROIs, it is helpful to improve the accuracy of the extracted ROI, thereby helping to improve the accuracy of the final processing result.

[0020] In combination with the first aspect, in some possible implementations of the first aspect, determining the first ROI according to the third text content includes: determining multiple ROIs from the first image according to the third text content; controlling a prompt device to prompt a user to select one or more ROIs from the multiple ROIs; and determining the first ROI in response to detecting a second input from the user, the second input instructing the user to select the first ROI.

[0021] Based on the above technical solution, when multiple ROIs are determined by the attributes of the target, the user can be prompted to select a certain ROI. In this way, the ROI expected by the user can be accurately obtained by combining the user's selection, which helps to improve the accuracy of the final processing result.

[0022] In combination with the first aspect, in some possible implementations of the first aspect, determining a first ROI from a first image based on a first text content includes: when the first text content does not include direction-related text content and does not include text content related to attributes of the target, obtaining a user's line of sight direction; determining the first ROI based on the line of sight direction.

[0023] Based on the above technical solution, when the target text content is not included in the text content, the ROI can be determined in combination with the user's line of sight, which helps to improve the accuracy of the extracted ROI, thereby helping to improve the accuracy of the final processing result.

[0024] In combination with the first aspect, in some possible implementations of the first aspect, the first input is a first voice input, and the method includes: obtaining a first gesture of the user when the user triggers the first voice input; wherein, based on the first input, determining a first ROI from the first image includes: determining the first ROI based on the first gesture.

[0025] Based on the above technical solution, determining the ROI by the user's gesture when triggering voice input helps to improve the accuracy of the extracted ROI, thereby helping to improve the accuracy of the final processing result.

[0026] In combination with the first aspect, in some possible implementations of the first aspect, obtaining a processing result of the first ROI includes: sending a first input and a first ROI to a cloud server; and receiving a processing result determined by the cloud server based on the first input and the first ROI.

[0027] Based on the above technical solution, taking the above human-computer interaction method executed by an electronic device as an example, the electronic device can send the first input and the first ROI to the cloud server, so that the cloud server obtains the processing result based on the first input and the first ROI. In this way, by leveraging the high computing power of the cloud server, it helps to reduce the delay when the user interacts with the electronic device. At the same time, since the electronic device only sends the first ROI instead of the first image to the cloud server, it helps to avoid unnecessary privacy leakage, and can achieve the role of protecting the privacy security of oneself and / or others.

[0028] In combination with the first aspect, in some possible implementations of the first aspect, obtaining a processing result of the first ROI includes: inputting the first input and the first ROI into a content generation model to obtain a processing result.

[0029] Based on the above technical solution, taking the above human-computer interaction method executed by an electronic device as an example, the processing result can be obtained by generating a model based on the first input and the first ROI input content. In this way, the processing result can be made more in line with the user's expectations, thereby helping to improve the user's human-computer interaction experience. At the same time, since only the ROI in the image is used when obtaining the processing result instead of the entire image, the computing overhead required for the electronic device to obtain the processing result can be reduced, thereby helping to reduce the power consumption of the electronic device.

[0030] In combination with the first aspect, in some possible implementations of the first aspect, the first input is a second gesture of the user, and determining the first ROI from the first image based on the first input includes: determining the first ROI based on the direction of the finger in the second gesture.

[0031] Based on the above technical solution, the direction of the finger in the second gesture can be used to determine the ROI from the image. In this way, the ROI can be determined only through gesture input without voice input or text input, thereby obtaining a processing result that meets the user's expectations. This helps to improve the intelligence of electronic devices, avoids tedious user input before obtaining processing results, and helps to improve the user's human-computer interaction experience.

[0032] In combination with the first aspect, in some possible implementations of the first aspect, the first input is a third gesture of the user, and based on the first input, a first ROI is determined from the first image, including: when the third gesture is a preset gesture, determining the first ROI based on the user's line of sight.

[0033] Based on the above technical solution, when the user's gesture is a preset gesture, the ROI can be determined based on the user's line of sight, so as to obtain a processing result that meets the user's expectations. This helps to improve the intelligence of electronic devices, avoids tedious user input before obtaining processing results, and helps to improve the user's human-computer interaction experience.

[0034] In combination with the first aspect, in some possible implementations of the first aspect, the first input is an input for a first button, and based on the first input, a first ROI is determined from the first image, including: obtaining a line of sight direction of the user when the input is for the first button; and determining the first ROI based on the line of sight direction.

[0035] Based on the above technical solution, the ROI is determined by the user's line of sight when the button is triggered, and then the processing result is obtained. In this way, the convenience of the user in obtaining the processing result can be improved, which helps to improve the intelligence of the electronic device.

[0036] In some possible implementations, obtaining the user's line of sight direction when inputting the first key includes: obtaining the user's line of sight direction after a preset time period from the inputting of the first key.

[0037] In combination with the first aspect, in some possible implementations of the first aspect, obtaining a processing result of the first ROI includes: sending the first ROI to a cloud server; and receiving a processing result determined by the cloud server based on the first ROI.

[0038] In combination with the first aspect, in some possible implementations of the first aspect, obtaining a processing result of the first ROI includes: inputting the first ROI into a content generation model to obtain a processing result.

[0039] In combination with the first aspect, in some possible implementations of the first aspect, obtaining the processing result of the ROI includes: determining the processing result according to the second target semantics included in the first input and the first ROI.

[0040] Based on the above technical solution, the processing result can be determined according to the second target semantics and the first ROI included in the first input. In this way, by determining the processing result through the target semantics and the ROI, the processing result can be made more in line with the user's expectations, which helps to improve the user's human-computer interaction experience. At the same time, the process of determining the processing result can be made invisible to the user.

[0041] In some possible implementations, the first target semantics and the second target semantics are different.

[0042] In a second aspect, the present application provides a human-computer interaction device, which includes: an acquisition unit, used to acquire a first image and a first input from a user; a determination unit, used to determine a first region of interest ROI from the first image based on the first input; the acquisition unit, also used to acquire a processing result of the ROI; and a control unit, used to control a prompt device to prompt the user with the processing result.

[0043] In combination with the second aspect, in some possible implementations of the second aspect, the first input is a first voice input or a first text input, and the determination unit is used to: input the first input and the first image into a target detection model to obtain a first ROI, and the output of the target detection model changes based on the change of the semantics of the first target in the first input.

[0044] In combination with the second aspect, in some possible implementations of the second aspect, the first input is a first voice input, and the determination unit is used to: determine the first text content according to the first voice input; and determine the first ROI from the first image according to the first text content.

[0045] In combination with the second aspect, in some possible implementations of the second aspect, the determination unit is used to: when the first text content includes second text content related to the direction, determine the first ROI according to the second text content; or, when the first text content includes third text content related to the attribute of the first target, determine the first ROI according to the third text content, and the first ROI includes the first target.

[0046] In combination with the second aspect, in some possible implementations of the second aspect, the determination unit is used to: determine multiple ROIs from the first image based on the third text content; and determine the first ROI from the multiple ROIs based on the user's line of sight when the first voice input is obtained.

[0047] In combination with the second aspect, in some possible implementations of the second aspect, the determination unit is used to: determine multiple ROIs from the first image according to the third text content; the control unit is used to control the prompt device to prompt the user to select one or more ROIs from the multiple ROIs; the determination unit is used to determine the first ROI in response to the user's second input, and the second input instructs the user to select the first ROI.

[0048] In combination with the second aspect, in some possible implementations of the second aspect, the acquisition unit is used to obtain the user's line of sight direction when the first text content does not include direction-related text content and does not include text content related to the attributes of the target; the determination unit is used to determine the first ROI based on the line of sight direction.

[0049] In combination with the second aspect, in some possible implementations of the second aspect, the first input is a first voice input, and the acquisition unit is further used to acquire a first gesture of the user when the user triggers the first voice input; the determination unit is used to determine the first ROI based on the first gesture.

[0050] In combination with the second aspect, in some possible implementations of the second aspect, the device also includes a sending unit, which is used to send the first input and the first ROI to the cloud server; and the acquiring unit is used to receive a processing result determined by the cloud server based on the first input and the first ROI.

[0051] In combination with the second aspect, in some possible implementations of the second aspect, the acquisition unit is used to: input the first input and the first ROI into the content generation model to obtain a processing result.

[0052] In combination with the second aspect, in some possible implementations of the second aspect, the first input is a second gesture of the user, and the determination unit is used to determine the first ROI according to the direction of the finger in the second gesture.

[0053] In combination with the second aspect, in some possible implementations of the second aspect, the first input is a third gesture of the user, and the determination unit is used to: when the third gesture is a preset gesture, determine the first ROI according to the line of sight of the user.

[0054] In combination with the second aspect, in some possible implementations of the second aspect, the first input is an input to a first button, and the acquisition unit is further used to obtain the user's line of sight direction when the input is to the first button; the determination unit is used to determine the first ROI based on the line of sight direction.

[0055] In combination with the second aspect, in some possible implementations of the second aspect, the device further includes a sending unit, the sending unit being configured to send the first ROI to the cloud server; and the acquiring unit being configured to receive a processing result determined by the cloud server based on the first ROI.

[0056] In combination with the second aspect, in some possible implementations of the second aspect, the acquiring unit is configured to: input the first ROI into the content generation model to obtain a processing result.

[0057] In combination with the second aspect, in some possible implementations of the second aspect, the determining unit is further configured to determine a processing result according to the second target semantics and the first ROI included in the first input.

[0058] In a third aspect, the present application provides a human-computer interaction device, which includes: a memory for storing a computer program; a processor for executing the computer program stored in the memory, so that the device performs any method described in any one of the above-mentioned first aspects.

[0059] In a fourth aspect, the present application provides an electronic device, which includes the device described in any one of the second aspect or the third aspect.

[0060] In a fifth aspect, the present application provides a cloud server, which includes the device described in any one of the second aspect or the third aspect above.

[0061] In a sixth aspect, the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is run on a computer, the computer executes any one of the methods in the first aspect.

[0062] In a seventh aspect, a computer program product is provided, the computer program product comprising: a computer program code, when the computer program code is run on a computer, the computer executes any one of the methods in the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] Figure 1 It is a schematic block diagram of an electronic device provided in an embodiment of the present application.

[0064] Figure 2 It is a schematic flow chart of the human-computer interaction method provided in an embodiment of the present application.

[0065] Figure 3 It is another schematic flow chart of the human-computer interaction method provided in an embodiment of the present application.

[0066] Figure 4 It is a schematic diagram of determining ROI provided in an embodiment of the present application.

[0067] Figure 5 It is another schematic flow chart of the human-computer interaction method provided in an embodiment of the present application.

[0068] Figure 6 It is another schematic flow chart of the human-computer interaction method provided in an embodiment of the present application.

[0069] Figure 7 It is a schematic diagram of dividing the first image into regions provided in an embodiment of the present application.

[0070] Figure 8 This is another schematic diagram of dividing the first image into regions provided in an embodiment of the present application.

[0071] Fig. 9 It is a schematic block diagram of a human-computer interaction device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0072] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application. Among them, in the description of the embodiments of the present application, unless otherwise specified, " / " means or, for example, A / B can mean A or B; "and / or" in this article is only a description of the association relationship of associated objects, indicating that three relationships can exist, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. "At least one" means one or more. For example, "at least one of A and B" is similar to "A and / or B", describing the association relationship of associated objects, indicating that three relationships can exist, for example, at least one of A and B can mean: A exists alone, A and B exist at the same time, and B exists alone.

[0073] In the embodiments of the present application, prefixes such as "first" and "second" are used only to distinguish different description objects, and have no limiting effect on the position, order, priority, quantity or content of the described objects. The use of prefixes such as ordinal numbers to distinguish description objects in the embodiments of the present application does not constitute a limitation on the described objects. For the statement of the described objects, please refer to the description in the context of the claims or embodiments, and no unnecessary limitation should be constituted due to the use of such prefixes. In addition, in the description of the present embodiment, unless otherwise specified, the meaning of "multiple" is two or more.

[0074] Figure 1 1 is a schematic block diagram of an electronic device 100 provided in an embodiment of the present application. The electronic device 100 may include a processor 110 , a memory 120 , and a camera 130 .

[0075] The processor 110 may include one or more processing units, for example, the processor 110 may include an application processor (AP), a modem processor, a graphics processor (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, or a neural-network processing unit (NPU), etc. Different processing units may be independent components or integrated into one or more processors.

[0076] The memory 120 may include an internal memory. The internal memory may be used to store one or more computer programs, which include instructions. The processor 110 may enable the electronic device 100 to execute the human-computer interaction method provided in some embodiments of the present application, run various applications, and process data by running the above instructions stored in the internal memory. The internal memory may include a program storage area and a data storage area. Among them, the program storage area may store an operating system; the program storage area may also store one or more applications, etc. The data storage area may store data created during the use of the electronic device 100. In addition, the internal memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more disks (or disk storage components), a flash memory component, a universal flash storage (UFS), etc. In some embodiments, the processor 110 may enable the electronic device 100 to execute the method provided in the embodiment of the present application, run various applications and process data by running instructions stored in the internal memory, and / or instructions stored in a memory provided in the processor 110.

[0077] Optionally, the electronic device 100 may include a display device 140. The display device 140 is used to display images, videos, a series of graphical user interfaces (GUIs), etc. The display device 140 may include a display screen or a projection device, etc.

[0078] The type of electronic device 100 is not limited in the embodiments of the present application. For example, according to some embodiments, the electronic device 100 may include a wearable device. Wearable devices include, for example, but are not limited to: head-mounted devices (such as helmets or hats, etc.), devices wearable on the ears (such as headphones), devices wearable on the wrist (such as watches), devices wearable on other parts (for example, electronic necklaces, medical monitoring equipment, or glasses, etc.), etc. According to some embodiments, the electronic device 100 may include a portable terminal. For example, the electronic device 100 may include, but is not limited to, mobile phones, general computing devices (such as laptop computers, or tablet computers, etc.), personal digital assistants, and the like. According to some embodiments, the electronic device 100 may include other types of end-side devices, such as personal computers, vehicle-mounted computers, vehicle-mounted computing platforms, vehicles, or smart home electronic products, etc.

[0079] Figure 2 A schematic flow chart of a human-computer interaction method 200 provided in an embodiment of the present application is shown. The method 200 may be executed by the electronic device 100; or, may be executed by the processor 110; or, may be executed by a cloud server. The method 200 includes:

[0080] S210: Acquire a first image and a first input of a user.

[0081] Exemplarily, the first input may be one or more of key input, voice input, text input or gesture input.

[0082] Exemplarily, the key input may be an input to a physical key or an input to a virtual key in the electronic device.

[0083] Exemplarily, the voice input may be voice content uttered by a user.

[0084] Exemplarily, the text input may be text content input by a user through an input device of the electronic device that can input text, such as a display screen.

[0085] Exemplarily, the gesture input may be a gesture made by a user.

[0086] Exemplarily, the first input may include an original input or a converted input. For example, the first input may include a text input converted from a voice input as the original input.

[0087] Optionally, acquiring the first image and the first input of the user includes: acquiring the first image in response to acquiring the first input.

[0088] For example, taking the method 200 being executed by a vehicle and the first input being a voice input as an example, when the user's voice input "help me confirm what brand the red vehicle in front is", is obtained, the vehicle can obtain the first image captured by the camera outside the cabin.

[0089] For example, the method 200 is performed by a headset equipped with a camera and the first input is a voice input. When the user's voice input "help me introduce the purpose of this tool" is obtained, the headset can obtain the first image captured by the camera.

[0090] Optionally, the first image may be an image taken some time or at a certain moment before the first input is obtained; or, the first image may be an image taken when the first input is obtained; or, the first image may be an image taken some time or at a certain moment after the first input is obtained.

[0091] S220: Determine a first region of interest ROI from the first image according to the first input.

[0092] Optionally, the first input is a first voice input or a first text input, and according to the first input, a first ROI is determined from the first image, including: inputting the first input and the first image into a target detection model to obtain the first ROI, and the output of the target detection model changes based on the change of the semantics of the first target in the first input. Since the input of the target detection model changes based on the change of the semantics of the target in the first input, the accuracy of the model reasoning result can be improved, and thus the accuracy of the final processing result can be improved.

[0093] The above target detection model can be trained by a training data set, which includes sample data. The sample data includes target keywords, sample images, and label information. The label information is used to indicate the ROI corresponding to the target keyword in the sample image.

[0094] Optionally, the object detection model may include an open vocabulary object detection model. The open vocabulary object detection model can detect categories that are not predefined during the training process. This can improve the versatility of the model.

[0095] Optionally, the object detection model may include a multimodal model.

[0096] Optionally, the first input is a first voice input, and determining a first ROI from the first image according to the first input includes: determining first text content according to the first voice input; and determining the first ROI from the first image according to the first text content.

[0097] Optionally, determining the first ROI from the first image according to the first text content includes: when the first text content includes the second text content related to the direction, determining the first ROI according to the second text content; or, when the first text content includes the third text content related to the attribute of the first target, determining the first ROI according to the third text content, wherein the first ROI includes the first target. In this way, extracting the ROI through the target text content helps to improve the accuracy of the extracted ROI, thereby helping to improve the accuracy of the final processing result.

[0098] Exemplarily, the attribute of the first object includes one or more of the color, size, brand, and type of the first object.

[0099] Optionally, the attribute of the first target may also be the distance between the first target and the electronic device 100 .

[0100] Optionally, determining the first ROI according to the third text content includes: determining multiple ROIs from the first image according to the third text content; and determining the first ROI from the multiple ROIs according to the user's sight direction when the first voice input is obtained. In this way, by combining the user's sight direction as the basis for extracting the ROI, it is helpful to improve the accuracy of the extracted ROI, thereby helping to improve the accuracy of the final processing result.

[0101] Exemplarily, the method 200 is performed by a vehicle and the first input is a voice input. When the user's voice input "help me confirm what brand the red vehicle in front is", the vehicle can obtain the image 1 captured by the camera outside the cockpit. Through the text content related to the attribute of the target in the voice input (for example, "red vehicle"), ROI1 and ROI2 can be extracted from the image 1, and ROI1 and ROI2 respectively include red vehicles. For example, the image 1 can be divided into multiple regions, ROI1 can be region a in multiple regions, and ROI2 can be region b in multiple regions. If the user's line of sight points to the region a when the user triggers the voice input, ROI1 can be selected as the target ROI. Optionally, determining the first ROI according to the third text content includes: determining multiple ROIs from the first image according to the third text content; controlling the prompt device to prompt the user to select one or more ROIs from the multiple ROIs; in response to detecting the second input of the user, determining the first ROI, and the second input instructs the user to select the first ROI. In this way, the ROI expected by the user can be accurately obtained by combining the user's input, which helps to improve the accuracy of the final processing result.

[0102] Exemplarily, take the case where the method 200 is executed by a vehicle and the first input is a voice input. When the user's voice input "Help me confirm what brand the red vehicle in front is", the vehicle can obtain image 2 captured by the camera outside the cockpit. Through the text content related to the attributes of the target in the voice input (for example, "red vehicle"), ROI3 and ROI4 can be extracted from image 2, and ROI3 and ROI4 respectively include red vehicles. For example, image 2 can be divided into multiple regions, ROI3 can be region c among multiple regions, and ROI4 can be region d among multiple regions, where region c is the region close to the left side of image 2, and region d is the region close to the left side of image 2. Figure 2 At this time, the vehicle can issue a prompt voice through the speaker in the cockpit, "Do you want to confirm the red vehicle on the left or the red vehicle on the right?" In response to receiving the user's voice reply "red vehicle on the left", ROI3 can be selected as the target ROI.

[0103] Optionally, determining a first ROI from the first image according to the first text content includes: when the first text content does not include direction-related text content and does not include text content related to attributes of the target, obtaining the user's line of sight direction; determining the first ROI according to the line of sight direction.

[0104] In the embodiment of the present application, when the target text content is not included in the text content, the ROI can be determined in combination with the user's line of sight direction, which helps to improve the accuracy of the extracted ROI, thereby helping to improve the accuracy of the final processing result.

[0105] Optionally, the first input is a first voice input, and the method includes: obtaining a first gesture of the user when the user triggers the first voice input; wherein, according to the first input, determining the first ROI from the first image includes: determining the ROI according to the first gesture. In this way, determining the ROI by the gesture of the user when the voice input is triggered helps to improve the accuracy of the extracted ROI, thereby helping to improve the accuracy of the final processing result.

[0106] Optionally, the first input is a second gesture of the user, and determining the first ROI from the first image according to the first input includes: determining the first ROI according to the direction of the finger in the second gesture. In this way, the ROI can be determined only through gesture input without voice input or text input, thereby obtaining a processing result that meets the user's expectations.

[0107] Optionally, the first input is a third gesture of the user, and determining the first ROI from the first image according to the first input includes: when the third gesture is a preset gesture, determining the first ROI according to the sight direction of the user.

[0108] In the embodiment of the present application, when the user's gesture is a preset gesture, the ROI can be determined based on the user's own line of sight, so as to obtain a processing result that meets the user's expectations. This helps to improve the intelligence of the electronic device, avoids the user's tedious input before obtaining the processing result, and helps to improve the user's human-computer interaction experience.

[0109] Exemplarily, the preset gesture is an OK gesture.

[0110] Optionally, the first input is an input for a first button, and determining the first ROI from the first image according to the first input includes: acquiring a sight direction of the user when inputting the first button; and determining the first ROI according to the sight direction.

[0111] In the embodiment of the present application, the ROI is determined by the user's line of sight when the button is triggered, and then the processing result is obtained. In this way, the convenience of the user in obtaining the processing result can be improved, which helps to improve the intelligence of the electronic device.

[0112] Optionally, obtaining the user's line of sight direction when inputting the first key includes: obtaining the user's line of sight direction after a preset time period from the inputting of the first key.

[0113] S230: Obtain a processing result of the first ROI.

[0114] Optionally, if the first input includes voice input or text input, obtaining the processing result of the first ROI includes: sending the first input and the first ROI to a cloud server; and receiving the processing result determined by the cloud server based on the first input and the first ROI.

[0115] Taking the above method 200 executed by the electronic device 100 and the first input including voice input as an example, the electronic device 100 can send the voice input and ROI to the cloud server, so that the cloud server obtains the processing result based on the voice input and ROI. In this way, by leveraging the high computing power of the cloud server, it helps to reduce the delay when the user interacts with the electronic device 100. At the same time, since the electronic device 100 only sends the ROI instead of the entire image to the cloud server, it helps to avoid unnecessary privacy leakage, and can achieve the role of protecting the privacy security of oneself and / or others.

[0116] Optionally, if the first input includes voice input or text input, obtaining the processing result of the first ROI includes: inputting the first input and the first ROI into a content generation model to obtain the processing result.

[0117] Taking the above method 200 executed by the electronic device 100 as an example, the processing result can be obtained by generating a model based on the voice input and the ROI input content. In this way, the processing result can be made more in line with the user's expectations, thereby helping to improve the user's human-computer interaction experience. At the same time, since only the ROI in the image is used when obtaining the processing result instead of the entire image, the computing overhead required for the electronic device 100 to obtain the processing result can be reduced, thereby helping to reduce the power consumption of the electronic device.

[0118] Exemplarily, the content generation model may be a multimodal model, such as a visual language model (VLM).

[0119] Exemplarily, the content generation model may be a region convolutional neural networks (RCNN) text content recognition model.

[0120] Exemplarily, the content generation model may be a Transformer model of bootstrapping language-image pre-training (BLIP).

[0121] Exemplarily, the content generation model may be a text to image model or an image to image model, for example, a diffusion transformer (DIT) Transformer model.

[0122] Exemplarily, the content generation model can be set on the electronic device side or the cloud server side based on the size of the model.

[0123] Exemplarily, the VLM may include a visual model (e.g., an image encoder) and a language model (e.g., a text encoder). For example, the VLM may use a pre-trained visual model and a language model, connected with a graph-text input feature alignment module, so that the language model can understand the image features and enter deeper question-answering reasoning.

[0124] Exemplarily, the language model may be a large language model (LLM), the image encoder may be a multilayer perceptron (MLP), and the image-text input feature alignment module may be a Q-Former model. The Q-Former model may be used in a cross-modal dialogue system to improve the interactivity and intelligence of the dialogue by understanding and generating mixed image-text dialogue content.

[0125] For example, taking the multimodal large model as a VLM, the text content corresponding to the first voice input and the ROI input into the VLM can obtain relevant scene understanding information.

[0126] The model of the above output processing result can be a downstream model. The downstream model is not limited to the above content generation model and multimodal model. The downstream model can also be various convolutional neural network (CNN) algorithms. Taking the CNN algorithm as a super-resolution algorithm as an example, when the user's voice input "help me save the photo of the target 5 in front" is obtained, the ROI corresponding to the target 5 can be obtained by inputting the voice input and the image collected by the camera into the target detection model. According to the super-resolution algorithm, the ROI can be super-resolved and clarified by a reasonable multiple to obtain the processed image.

[0127] Optionally, if the first input does not include voice input and does not include text input, obtaining the processing result of the first ROI includes: sending the first ROI to a cloud server; and receiving the processing result determined by the cloud server based on the first ROI.

[0128] Optionally, if the first input does not include voice input and does not include text input, obtaining the processing result of the first ROI includes: inputting the first ROI into a content generation model to obtain the processing result.

[0129] Exemplarily, the content generation model may be a model of generating text from images, or a model of generating images from images.

[0130] Optionally, obtaining the processing result of the first ROI includes: determining the processing result according to the second target semantics included in the first input and the first ROI.

[0131] The first target semantics and the second target semantics may be the same or different.

[0132] Exemplarily, the first input includes a voice input "What is this?" and a user's gesture input (e.g., the index finger points to the left front of the user) when acquiring the voice input, then the first target semantics can be obtained by the index finger pointing, such as the first target semantics can be "left front". The second target semantics can be the text content corresponding to the voice input, such as the second target semantics is "What is this".

[0133] S240, controlling the prompting device to prompt the user with the processing result.

[0134] Optionally, taking the method 200 being executed by the electronic device 100 as an example, before the control prompt device prompts the user with the processing result, the method 200 further includes: receiving the processing result sent by the cloud server.

[0135] Optionally, taking S230 being executed by the cloud server as an example, controlling the prompting device to prompt the user with the processing result includes: sending the processing result to the electronic device so that the electronic device prompts the user with the processing result.

[0136] Exemplarily, taking the processing result as an example, the electronic device 100 can display the picture through the display device 140 .

[0137] Exemplarily, taking the processing result as text content, the electronic device 100 may display the text content through the display device 140; or, the electronic device 100 may convert the text content into voice content and broadcast the voice content to the user through a speaker.

[0138] In the embodiment of the present application, the ROI is determined from the image according to the user's input and the processing result of the ROI is obtained, so that the processing result can be more in line with the user's expectations, thereby helping to improve the user's human-computer interaction experience. At the same time, since only the ROI in the image is used instead of the entire image when obtaining the processing result, the computational overhead required to obtain the processing result can be reduced. In addition, since only the ROI is used instead of the entire image, unnecessary privacy leakage is avoided in the process of obtaining the processing result, and the privacy security of oneself and / or others can be protected.

[0139] Figure 3 A schematic flow chart of a human-computer interaction method 300 provided in an embodiment of the present application is shown. The method 300 may be executed by the electronic device 100 described above. The method 300 includes:

[0140] S301, obtaining a first image and a first voice input of a user.

[0141] Exemplarily, taking the case where method 300 is executed by a vehicle, images captured by a camera of the vehicle and voice signals captured by a microphone may be acquired.

[0142] Optionally, acquiring the first image and the user's first voice input includes: acquiring the first image in response to acquiring the first voice input.

[0143] S302: Determine a first region of interest ROI from a first image according to the first voice input.

[0144] Optionally, determining a first region of interest ROI from the first image according to the first voice input includes: identifying one or more targets in the first image to obtain a recognition result; and determining the first ROI according to the first voice input and the recognition result.

[0145] For example, Figure 4 A schematic diagram of determining ROI provided in an embodiment of the present application is shown.

[0146] like Figure 4 As shown, taking the execution of method 300 by a vehicle as an example, in response to obtaining the voice input "What is the gray object in front?" detected by the microphone in the vehicle cabin, the vehicle can obtain image 3 captured by the camera outside the cabin. The vehicle can recognize image 3, for example, by using an image segmentation algorithm to obtain a recognition result, and the recognition result indicates multiple targets in image 3, for example, targets 1-4. By parsing the text content corresponding to the voice input, it can be obtained that the user's intention is "confirm what the gray object in front is", and the vehicle can confirm that target 4 is a gray object from targets 1-4 based on the user's intention and the recognition result, so that the area where target 4 is located in image 3 (such as Figure 4 The dashed box in the figure was used as the ROI.

[0147] Exemplarily, taking the method 300 executed by a helmet as an example, the helmet may be provided with a camera and a microphone. In response to obtaining the voice input "Is there a risk of falling rocks on the mountain ahead" detected by the microphone on the helmet, the helmet may obtain image 4 captured by the camera. The helmet may recognize image 4 captured by the camera, for example, by obtaining a recognition result through an image segmentation algorithm, and the recognition result indicates multiple targets in image 4, including vehicles, mountains, lane lines, etc. By parsing the text content corresponding to the voice input, it can be obtained that the user's intention is "to confirm whether there is a risk of falling rocks on the mountain ahead", and the vehicle can confirm the mountain from the multiple targets based on the user's intention and the recognition result, so that the area where the mountain is located in image 4 can be used as ROI.

[0148] Exemplarily, the above multimodal model for determining ROI may include a multimodal alignment model and a detection CNN model.

[0149] Exemplarily, according to the first voice input, determining the first region of interest ROI from the first image includes: inputting the first voice input or the text content corresponding to the first voice input into the corresponding encoder of the multimodal alignment model to obtain a text encoding result, and then inputting the text encoding result and the first image into the detection CNN model, so as to obtain the first ROI. The multimodal alignment model and the detection CNN model can understand the semantics and detect the target expressed by the semantics in the first image, so as to obtain the first ROI.

[0150] Exemplarily, the multimodal alignment model is a contrastive language-image pre-training (CLIP) model.

[0151] Exemplarily, according to the first voice input, determining a first region of interest ROI from the first image includes: segmenting the first image to obtain multiple regions; inputting the first voice input and the multiple regions into a multimodal model to obtain the first ROI; or, inputting text content corresponding to the first voice input and the multiple regions into a multimodal model to obtain the first ROI; or, inputting target semantics in the text content corresponding to the first voice input and the multiple regions into a multimodal model to obtain the first ROI.

[0152] S303: Send the first voice input and the first ROI to the cloud server.

[0153] Optionally, sending the first voice input and the first ROI to the cloud server includes: sending the first voice input and the first ROI to the cloud server; or sending text content corresponding to the first voice input and the first ROI to the cloud server.

[0154] Optionally, a content generation model is stored in the cloud server, the input of the content generation model may be multimodal input, such as text content and image (or, voice and image; or, text content, voice and image), and the output may be a processing result.

[0155] Optionally, the content generation model may be a VLM.

[0156] For example, Figure 4 Taking the scene shown as an example, the text content input to VLM is "What is the gray object in front?" and the image input to VLM is the above ROI, the output of VLM can be the text content "The gray object in front is a rock. The distance between the rock and the vehicle is close, please pay attention to driving safety."

[0157] Exemplarily, the content generation model may also be other models, such as CNN.

[0158] S304, receiving a processing result for the first voice input and the first ROI sent by the cloud server.

[0159] Exemplarily, the processing result includes the text content output by the above-mentioned VLM.

[0160] Optionally, the cloud server may determine the type of the processing result based on text content corresponding to the first voice input.

[0161] Exemplarily, the first voice input is “What’s in front? Please introduce it to me”, and at this time the content generation model can output text content related to the ROI.

[0162] Exemplarily, the first voice input is "What's in front? Help me turn it into Pikachu style." At this time, the content generation model can use the image-to-image technology to output a picture after the ROI is style-converted.

[0163] Optionally, the content generation model may integrate an optical character recognition (OCR) function, so that the content generation model may use the OCR recognition result for the ROI as the processing result. Alternatively, the content generation model may also include a text to text model, and the content generation model may input the OCR recognition result into the text to text model to obtain the processing result.

[0164] In the above S303-S304, the content generation model is located in the cloud server as an example, and the embodiments of the present application are not limited to this. The multimodal large model can also be located in an electronic device. The electronic device can input the first voice and the first ROI into the content generation model to obtain the processing result.

[0165] S305, controlling the prompting device to prompt the processing result.

[0166] For example, Figure 4 Take the scenario shown in the figure as an example. After obtaining the processing results sent by the cloud server, the vehicle can first convert the text content into voice information, and then send a voice message through the speaker in the cockpit: "The gray object in front is a rock. The distance between the rock and the vehicle is close. Please pay attention to driving safety."

[0167] In an embodiment of the present application, the ROI can be determined from the image based on the user's voice input, and then the voice input and the ROI can be sent to the cloud server. In this way, when the cloud server uses the content generation model to generate content, the amount of calculation of the content generation model can be reduced, thereby helping to reduce the time delay between the electronic device obtaining the voice input and the image and the electronic device prompting the user with the processing result, which helps to improve the user experience. At the same time, by determining the ROI from the image, the user can obtain a processing result that better meets the needs, and can also avoid unnecessary privacy leakage caused by transmitting all the images collected by the camera to the cloud server, which can achieve the purpose of protecting the privacy of oneself and / or others.

[0168] Figure 5 A schematic flow chart of a human-computer interaction method 500 provided in an embodiment of the present application is shown. The method 500 may be executed by the electronic device 100 described above. The method 500 includes:

[0169] S501: Acquire a first image and a first input from a user.

[0170] Exemplarily, the first input is key input, voice input or gesture input.

[0171] S502: When the first input satisfies a preset condition, determine a first ROI from the first image according to the sight direction of the user.

[0172] Optionally, when the first input meets a preset condition, determining the first ROI from the first image according to the user's line of sight direction includes: when the first input is a key input by the user, determining the first ROI from the first image according to the user's line of sight direction.

[0173] Exemplarily, the key input may be an input from a user on a physical key or a virtual key.

[0174] Optionally, the first input is a key input, and acquiring the first image and the first input includes: determining the first ROI according to the user's sight direction after a preset time period from when the user triggers the key input.

[0175] In an embodiment of the present application, the user's line of sight may be on the button when triggering the key input. By determining the first ROI in the first image through the line of sight obtained after a preset time from the time when the user triggers the key input, the obtained ROI can be more in line with the user's needs, thereby improving the accuracy of the processing results and helping to improve the user's human-computer interaction experience.

[0176] Optionally, before determining the first ROI from the first image according to the user's line of sight direction, the method 500 also includes: obtaining a voice input from the user; wherein, when the first input satisfies a preset condition, determining the first ROI from the first image according to the user's line of sight direction, including: when the text content corresponding to the voice input does not include text content related to the orientation and does not include text content related to the attributes of the target, determining the first ROI from the first image according to the user's line of sight direction.

[0177] For example, taking the target as a vehicle, the attributes of the target include but are not limited to the color, brand, size, type (e.g., sedan, SUV, truck), etc. of the vehicle.

[0178] Exemplarily, the voice input is “What is this?” The electronic device may determine that the voice input does not include text content related to the orientation and does not include text content related to the attributes of the target, so that the first ROI may be determined according to the user's sight direction.

[0179] Optionally, obtaining the sight line direction of the user includes: obtaining a facial image of the user; and determining the sight line direction of the user based on the facial image.

[0180] Optionally, obtaining the user's sight direction includes: determining the user's sight direction through eye tracking.

[0181] For example, the electronic device is a vehicle. When the user's voice input "What is this?" is detected by the microphone in the cabin of the vehicle, the vehicle can obtain an image 5 captured by the camera outside the cabin and an image 6 captured by the camera inside the cabin, and the image 6 includes the user's face image when the voice input is triggered. The vehicle can determine the user's line of sight based on the image 6. Thus, the ROI can be determined from the image 5 based on the user's line of sight.

[0182] For example, the electronic device is a smart headset. When the user's voice input "What is this?" is detected through the microphone of the smart headset, the smart headset can control the camera on the smart headset to capture image 7. The smart headset can determine the ROI from the image 7 based on the image 7 and the relationship between the image captured by the pre-calibrated camera and the user's line of sight.

[0183] Exemplarily, the user's sight direction can be calibrated with the camera that captures the first image. After obtaining the user's sight direction, a selection box centered on the gaze point can be selected based on the user's sight direction (eg, eye focus area), thereby determining the ROI.

[0184] For example, taking the first input including a user's gesture input as an example, the user's finger pointing can be calibrated with the camera that captures the first image. After obtaining the user's finger pointing, a selection box can be selected based on the user's finger pointing point in the first image as the center, thereby determining the ROI.

[0185] Optionally, before determining the first ROI from the first image according to the user's line of sight direction, the method 500 also includes: obtaining the user's gesture; wherein, when the first input meets a preset condition, determining the first ROI from the first image according to the user's line of sight direction, including: when the user's gesture is a preset gesture, determining the first ROI from the first image according to the user's line of sight direction.

[0186] Exemplarily, the preset gesture may be an OK gesture.

[0187] In the above S502, when the first input meets the preset condition, the first ROI is determined from the first image according to the user's sight direction. In the embodiment of the present application, in S501, only the first image may be acquired, so that the first ROI can be determined directly from the first image based on the user's sight direction.

[0188] S503: Send the first ROI to the cloud server.

[0189] Optionally, the first input is voice input or text input, and sending the first ROI to the cloud server includes: sending the voice input and the first ROI to the cloud server; or sending the text input and the first ROI to the cloud server.

[0190] S504: Receive a processing result for the first ROI sent by the cloud server.

[0191] Optionally, the first input is voice input or text input, and receiving a processing result for the first ROI sent by the cloud server includes: receiving a processing result for the voice input and the first ROI sent by the cloud server; or receiving a processing result for the text input and the first ROI sent by the cloud server.

[0192] S505, controlling the prompting device to prompt the processing result.

[0193] The above S503-S505 can refer to the description of the above S303-S305, which will not be repeated here.

[0194] Figure 6 A schematic flow chart of a human-computer interaction method 600 provided in an embodiment of the present application is shown. The method 600 may be executed by the electronic device 100 described above. The method 600 includes:

[0195] S601: Acquire a first image and a first input from a user.

[0196] Exemplarily, the first input is voice input or gesture input.

[0197] S602: Determine a first direction according to the first input.

[0198] Optionally, the first input is a voice input, and determining the first direction according to the first input includes: determining the first direction according to text content related to the direction in text content corresponding to the voice input.

[0199] For example, Figure 7 A schematic diagram of dividing a first image into regions provided in an embodiment of the present application is shown. The first image is divided into regions 1 to 9. Table 1 shows a correspondence between a direction and a corresponding region.

[0200] Table 1

[0201] direction area Ahead Area 5 Left side (or left side) Region 1, Region 4 and Region 7 Right (or right side) Region 3, Region 6 and Region 9 above (or above) Regions 1-3 Below (or below) Zones 7-9 Top left corner Region 1 Top right corner Area 3 Lower left corner Area 7 Bottom right corner Area 9 … …

[0202] Exemplarily, the above region c may include region 1 , region 4 , and region 7 in image 2 , and region d may include region 3 , region 6 , and region 9 in image 2 .

[0203] The above division is merely illustrative, and the present application embodiment does not specifically limit this. For example, the upper left corner may also correspond to area 1 and area 4. For example, the front may also correspond to areas 4-6.

[0204] above Figure 7In the example of dividing an image into 9 regions, the present application embodiment does not specifically limit the number of regions after division. For example, the image can be divided into 2 or more regions.

[0205] above Figure 7 In the description, the non-overlapping of the divided regions is used as an example, and the present application embodiment does not specifically limit this. Two adjacent regions may also partially overlap.

[0206] Optionally, the first input is a first gesture input of the user, and determining the first direction according to the first input includes: determining the first direction according to a finger direction corresponding to the first gesture.

[0207] Exemplarily, the first gesture of the user is an input of the user raising his index finger, and the direction of the index finger can be used as the first direction.

[0208] Optionally, the first input is a first gesture input of the user, and determining the first direction according to the first input includes: determining the first direction according to the first gesture and a mapping relationship.

[0209] Exemplarily, Table 2 shows the above mapping relationship.

[0210] Table 2

[0211] gesture direction Gesture A left Gesture B right Gesture C above … …

[0212] For example, when the user's gesture A is acquired, it can be determined that the corresponding direction is the left side. At this time, the first ROI can be determined from the first image in combination with the corresponding relationship shown in Table 1 above.

[0213] S603: Determine a first ROI from the first image according to the first direction.

[0214] Exemplarily, when the voice input is “what is on the left?”, it can be determined that the text content corresponding to the voice input includes the direction “left.” Based on the corresponding relationship shown in Table 1 above, region 1, region 4, and region 7 can be extracted from the first image as the first ROI.

[0215] Optionally, determining the first ROI from the first image includes: determining a second ROI according to the first direction; when the second ROI includes a part of the first target, acquiring a region where the first target is located from the first image; and determining the first ROI according to the region where the first ROI is located and the second ROI.

[0216] For example, Figure 8Another schematic diagram of dividing the first image into regions is shown. The first image includes target a. When the voice input is "What's on the left?", it can be determined that the text content corresponding to the voice input includes the direction "left". Based on the correspondence shown in Table 1 above, region 1, region 4 and region 7 can be extracted from the first image as the second ROI. When determining the part including target a in the second ROI, the region where the first target is located can be obtained from the first image through an image segmentation algorithm, for example, region 4, region 5, region 7 and region 8. At this time, the electronic device can use the union of the region where the first target is located and the region corresponding to the second ROI as the first ROI, that is, the first ROI includes region 1, region 4, region 5, region 7 and region 8.

[0217] S604: Send the first ROI to the cloud server.

[0218] Optionally, the first input is voice input or text input, and sending the first ROI to the cloud server includes: sending the voice input and the first ROI to the cloud server; or sending the text input and the first ROI to the cloud server.

[0219] S605: Receive a processing result for the first ROI sent by the cloud server.

[0220] Optionally, the first input is voice input or text input, and receiving a processing result for the first ROI sent by the cloud server includes: receiving a processing result for the voice input and the first ROI sent by the cloud server; or receiving a processing result for the text input and the first ROI sent by the cloud server.

[0221] S606, controlling the prompting device to prompt the processing result.

[0222] The above S604-S606 can refer to the description of the above S303-S305, which will not be repeated here.

[0223] Fig. 9 A schematic block diagram of a human-computer interaction device 900 provided in an embodiment of the present application is shown. The device 900 includes: an acquisition unit 910, which is used to acquire a first image and a first input of a user; a determination unit 920, which is used to determine a first region of interest ROI from the first image according to the first input; the acquisition unit 910 is also used to acquire a processing result of the ROI; and a control unit 930, which is used to control a prompting device to prompt the user with the processing result.

[0224] Optionally, the first input is a first voice input or a first text input, and the determination unit 920 is used to: input the first input and the first image into the target detection model to obtain a first ROI, and the output of the target detection model changes based on the change of the first target semantics in the first input.

[0225] Optionally, the first input is a first voice input, and the determination unit 920 is used to: determine the first text content according to the first voice input; and determine the first ROI from the first image according to the first text content.

[0226] Optionally, the determination unit 920 is used to: when the first text content includes second text content related to the direction, determine the first ROI according to the second text content; or, when the first text content includes third text content related to the attribute of the first target, determine the first ROI according to the third text content, and the first ROI includes the first target.

[0227] Optionally, the determination unit 920 is used to: determine multiple ROIs from the first image according to the third text content; and determine the first ROI from the multiple ROIs according to the user's line of sight when the first voice input is obtained.

[0228] Optionally, the determination unit 920 is used to: determine multiple ROIs from the first image according to the third text content; the control unit is used to control the prompt device to prompt the user to select one or more ROIs from the multiple ROIs; the determination unit is used to determine the first ROI in response to the user's second input, and the second input instructs the user to select the first ROI.

[0229] Optionally, the acquisition unit 910 is used to acquire the user's sight direction when the first text content does not include direction-related text content and does not include text content related to the attributes of the target; the determination unit is used to determine the first ROI according to the sight direction.

[0230] Optionally, the first input is a first voice input, and the acquisition unit 910 is further used to acquire a first gesture of the user when the user triggers the first voice input; the determination unit is used to determine the first ROI according to the first gesture.

[0231] Optionally, the device 900 further includes a sending unit, which is used to send the first input and the first ROI to the cloud server; and an acquiring unit, which is used to receive a processing result determined by the cloud server based on the first input and the first ROI.

[0232] Optionally, the acquisition unit 910 is configured to: input the first input and the first ROI into a content generation model to obtain a processing result.

[0233] Optionally, the first input is a second gesture of the user, and the determination unit 920 is configured to determine the first ROI according to a direction of a finger in the second gesture.

[0234] Optionally, the first input is a third gesture of the user, and the determination unit 920 is used to determine the first ROI according to the sight direction of the user when the third gesture is a preset gesture.

[0235] Optionally, the first input is an input to a first key, and the acquisition unit 910 is further configured to acquire a sight line direction of the user when inputting the first key; and the determination unit is configured to determine the first ROI according to the sight line direction.

[0236] Optionally, the device 900 further includes a sending unit, which is used to send the first ROI to the cloud server; and an acquiring unit, which is used to receive a processing result determined by the cloud server based on the first ROI.

[0237] Optionally, the acquisition unit 910 is configured to: input the first ROI into a content generation model to obtain a processing result.

[0238] Optionally, the determining unit 920 is further configured to determine a processing result according to the second target semantics and the first ROI included in the first input.

[0239] An embodiment of the present application further provides a human-computer interaction device, which includes: a memory for storing a computer program; and a processor for executing the computer program stored in the memory, so that the device executes any one of the methods described in the first aspect above.

[0240] An embodiment of the present application further provides an electronic device, which includes the above-mentioned human-computer interaction device 900.

[0241] The embodiment of the present application also provides a cloud server, which includes the above-mentioned human-computer interaction device 900.

[0242] An embodiment of the present application further provides a computer program product, which includes: a computer program code, and when the computer program code is executed on a computer, the computer executes the human-computer interaction method in the above embodiment.

[0243] An embodiment of the present application further provides a computer-readable medium, wherein the computer-readable medium stores a program code. When the computer program code runs on a computer, the computer executes the human-computer interaction method in the above embodiment.

[0244] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0245] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0246] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0247] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0248] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0249] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0250] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A human-computer interaction method, characterized in that: include: Acquire a first image and a first input from a user; Determine a first region of interest ROI from the first image according to the first input; Obtaining a processing result of the first ROI; The control prompting device prompts the user of the processing result.

2. The method according to claim 1, characterized in that The first input is a first voice input or a first text input, and determining a first ROI from the first image according to the first input includes: The first input and the first image are input into an object detection model to obtain the first ROI, and an output of the object detection model changes based on a change in semantics of a first object in the first input.

3. The method according to claim 1 or 2, characterized in that: The first input is a first speech input, and determining a first ROI from the first image according to the first input includes: Determining first text content according to the first voice input; The first ROI is determined from the first image according to the first text content.

4. The method according to claim 3, characterized in that The determining the first ROI from the first image according to the first text content includes: When the first text content includes second text content related to direction, determining the first ROI according to the second text content; or, When the first text content includes third text content related to the attribute of the first target, the first ROI is determined according to the third text content, and the first ROI includes the first target.

5. The method according to claim 4, characterized in that The determining the first ROI according to the third text content includes: Determining a plurality of ROIs from the first image according to the third text content; The first ROI is determined from the multiple ROIs according to the user's sight direction when the first voice input is acquired.

6. The method according to claim 4, characterized in that The determining the first ROI according to the third text content includes: Determining a plurality of ROIs from the first image according to the third text content; Controlling the prompting device to prompt the user to select one or more ROIs from the multiple ROIs; The first ROI is determined in response to detecting a second input from a user, the second input indicating that the user selected the first ROI.

7. The method according to claim 3, characterized in that The determining the first ROI from the first image according to the first text content includes: When the first text content does not include text content related to the direction and does not include text content related to the attribute of the target, obtaining the sight direction of the user; The first ROI is determined according to the sight line direction.

8. The method according to claim 1 or 2, characterized in that: The first input is a first voice input, and the method includes: Acquire a first gesture of the user when the user triggers the first voice input; Wherein, determining a first ROI from the first image according to the first input includes: The first ROI is determined according to the first gesture.

9. The method according to any one of claims 2 to 8, characterized in that The obtaining the processing result of the first ROI includes: Sending the first input and the first ROI to a cloud server; The processing result determined by the cloud server based on the first input and the first ROI is received.

10. The method according to any one of claims 2 to 8, characterized in that The obtaining the processing result of the first ROI includes: The first input and the first ROI are input into a content generation model to obtain the processing result.

11. The method according to claim 1, characterized in that: The first input is a second gesture of the user, and determining a first ROI from the first image according to the first input includes: The first ROI is determined according to the direction of the finger in the second gesture.

12. The method according to claim 1, characterized in that The first input is a third gesture of the user, and determining a first ROI from the first image according to the first input includes: When the third gesture is a preset gesture, the first ROI is determined according to the sight direction of the user.

13. The method according to claim 1, characterized in that The first input is an input for a first button, and determining a first ROI from the first image according to the first input includes: Acquire the user's sight direction when inputting the first button; The first ROI is determined according to the sight line direction.

14. The method according to any one of claims 11 to 13, characterized in that The obtaining the processing result of the first ROI includes: Sending the first ROI to a cloud server; The processing result determined by the cloud server based on the first ROI is received.

15. The method according to any one of claims 11 to 13, characterized in that The obtaining the processing result of the first ROI includes: The first ROI is input into a content generation model to obtain the processing result.

16. The method according to any one of claims 1 to 15, characterized in that The obtaining the processing result of the first ROI includes: The processing result is determined according to the second target semantics included in the first input and the first ROI.

17. A human-computer interaction device, characterized in that: The method comprises a unit or a module for executing the method according to any one of claims 1 to 16.

18. A human-computer interaction device, characterized in that: include: Memory for storing computer programs; A processor, configured to execute the computer program stored in the memory, so that the apparatus performs the method according to any one of claims 1 to 16.

19. An electronic device, characterized in that: The electronic device includes the human-computer interaction device shown in claim 17 or 18.

20. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a computer, the method according to any one of claims 1 to 16 is implemented.

21. A computer program product, characterized in that The computer program product comprises: a computer program code, and when the computer program code is run on a computer, the method according to any one of claims 1 to 16 is implemented.