Man-machine interaction system and method
By configuring an image input source on a smart TV to acquire and process image data, the problem of smart TVs not supporting image input is solved, enabling smart TVs to intelligently process and respond to images, expanding application scenarios, and improving user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN SKYWORTH DISPLAY TECH CO LTD
- Filing Date
- 2025-11-17
- Publication Date
- 2026-04-21
AI Technical Summary
Existing smart TVs do not support image input and processing, resulting in outdated human-computer interaction functions and a poor user experience.
By configuring an image input source on a smart TV, image data is acquired, and content recognition is performed in the interactive interface of the smart agent application. The corresponding smart agent service is then invoked to generate and output response information.
It enables smart TVs to intelligently process and respond to images, breaking through the limitations of image input, expanding application scenarios, and improving user experience.
Smart Images

Figure CN121900611A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a human-computer interaction system and method. Background Technology
[0002] With the development of the internet, the application of artificial intelligence (AI) technology is becoming increasingly widespread. The input and output forms of AI applications are becoming more diverse. The human-computer interaction functions of intelligent applications allow users to search for content or operate smart devices through voice, text, images, and other means, bringing greater convenience to users. Currently, intelligent applications supporting human-computer interaction functions are ubiquitous on smart electronic devices such as mobile phones, tablets, and computers.
[0003] However, existing smart TVs still only support voice or text input, not image input. This has resulted in the lagging application of artificial intelligence technology in smart TVs, making it impossible for users to recognize and process image content, leading to a poor user experience. Summary of the Invention
[0004] This application provides a human-computer interaction system and method to solve the technical problem that existing first devices (smart TVs) do not support image input and image processing, resulting in outdated human-computer interaction functions and poor user experience on smart TVs.
[0005] In a first aspect, this application provides a human-computer interaction system, the system comprising: an intelligent agent server, a first device configured with an image input source, the first device integrating an intelligent agent application, and the intelligent agent server integrating at least one intelligent agent service; The first device acquires image data of the target object through the image input source; In the interactive interface of the intelligent agent application of the first device, the image data is used as dialogue input information so that the intelligent agent application can perform content recognition on the image data and call the corresponding intelligent agent service based on the content recognition result. The intelligent agent service generates response information in response to the image data; The first device outputs the response information.
[0006] In one possible implementation, the system further includes: a control terminal for the first device, wherein the control terminal integrates an image acquisition component; The control terminal, in response to an image acquisition operation, acquires image data through an integrated image acquisition component and sends the image data to the first device.
[0007] In one possible implementation, the system further includes: an image acquisition component connected to the first device; The first device controls an image acquisition component connected to the first device to acquire image data and obtain the image data.
[0008] In one possible implementation, the system further includes: a second device communicatively connected to the first device; The second device acquires image data in response to an image acquisition operation; or selects image data from a local image library in response to an image selection operation. The image data is sent to the first device.
[0009] In one possible implementation, the first device integrates a screenshot function component; The first device uses an integrated screenshot function component to capture the currently displayed screen and generate image data.
[0010] Secondly, this application provides a human-computer interaction method applied to a first device configured with an image input source, the method comprising: Image data is acquired through the image input source; In the interactive interface of the intelligent agent application of the first device, the image data is used as dialogue input information so that the intelligent agent application can perform content recognition on the image data and call the corresponding intelligent agent service based on the content recognition result. The response information generated by the intelligent agent service in response to the image data is output through the first device.
[0011] In one possible implementation, the image input source includes the control terminal of the first device, and the control terminal integrates an image acquisition component; The step of acquiring image data through the image input source includes: The control terminal receives image data sent to the first device, wherein the image data is image data acquired by the control terminal in response to an image acquisition operation through an image acquisition component integrated on the control terminal.
[0012] In one possible implementation, the image input source includes an image acquisition component connected to the first device; The step of acquiring image data through the image input source includes: In response to an image acquisition operation triggered on the first device, the image acquisition component connected to the first device is controlled to acquire image data and obtain the image data.
[0013] In one possible implementation, the image input source includes a second device that is communicatively connected to the first device; The step of acquiring image data through the image input source includes: The device receives image data sent from the second device to the first device, wherein the image data is image data acquired by the second device in response to an image acquisition operation, or image data selected by the second device from a local image library in response to an image selection operation.
[0014] In one possible implementation, the image input source includes a screenshot function component integrated into the first device; The step of acquiring image data through the image input source includes: Obtain image data generated by the first device taking a screenshot of the currently displayed screen through the screenshot function component.
[0015] In one possible implementation, outputting the response information generated by the intelligent agent service in response to the image data via the first device includes: The response information generated by the intelligent agent service in response to the image data is output in a modality that combines or integrates any one or at least two of the following forms: text, graphics, voice broadcast, or video.
[0016] Thirdly, this application provides a human-computer interaction device, applied to a first device configured with an image input source, the device comprising: An image acquisition module is used to acquire image data from the image input source; The image processing module is used to take the image data as dialogue input information in the interactive interface of the intelligent agent application of the first device, so that the intelligent agent application can perform content recognition on the image data and call the corresponding intelligent agent service based on the content recognition result. The image output module is used to output the response information generated by the intelligent agent service in response to the image data through the first device.
[0017] In one possible implementation, the image input source includes the control terminal of the first device, and the control terminal integrates an image acquisition component; The image acquisition module is specifically used for: The control terminal receives image data sent to the first device, wherein the image data is image data acquired by the control terminal in response to an image acquisition operation through an image acquisition component integrated on the control terminal.
[0018] In one possible implementation, the image input source includes an image acquisition component connected to the first device; The image acquisition module is specifically used for: In response to an image acquisition operation triggered on the first device, the image acquisition component connected to the first device is controlled to acquire image data and obtain the image data.
[0019] In one possible implementation, the image input source includes a second device that is communicatively connected to the first device; The image acquisition module is specifically used for: The device receives image data sent from the second device to the first device, wherein the image data is image data acquired by the second device in response to an image acquisition operation, or image data selected by the second device from a local image library in response to an image selection operation.
[0020] In one possible implementation, the image input source includes a screenshot function component integrated into the first device; The image acquisition module is specifically used for: Obtain image data generated by the first device taking a screenshot of the currently displayed screen through the screenshot function component.
[0021] In one possible implementation, the image output module is specifically used for: The response information generated by the intelligent agent service in response to the image data is output in a modality that combines or integrates any one or at least two of the following forms: text, graphics, voice broadcast, or video.
[0022] Fourthly, this application provides an electronic device, including: a processor and a memory, wherein the processor is configured to execute a human-computer interaction program stored in the memory to implement the human-computer interaction method described in any one of the first aspects.
[0023] Fifthly, this application provides a storage medium storing one or more programs that can be executed by one or more processors to implement the human-computer interaction method described in any of the first aspects.
[0024] Compared with the prior art, the technical solution provided in this application has the following advantages: The method provided in this application is applied to a first device configured with an image input source. Image data is acquired through the image input source. In the interactive interface of the intelligent agent application of the first device, the image data is used as dialogue input information. The intelligent agent application performs content recognition on the image data and calls the corresponding intelligent agent service based on the content recognition result. The response information generated by the intelligent agent service in response to the image data is then output through the first device. By configuring an image input source for the first device, the lack of image input capability in the first device (such as a smart TV) is compensated, realizing the first device's intelligent image processing and response information output capability. This breaks through the limitation that the first device itself does not have image input capability, expanding a new paradigm of vision-based AI interaction, thereby broadening the application scenarios of the first device and significantly improving the user experience through multimodal intelligent interaction. Attached Figure Description
[0025] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0026] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0027] One or more embodiments are illustrated by way of example with reference numerals in the accompanying drawings. These illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are denoted as similar elements. Unless otherwise stated, the figures in the drawings are not to be limited by scale.
[0028] Figure 1 A structural block diagram of a human-computer interaction system provided in an embodiment of this application; Figure 2 A flowchart illustrating an embodiment of a human-computer interaction method provided in this application; Figure 3 A flowchart illustrating another embodiment of the human-computer interaction method provided in this application; Figure 4 This application provides a schematic diagram of a system framework where the image input source is a remote control. Figure 5 A flowchart illustrating another embodiment of a human-computer interaction method provided in this application; Figure 6A flowchart illustrating another embodiment of the human-computer interaction method provided in this application; Figure 7 A structural block diagram of a human-computer interaction device provided in an embodiment of this application; Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0029] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0030] The following disclosure provides numerous different embodiments or examples for implementing various structures of this application. To simplify the disclosure, specific examples of components and arrangements are described below. These are merely examples and are not intended to limit the scope of this application. Furthermore, reference numerals and / or letters may be repeated in different examples. Such repetition is for simplification and clarity and does not in itself indicate a relationship between the various embodiments and / or arrangements discussed.
[0031] To address the technical problem that existing smart TVs lack image input and processing capabilities, resulting in outdated human-computer interaction functions and poor user experience, this application provides a human-computer interaction system and method. Through multi-source image acquisition, it enables the smart TV to input images. Automated intelligent processing and response information output enable the TV to automatically recognize and respond to images, overcoming the input limitations of traditional TVs and optimizing the convenience and stability of image input. Simultaneously, leveraging the large screen advantage of TVs, it redefines the visual interaction capabilities of TVs, covering various application scenarios, enhancing user experience, and further realizing a high level of intelligence for TV devices.
[0032] Figure 1 The present application provides a structural block diagram of a human-computer interaction system. The system includes: an intelligent agent service terminal 10, a first device 20 configured with an image input source 21, an intelligent agent application 22 integrated on the first device, and at least one intelligent agent service 1111 integrated on the intelligent agent service terminal 10.
[0033] In this process, the first device 20 acquires image data of the target object through image input source 21. In the interactive interface of the intelligent agent application 22 on the first device 20, the image data is used as dialogue input information. The intelligent agent application 22 performs content recognition on the image data and, based on the content recognition result, calls the corresponding intelligent agent service 11, outputting the response information generated by the intelligent agent service 11 based on the image data. Meanwhile, the intelligent agent service 11 generates response information based on the image data.
[0034] The human-computer interaction system of this application embodiment also includes a control terminal 50 of the first device, and the control terminal integrates an image acquisition component 51. In this case, the image input source 21 of the first device 20 is the control terminal of the first device 20, such as a remote control. Specifically, under this image input source 21, the control terminal of the first device 20 responds to the image acquisition operation by acquiring image data through the image acquisition component integrated on the control terminal, and sends the image data to the first device 20.
[0035] The human-computer interaction system of this application embodiment also includes an image acquisition component 40 connected to the first device. In this case, the image input source 21 of the first device is the image acquisition component connected to the first device, for example, an external camera connected to the first device 20 (television). Under this image input source 21, the first device 20 controls the image acquisition component connected to the first device to acquire image data and obtain the image data.
[0036] The human-computer interaction system of this application embodiment also includes a second device 30 communicatively connected to the first device. In this case, the image input source 21 of the first device is the second device 30 communicatively connected to the first device, such as a mobile phone connected to the first device 20 (television). Under this image input source 21, the second device 30 acquires image data in response to an image acquisition operation; or, in response to an image selection operation, selects image data from a local image library; and sends the image data to the first device 20.
[0037] The human-computer interaction system in this embodiment of the application also includes a screenshot function component 211 integrated into the first device. In this case, the image input source 21 of the first device is the screenshot function component 211 integrated into the first device itself, such as a screenshot component on a television. With this image input source, the first device 20 uses the integrated screenshot function component 211 to capture the currently displayed screen and generate image data.
[0038] pass Figure 1The detailed description of the illustrated embodiments shows that the human-computer interaction system provided in this application, based on multi-source image input, intelligent agent collaboration, and cross-device linkage design, completely overcomes the pain points of traditional television's limitations in image input and application scenarios. Furthermore, it redefines the full-scene capabilities of television's human-computer interaction, covering multiple application scenarios, and, based on multi-device collaborative linkage, activates the unique value of smart TV human-computer interaction systems in home scenarios.
[0039] Figure 2 A flowchart illustrating an embodiment of a human-computer interaction method provided in this application is applied to the above-mentioned method. Figure 1 The human-computer interaction system shown includes the following steps: The first device 20 can refer to an electronic device capable of applying the human-computer interaction method provided in the embodiments of this application. Essentially, it is the executing entity of the human-computer interaction method provided in the embodiments of this application, and also the core interaction device of the human-computer interaction method provided in the embodiments of this application. For example, the first device 20 can be an electronic device such as a smart box or television with AI processing capabilities; this embodiment of the application does not limit this. For ease of description, a smart television (the first device) will be used as an example for explanation below.
[0040] Step 201: Obtain image data through the image input source.
[0041] The image input source is a hardware device and software component used to acquire image data. For example, the image input source can be one or more of the following: an integrated remote control with a camera, a USB external camera, a standalone wireless camera, a mobile phone, or a screenshot function software component integrated within the first device. This is merely an example, and the embodiments of this application do not impose limitations.
[0042] Image data can refer to digital image files acquired by an image input source, containing visual information that the user needs AI to process (such as homework questions, medicine boxes, food ingredients, etc.). In addition, image data may also include shooting parameter information (such as resolution and timestamp).
[0043] Since the image input source can be various hardware devices and software components capable of acquiring image data, the specific methods of acquiring image data and the interaction methods with the first device 20 differ depending on the hardware device and software component. For example, when the image input source is a remote control, the first device 20 can directly obtain the image uploaded by the remote control. The remote control needs to respond to the user clicking the shutter button, controlling the camera integrated on the remote control to perform an image acquisition operation, thereby acquiring the corresponding image data. As another example, when the image input source is an external camera, the first device 20 needs to control the camera connected to the first device 20 to perform an image acquisition operation based on the user-triggered image acquisition operation, that is, to execute a shutter command and acquire the corresponding image data.
[0044] Therefore, the way users interact with the first device 20 differs depending on the type of image input source, meaning the specific application scenarios differ, and the way the first device 20 collects and acquires image data of the target object specified by the user also differs. Here, we will only use the screenshot function component 211 integrated into the first device itself as an example for explanation. The specific methods for collecting and acquiring image data corresponding to other image input sources will be explained in the relevant embodiments below.
[0045] The aforementioned target object can refer to a specific thing or information carrier that the user wants to process through the first device 20 and that needs to be collected through an image input source; essentially, it is the core content of the image data. For example, the target object can refer to a specific item in the physical world, such as a student's homework sheet, an elderly person's medicine box, or kitchen ingredients. The core function of the target object is to serve as a carrier of the user's intention to interact with artificial intelligence. By collecting an image of the target object, the user conveys the need to process relevant information about that object to the artificial intelligence application on the first device 20. For example, if a user captures a screenshot of a TV series, they may want the intelligent agent application 22 to identify the actors' names. This is merely an example, and the embodiments of this application do not impose limitations on this.
[0046] In one embodiment, in the application scenario where the image input source of the first device 20 is its own integrated screenshot function component 211, the specific implementation method for obtaining image data through the image input source is as follows: obtain the image data generated by the first device 20 taking a screenshot of the current display screen through the screenshot function component 211.
[0047] The screenshot function component 211 can refer to an internal module integrated into the first device 20 (such as a smart TV), used to capture the real-time image displayed on the screen of the first device itself and generate digital image data, i.e., a screenshot, as a special image input source. Its core function is to convert the virtual content (such as the interface or screen) displayed on the first device 20 itself into image data that can be processed by the AI agent, without relying on external devices (such as remote control cameras or mobile phones), enabling rapid input of target objects within the device. For example, the screenshot function component 211 captures the target object currently displayed on the screen of the first device 20, generates image data, and transmits it to the intelligent agent application 22 as dialogue input information for the intelligent agent application. The aforementioned target object is the target object within the first device 20 captured by the screenshot function component 211; therefore, the screenshot function component 211 can be an internal image input channel built into the first device.
[0048] For example, a user triggers a screenshot command by operating the control terminal of the smart TV, i.e., the remote control. The remote control sends the screenshot command triggered by the user to the smart TV through a signal communicating with the TV. Based on the received screenshot command, the smart TV uses its own screenshot function component 211 to take a screenshot of the current display interface of the smart TV and generate a corresponding screenshot image. Then, the screenshot image generated by its own screenshot function component 211 is temporarily stored in the cache area for the smart agent to process the screenshot image.
[0049] In the aforementioned application scenario of the screenshot function component 211 on the first device, for the content displayed on the TV itself, there is no need to operate external devices. The image can be captured directly through the screenshot component, simplifying the interaction process of the target object and bringing great convenience to the user.
[0050] Step 202: In the interactive interface of the intelligent agent application of the first device, the image data is used as dialogue input information so that the intelligent agent application can perform content recognition on the image data and call the corresponding intelligent agent service based on the content recognition result.
[0051] The intelligent agent application 22 can refer to the AI interaction program running on the first device 20 (smart TV), which has the ability to receive input information, invoke AI capabilities, and generate responses, and serves as the entry point for users to interact with AI services.
[0052] The interactive interface refers to the visual interface displayed on a TV screen for an intelligent application, including an image upload portal, dialogue input boxes, and response display areas, supporting user operations. For example, it allows users to click the photo button and view AI responses on the interface.
[0053] Dialogue input information can refer to the content that the user inputs to the intelligent agent application. For example, dialogue input information can refer to the image data obtained in step 101 of the above embodiment of this application. This image data, which is different from traditional text / voice input, serves as the original basis for AI processing of the smart TV.
[0054] The intelligent agent service 11 can refer to specialized AI service modules targeting specific application scenarios, such as homework tutoring agents, health query agents, and food recognition agents. These different intelligent agent service 11 modules can be automatically invoked by the intelligent agent application based on the content recognition results of the image data. For example, if the image data in the dialogue input information is recognized as a test paper, the homework tutoring agent can be invoked; if the recognized image data is a medicine box, the health query agent can be invoked. This is merely an example, and the embodiments of this application do not impose any limitations.
[0055] In one embodiment, in response to the first device 20 acquiring image data, the acquired image data is directly input into the dialog input box in the interactive interface of the first device's intelligent agent application 22. When the user confirms the use of the image data as input, the intelligent agent performs preliminary content recognition on the image data and calls the corresponding intelligent agent service 11 to process the image data based on the content recognition result.
[0056] For example, a user controls the first device 20 via remote control to take a screenshot of the current display interface (e.g., a portrait of a celebrity) and generate a screenshot image. Upon detecting that the first device has generated a screenshot image, the screenshot image is automatically input into the human-computer interaction interface of the first device's intelligent agent application without user input. If the user confirms the screenshot image as image data by clicking to confirm or through other means, the intelligent agent application performs preliminary content recognition on the image data. If the image data is identified as a portrait, the intelligent agent service 11 corresponding to the portrait is invoked to process the image data.
[0057] Furthermore, when the intelligent application performs preliminary recognition of image data and identifies image data from different scenes, such as simultaneously identifying a medicine box and an instruction manual, the intelligent agent application can prioritize calling the intelligent agents that the user has used most frequently in the past, based on the frequency of the intelligent agent service 11 called by the user in the past. This ensures that the final processing result is the result the user wants and improves the user experience.
[0058] Step 203: Output the response information generated by the intelligent agent service in response to the image data through the first device.
[0059] The response information can refer to the output result generated by the intelligent agent service in response to image data. Furthermore, the response information supports multiple forms: text (such as problem-solving steps, drug instructions), voice (such as broadcasting answers), video (such as operation demonstrations), or multimodal response information composed of any one or more of the above forms. This is merely an example, and the embodiments of this application do not impose any limitations.
[0060] In one embodiment, the specific implementation of outputting the response information generated by the intelligent agent service in response to the image data through the first device 20 is as follows: the response information generated by the intelligent agent service in response to the image data is output in a modality of any one or at least a combination of two of the following: text, graphics, voice broadcast or video.
[0061] For example, the intelligent agent service generates multimodal response information based on image data. This multimodal response information may include text and voice information. This information is then output to the first device 20 in a predefined manner. For instance, text information is displayed in a scrolling manner on the display interface of the first device 20, and the font size can be set by the user. Furthermore, the television speaker synchronously plays the voice and text content at the same volume as the current television volume, supporting pause or replay functions. Further, if video content is included, the video content is played on the interface.
[0062] Furthermore, the real-time example of this application also supports users clicking buttons such as "Follow-up Question," "Favorite," and "Replay" on the interactive interface via remote control, facilitating secondary interactions. The solution provided by this application allows elderly family members to input images from various image input sources and obtain recognition results. Simultaneously, the display method of the recognition results can be controlled, such as font size, making it easier for the elderly to recognize visual objects around them and providing a unique user experience. In addition, this application also supports users simultaneously inputting multimodal input information such as images, text, and voice. Based on the user's multimodal input information, the final response information is obtained, meeting various user needs.
[0063] The method provided in this application embodiment is applied to a first device configured with an image input source. Image data is acquired through the image input source. In the interactive interface of the first device's intelligent agent application, the image data is used as dialogue input information. The intelligent agent application 22 performs content recognition on the image data and calls the corresponding intelligent agent service based on the content recognition result. The response information generated by the intelligent agent service in response to the image data is then output through the first device. By configuring an image input source for the first device, the lack of image input capability in the first device (such as a smart TV) is compensated for. This enables the first device to intelligently process images and output response information, breaking through the limitation that the first device itself lacks image input capability. It expands the new paradigm of vision-based AI interaction, thereby broadening the application scenarios of the first device and significantly improving the user experience through multimodal intelligent interaction.
[0064] Figure 3 A flowchart illustrating another embodiment of the human-computer interaction method provided in this application is shown below. Figure 2 Based on the illustrated process, this section mainly describes how to acquire image data when the image input source for the first device is the control terminal of the first device, including the following steps: Step 301: Receive image data sent by the control terminal to the first device, wherein the image data is image data acquired by the control terminal in response to the image acquisition operation through the image acquisition component integrated on the control terminal.
[0065] The control terminal 50 can refer to a peripheral device that is paired with the first device 20 (smart TV), used to control the TV and has image acquisition capabilities. In this scenario, it can refer to an integrated remote control with a camera, which is different from a traditional remote control without a camera. It is the core carrier connecting user operation and TV interaction.
[0066] The image acquisition component 51 can refer to a hardware module integrated inside the control terminal. It is the core hardware foundation for the control terminal to realize image acquisition. For example, the image acquisition component can be set on the back of the remote control, so that the user can pick up the remote control to take pictures. In addition, the component can also be set on the front of the remote control. This is just an example, and the embodiments of this application do not limit it.
[0067] Image acquisition can be a specific action that the user triggers to control the terminal to take pictures, such as pressing the physical shutter button on the remote control, touching the virtual shutter button on the remote control, or triggering it through voice commands (such as taking pictures with the remote control). In essence, it can be a direct command to start image acquisition.
[0068] In one embodiment, the user triggers an image acquisition operation on the control terminal. The control terminal controls its own image acquisition component to acquire image data according to the acquisition operation triggered by the user, preprocesses the acquired image data, and sends the preprocessed image data to the first device.
[0069] For example, if a user wants to identify a medicine box on a table, the user can pick up the remote control, point the camera on the back of the remote at the medicine box, and press the shutter button on the remote or input a photo-taking command via voice. The remote control will then preprocess the image data of the medicine box, for example, by converting the image data to JPG encoding format and packaging and encrypting it. The remote control will then upload the preprocessed image data to the relevant human-computer interaction interface of the first device 20. The first device will then display the image data collected by the user so that the user can view and confirm the image data.
[0070] Specifically, Figure 4 A schematic diagram of a system framework where the image input source is a remote control, provided in an embodiment of this application, is shown below. Figure 4 As shown, the remote control, serving as the image input source, can be equipped with a camera and a shutter button. The user points the camera at the target object and presses the shutter button. The remote control's CPU (Central Processing Unit) then instructs the camera module to take a picture. The captured image is then processed to obtain the corresponding image signal, which is input to an image encoder. This encoder encodes the image signal into an image-encoded signal that can be transmitted wirelessly from the remote control to the TV system's wireless receiver. Upon receiving the image-encoded signal, the TV system decodes it using an image decoder. The decoded image is automatically input into the dialogue input box of the smart agent application. Simultaneously, the smart agent application in the TV system uploads the image data to a cloud AI server for content recognition and intent understanding. The cloud AI server then calls upon the corresponding smart agent to process and respond based on the specific content of the image data, and outputs the response through the TV device. For example, if the image data is a test paper, the AI agent is invoked for image processing; if the image data is a medicine box, the AI health agent is invoked for image processing; if the image data is a dish, the AI food agent is invoked for image processing, and so on. The AI agent answers the user's questions about the image data, ultimately obtaining a satisfactory answer from the user, which is then output through the television device.
[0071] Step 302: In the interactive interface of the intelligent agent application of the first device, the image data is used as dialogue input information so that the intelligent agent application can perform content recognition on the image data and call the corresponding intelligent agent service based on the content recognition result.
[0072] Step 303: Output the response information generated by the intelligent agent service in response to the image data through the first device.
[0073] For steps 302-303 above, please refer to the above. Figure 2 The detailed description of the relevant embodiments will not be repeated here.
[0074] pass Figure 3 The illustrated embodiment integrates the image acquisition component into the TV's control terminal, aligning with TV users' operating habits and lowering the barrier to entry for users to use the TV's image input function. It is suitable for various home scenarios, such as home learning, elderly health, and home entertainment, compensating for the TV's inability to perceive and recognize objects at close range. This lays the foundation for the TV's human-computer interaction function and enhances the user experience. Furthermore, integrating the image acquisition component into the TV control terminal utilizes the control terminal's wireless transmission function, ensuring the integrity of transmitted data while guaranteeing the stability and real-time performance of image input.
[0075] Figure 5 A flowchart illustrating another embodiment of the human-computer interaction method provided in this application is shown below. Figure 2 Based on the illustrated process, this section mainly describes how to acquire image data when the image input source of the first device is an image acquisition component connected to the first device, including the following steps: The image acquisition component connected to the first device mentioned above can refer to external image acquisition hardware that establishes a stable connection with the TV via wired (e.g., USB interface) or wireless (e.g., WiFi, Bluetooth, or Wi-Fi). For example, the image acquisition component connected to the first device 20 can be a USB external camera, a TV-specific PTZ camera, etc., which has independent image capture capabilities.
[0076] Step 501: In response to an image acquisition operation triggered on the first device, control the image acquisition component connected to the first device to acquire image data and obtain the image data.
[0077] In one embodiment, when a user triggers an image acquisition command in the smart agent application interface of the TV, the device management module of the first device detects the image acquisition command and checks the connection status of the image acquisition component 40 connected to the first device. If the connection status is normal, the module controls the image acquisition component to perform the acquisition operation.
[0078] For example, a user can click the camera button on the TV's smart application interface via the TV's control terminal, such as a remote control, or send an image capture voice command via the TV's control terminal, transmitting the image capture command to the TV via infrared using the remote control. Furthermore, the user can also trigger the image capture operation via a button on the first device 20, causing the TV to generate a corresponding image capture command in response to the user's clicked capture button. The TV's device management module detects the image capture command and checks if the external image capture component 40 is properly connected to the TV. If a connection abnormality is detected, a prompt will pop up on the TV's interactive interface, such as "Please check if the external camera is connected or if the connection is faulty," guiding the user to troubleshoot the connection problem. If the camera connection is detected as normal, a capture command is sent to the camera, controlling the camera to capture image data. The TV then acquires the image data and performs preprocessing.
[0079] Furthermore, in this embodiment, the image acquisition component 40 connected to the television can be a movable camera, which can cover a wider visual range and facilitate users to collect various image data. At the same time, it can free up the user's hands. For example, the user can control the external camera of the first device to automatically collect image data of relevant locations through voice control, further improving the user experience.
[0080] Step 502: In the interactive interface of the intelligent agent application of the first device, the image data is used as dialogue input information so that the intelligent agent application can perform content recognition on the image data and call the corresponding intelligent agent service based on the content recognition result.
[0081] Step 503: Output the response information generated by the intelligent agent service in response to the image data through the first device.
[0082] For steps 502-503 above, please refer to the above. Figure 2 The detailed description of the relevant embodiments will not be repeated here.
[0083] pass Figure 5 The detailed description of the embodiment describes a method in which a first device controls an external image acquisition component 40 to acquire images. By utilizing the mobility of the image acquisition component 40, more image acquisition positions can be covered. At the same time, based on the cloud capabilities of the camera, the hands can be completely freed to perform image acquisition operations, thereby improving the flexibility of image input and simplifying user operations.
[0084] Figure 6 A flowchart illustrating another embodiment of the human-computer interaction method provided in this application is shown below. Figure 2 Based on the illustrated process, this section mainly describes how to acquire image data when the image input source of the first device is a second device 30 that is communicatively connected to the first device, including the following steps: The second device 30, which communicates with the first device, can be an external electronic device that establishes a connection with the first device 20 (smart TV) via wireless communication (such as WiFi, Bluetooth, or Wi-Fi) or a wired network. It possesses independent image acquisition, storage, or transmission capabilities and serves as an auxiliary terminal for TV image input. For example, the second device could be a smartphone, tablet, laptop, smart camera, etc., that establishes a communication connection with the first device.
[0085] Step 601: Receive image data sent by the second device to the first device, wherein the image data is image data acquired by the second device in response to an image acquisition operation, or image data selected by the second device from a local image library in response to an image selection operation.
[0086] The aforementioned image acquisition operation can refer to the image shooting action triggered by the user on a second device, such as clicking the shutter button of a mobile phone camera app or the camera shortcut on a tablet, to generate real-time image data.
[0087] Image selection can refer to the action of a user selecting an existing image from a locally stored image library (such as a phone's photo album or a computer's picture folder) on a second device. For example, selecting a previously taken photo from the phone's photo album as image data to be transmitted to the TV.
[0088] In one embodiment, the user selects a corresponding image input source on the first device 20. The first device sends a capture command to the second device and displays a pop-up prompt on the second device. Based on the pop-up, the user on the second device either captures an image in real time or selects a relevant image from the second device's gallery as image data and transmits it to the dialog input box of the smart agent application on the first device 20.
[0089] For example, a user selects the corresponding image input source—a second device connected to the first device—from the smart agent application interface of the first device via a remote control. At this time, the first device 20, i.e., the TV, detects the currently connected second device. Upon detection, it sends an image upload command to the second device and displays a pop-up window on the second device's interface. The user selects or captures an image from the second device's gallery through the image upload pop-up window and uploads it as image data to the dialog input box of the smart agent application on the first device. The second device can be a mobile phone.
[0090] In another embodiment, the user opens a second agent application interface that is communicatively connected to the first device 20, and uploads image data collected in real time by the second device or selected from the gallery to the dialog input box of the second agent application interface.
[0091] For example, users can connect their mobile phones to the smart agent application on the TV in advance. Specifically, they can download the same mobile application as the smart agent application on the TV and log in with the same account information as the smart agent application on the TV. Then, they can directly select pictures from their phone's album or take photos in real time and upload them to the dialog input box of the mobile application. The smart agent application on the TV will then display the same interface and data as the mobile application in real time based on the relevant data uploaded by the mobile application.
[0092] Step 602: In the interactive interface of the intelligent agent application of the first device, the image data is used as dialogue input information so that the intelligent agent application can perform content recognition on the image data and call the corresponding intelligent agent service based on the content recognition result.
[0093] Step 603: Output the response information generated by the intelligent agent service in response to the image data through the first device.
[0094] For steps 602-603 above, please refer to the above. Figure 2 The detailed description of the relevant embodiments will not be repeated here.
[0095] pass Figure 6 The detailed description of the illustrated embodiment shows that by transmitting image data from the second device to the first device, the widespread availability of the second device is fully utilized to provide a flexible, complementary, and scenario-adaptive solution for television image input. Simultaneously, by using the first device 20 to send image upload commands to the second device and enabling account and data interoperability, multimodal interactive collaboration between the first and second devices is achieved, improving the smoothness of cross-device operation and further enhancing the fluency of the smart TV's human-computer interaction functions.
[0096] Figure 7 A block diagram of a human-computer interaction device provided in this application is applied to a first device configured with an image input source. Referring to Figure 7, the device includes: Image acquisition module 71 is used to acquire image data through the image input source; The image processing module 72 is used to use the image data as dialogue input information in the interactive interface of the intelligent agent application of the first device, so that the intelligent agent application can perform content recognition on the image data and call the corresponding intelligent agent service based on the content recognition result. The image output module 73 is used to output the response information generated by the intelligent agent service in response to the image data through the first device.
[0097] In one possible implementation, the image input source includes the control terminal of the first device, and the control terminal integrates an image acquisition component; The image acquisition module 71 is specifically used for: The control terminal receives image data sent to the first device, wherein the image data is image data acquired by the control terminal in response to an image acquisition operation through an image acquisition component integrated on the control terminal.
[0098] In one possible implementation, the image input source includes an image acquisition component connected to the first device; The image acquisition module 71 is specifically used for: In response to an image acquisition operation triggered on the first device, the image acquisition component connected to the first device is controlled to acquire image data and obtain the image data.
[0099] In one possible implementation, the image input source includes a second device that is communicatively connected to the first device; The image acquisition module 71 is specifically used for: The device receives image data sent from the second device to the first device, wherein the image data is image data acquired by the second device in response to an image acquisition operation, or image data selected by the second device from a local image library in response to an image selection operation.
[0100] In one possible implementation, the image input source includes a screenshot function component integrated into the first device; The image acquisition module 71 is specifically used for: Obtain image data generated by the first device taking a screenshot of the currently displayed screen through the screenshot function component.
[0101] In one possible implementation, the image output module 73 is specifically used for: The response information generated by the intelligent agent service in response to the image data is output in a modality that combines or integrates any one or at least two of the following forms: text, graphics, voice broadcast, or video.
[0102] like Figure 8 As shown in the figure, this application provides an electronic device, including a processor 111, a communication interface 112, a memory 113, and a communication bus 114, wherein the processor 111, the communication interface 112, and the memory 113 communicate with each other through the communication bus 114. Memory 113 is used to store computer programs; In one embodiment of this application, when the processor 111 executes a program stored in the memory 113, it implements the human-computer interaction method provided in any of the foregoing method embodiments, including: Image data is acquired through the image input source; In the interactive interface of the intelligent agent application of the first device, the image data is used as dialogue input information so that the intelligent agent application can perform content recognition on the image data and call the corresponding intelligent agent service based on the content recognition result. The response information generated by the intelligent agent service in response to the image data is output through the first device.
[0103] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the human-computer interaction method provided in any of the foregoing method embodiments.
[0104] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0105] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented using software plus a general-purpose hardware platform, or of course, using hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0106] It should be understood that the terminology used herein is for the purpose of describing particular exemplary embodiments only and is not intended to be limiting. Unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “described” as used herein may also mean including the plural forms. The terms “comprising,” “including,” “containing,” and “having” are inclusive and therefore indicate the presence of the stated features, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, elements, components, and / or combinations thereof. The method steps, processes, and operations described herein are not construed as requiring them to be performed in a particular order described or illustrated unless the order of performance is explicitly indicated. It should also be understood that additional or alternative steps may be used.
[0107] The above description is merely a specific embodiment of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A human-computer interaction system, characterized in that, The system includes: an intelligent agent server, a first device configured with an image input source, the first device integrating an intelligent agent application, and the intelligent agent server integrating at least one intelligent agent service. The first device acquires image data of the target object through the image input source; In the interactive interface of the intelligent agent application of the first device, the image data is used as dialogue input information so that the intelligent agent application can perform content recognition on the image data and call the corresponding intelligent agent service based on the content recognition result. The intelligent agent service generates response information in response to the image data; The first device outputs the response information.
2. The system according to claim 1, characterized in that, The system further includes: a control terminal for the first device, wherein the control terminal integrates an image acquisition component; The control terminal, in response to an image acquisition operation, acquires image data through an integrated image acquisition component and sends the image data to the first device.
3. The system according to claim 1, characterized in that, The system further includes: an image acquisition component connected to the first device; The first device controls an image acquisition component connected to the first device to acquire image data and obtain the image data.
4. The system according to claim 1, characterized in that, The system further includes: a second device that is communicatively connected to the first device; The second device acquires image data in response to an image acquisition operation; or selects image data from a local image library in response to an image selection operation. The image data is sent to the first device.
5. The system according to claim 1, characterized in that, The first device integrates a screenshot function component; The first device uses an integrated screenshot function component to capture the currently displayed screen and generate image data.
6. A human-computer interaction method, characterized in that, The method, applied to the human-computer interaction system according to any one of claims 1-5, comprises: Image data is acquired through the image input source; In the interactive interface of the intelligent agent application of the first device, the image data is used as dialogue input information so that the intelligent agent application can perform content recognition on the image data and call the corresponding intelligent agent service based on the content recognition result. The response information generated by the intelligent agent service in response to the image data is output through the first device.
7. The method according to claim 6, characterized in that, The image input source includes the control terminal of the first device, and the control terminal integrates an image acquisition component. The step of acquiring image data through the image input source includes: The control terminal receives image data sent to the first device, wherein the image data is image data acquired by the control terminal in response to an image acquisition operation through an image acquisition component integrated on the control terminal.
8. The method according to claim 6, characterized in that, The image input source includes an image acquisition component connected to the first device; The step of acquiring image data through the image input source includes: In response to an image acquisition operation triggered on the first device, the image acquisition component connected to the first device is controlled to acquire image data and obtain the image data.
9. The method according to claim 6, characterized in that, The image input source includes a second device that is communicatively connected to the first device; The step of acquiring image data through the image input source includes: The device receives image data sent from the second device to the first device, wherein the image data is image data acquired by the second device in response to an image acquisition operation, or image data selected by the second device from a local image library in response to an image selection operation.
10. The method according to claim 6, characterized in that, The image input source includes the screenshot function component integrated into the first device; The step of acquiring image data through the image input source includes: Obtain image data generated by the first device taking a screenshot of the currently displayed screen through the screenshot function component.