Information processing method and device, equipment and storage medium

By using a processing module that receives images in the interactive interface and automatically matches them with interactive scenarios, the problem of poor adaptability of interactive scenarios in existing technologies is solved, thereby improving user experience and diversifying interactive scenarios.

CN119781888BActive Publication Date: 2025-10-17BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411987761.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-10-17
Estimated Expiration
2044-12-31

AI Technical Summary

Technical Problem

In existing technologies, generative models cannot automatically identify users' interaction scenario needs, resulting in poor adaptability to interaction scenarios and a poor user experience.

Method used

By providing image input controls in the interactive interface, the system receives images input by users, automatically determines the matching interactive scene based on the image, calls the corresponding processing module to generate response content, and supports diverse interactive scene selection and expansion.

Benefits of technology

It enables accurate identification of target interaction scenarios based on user input images and generates appropriate response content, thereby improving user experience and the diversity and scalability of interaction scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119781888B_ABST
    Figure CN119781888B_ABST
Patent Text Reader

Abstract

According to embodiments of the present disclosure, information processing methods, apparatuses, devices and storage media are provided. The method includes presenting an interaction interface with a virtual object, the interaction interface including an image input control; receiving a first input image via the image input control; and in response to the first input image being determined to match a first interaction scenario, providing first response content for the first input image in the interaction interface, the first response content being generated by a first processing module corresponding to the first interaction scenario. Thus, embodiments of the present disclosure can automatically identify an interaction scenario matching an input image, ensure the adaptability between the generated response content and the input image and its interaction scenario, and improve user experience.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Example embodiments of the present disclosure generally relate to the field of computers, and in particular, to an information processing method, apparatus, device and computer readable storage medium. BACKGROUND

[0002] With the development of computer technology, the Internet has become an important platform for people to interact with information, and various forms of electronic devices can greatly enrich people's daily life. For example, some electronic devices can provide virtual scenes for users, such as users interacting with virtual characters, etc.

[0003] In such a virtual scene, it is difficult to automatically and accurately determine what kind of scene interaction the user wants for the information input by the user, which affects the user's interaction experience. SUMMARY

[0004] In a first aspect of the present disclosure, an information processing method is provided. The method comprises: presenting an interaction interface with a virtual object, the interaction interface comprising an image input control; receiving a first input image via the image input control; and in response to the first input image being determined to match a first interactive scene, providing first response content for the first input image in the interaction interface, the first response content being generated by a first processing module corresponding to the first interactive scene.

[0005] In a second aspect of the present disclosure, an information processing apparatus is provided. The apparatus comprises: an interface presentation module, an image receiving module and a first content providing module, wherein the presentation module is configured to present an interaction interface with a virtual object, the interaction interface comprising an image input control; the receiving module is configured to receive a first input image via the image input control; and the first content providing module is configured to, in response to the first input image being determined to match a first interactive scene, provide first response content for the first input image in the interaction interface, the first response content being generated by a first processing module corresponding to the first interactive scene.

[0006] In a third aspect of the present disclosure, an electronic device is provided. The device comprises at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. The instructions, when executed by the at least one processing unit, cause the device to perform the method of the first aspect.

[0007] In a fourth aspect of the present disclosure, a computer readable storage medium is provided. The computer readable storage medium has stored thereon a computer program, the computer program being executable by a processor to implement the method of the first aspect.

[0008] In a fifth aspect of the present disclosure, a computer program product is provided. The computer program product is tangibly stored in a computer storage medium and comprises computer-executable instructions that, when executed by a device, cause the device to perform the method provided according to the first aspect.

[0009] It should be understood that nothing in the Summary is to be construed as a limitation on the scope of the embodiments of the present disclosure. Other features, aspects, and advantages of the present disclosure will become apparent from the following description, which is given by way of example in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS

[0010] The above and other features, aspects, and advantages of embodiments of the present disclosure will become more apparent from the following detailed description when taken in conjunction with the accompanying drawings. In the drawings, like reference numerals refer to like elements, in which:

[0011] FIG. 1 A schematic diagram showing an example environment in which embodiments according to the present disclosure can be implemented is shown;

[0012] FIGS. 2A-2C An example interface according to some embodiments of the present disclosure is shown;

[0013] FIGS. 3A-3C An example scenario according to some embodiments of the present disclosure is shown;

[0014] FIG. 4 A flowchart showing an example process of information processing according to some embodiments of the present disclosure is shown;

[0015] FIG. 5 A schematic block diagram showing an example apparatus for information processing according to some embodiments of the present disclosure is shown; and

[0016] FIG. 6 A block diagram of an electronic device capable of implementing a number of embodiments of the present disclosure is shown. DETAILED DESCRIPTION

[0017] Embodiments of the present disclosure will be described below in greater detail with reference to the accompanying drawings. While certain embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be embodied in various forms and should not be construed as being limited to the embodiments set forth herein; rather, these embodiments are provided so that the present disclosure will be more thoroughly and completely understood. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of the present disclosure.

[0018] It should be noted that the titles of any sections / sub-sections provided herein are not limiting. Various embodiments are described throughout this document and any type of embodiment can be included under any section / sub-section. Furthermore, embodiments described in any section / sub-section can be combined with any other embodiments described in the same section / sub-section and / or different section / sub-section in any manner.

[0019] In the description of embodiments of the disclosure, the term "includes" and its conjugates are open-ended, i.e., "includes but is not limited to". The term "based on" is intended to mean "based, at least in part, on" The term "one embodiment" or "an embodiment" is intended to mean "at least one embodiment". The term "some embodiments" is intended to mean "at least some embodiments". Other explicit or implicit definitions can also be included below. The terms "first", "second", etc. can refer to different or the same objects. Other explicit and implicit definitions can also be included below.

[0020] Data of users, acquisition and / or use of data, etc. can be involved in embodiments of the disclosure. These aspects all comply with corresponding laws and regulations and relevant provisions. In embodiments of the disclosure, all data collection, acquisition, processing, processing, forwarding, use, etc. are carried out on the premise that the user is aware of and confirms. Accordingly, when implementing embodiments of the disclosure, the type of data or information that can be involved, the use range, the use scenario, etc. should be notified to the user and the authorization of the user should be obtained according to relevant laws and regulations through appropriate means. The specific notification and / or authorization mode can vary according to the actual situation and application scenario, and the scope of the disclosure is not limited in this respect.

[0021] In the specification and embodiments, if personal information processing is involved, it will be processed on the premise of legality (for example, obtaining the consent of the subject of personal information, or being necessary for the performance of a contract, etc.), and only within the prescribed or agreed range. Users refuse to process personal information other than the necessary information required for basic functions, which will not affect the user's use of basic functions.

[0022] Information processing is an important part of the online information interaction process. With the development of artificial intelligence technology, the interaction scenarios between people and virtual objects are also expanding. For example, in the field of artificial intelligence, people have more and more interaction demands for general medical treatment, such as medical consultation, health encyclopedia, report interpretation, symptom identification, medication consultation, and other scene requirements.

[0023] In the related art, a generative model supports general camera uploaded image class data, but cannot automatically identify the interactive scene requirement of a user according to a received image, that is, cannot automatically and accurately identify and call a matched scene sub-model to adapt to the user requirement; moreover, the interactive scene that can be adapted is relatively single, and it is difficult to meet the diversified requirement of the user, and the user experience is poor.

[0024] Embodiments of the present disclosure provide a scheme for information processing. According to the scheme, an interactive interface with a virtual object can be presented, the interactive interface including an image input control; a first input image is received via the image input control; and in response to the first input image being determined to match a first interactive scene, first response content for the first input image is provided in the interactive interface, the first response content being generated by a first processing module corresponding to the first interactive scene.

[0025] According to the scheme provided by the embodiments of the present disclosure, the interactive scene matching the image input by the user can be automatically determined, so that the corresponding processing module is called to process the input image, to accurately generate and provide the corresponding response content, effectively ensuring the accuracy of the provided response content, and improving the user experience.

[0026] In some embodiments of the present disclosure, a set of interactive entrances corresponding to a set of preset interactive scenes can also be provided in the interactive interface of the scheme, the set of preset interactive scenes including the first interactive scene, so as to improve the diversity of the interactive scene, to meet the diversity of the user interactive requirement, and to improve the expandability of the interactive scene, and to improve the user experience.

[0027] Various example implementations of the scheme are described in further detail below in conjunction with the accompanying drawings.

[0028] Example Environment

[0029] FIG. 1 A schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented is shown. As shown, the example environment 100 can include an electronic device 110. FIG. 1

[0030] In this example environment 100, the electronic device 110 can have an application 120 running that supports interface interaction. The application 120 can be any suitable type of application for interface interaction, examples of which can include, but are not limited to, a chat application, a virtual interaction application, or other suitable application. A user 140 can interact with the application 120 via the electronic device 110 and / or its attached devices.

[0031] In FIG. 1 ​If the application 120 is active, the electronic device 110 can present, through the application 120, an interface 150 for supporting interaction with the virtual object in the environment 100.

[0032] In some embodiments, the electronic device 110 communicates with the server 130 to implement provisioning of services for the application 120. The electronic device 110 can be any type of mobile terminal, fixed terminal, or portable terminal including a mobile handset, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, a palmtop computer, a portable gaming terminal, a VR / AR device, a Personal Communication System (PCS) device, a personal navigation device, a Personal Digital Assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an electronic book device, a game device, or any combination thereof, including an accessory or peripheral device for the foregoing, or any combination thereof. In some embodiments, the electronic device 110 can also support any type of interface to the user (such as "wearable" circuitry, etc.).

[0033] The server 130 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks, and basic cloud computing services such as big data and artificial intelligence platforms. The server 130 may, for example, include a computing system / server, such as a mainframe, an edge computing node, a computing device in a cloud environment, and the like. The server 130 can provide background services for the application 120 in the electronic device 110 that supports interaction with the virtual object.

[0034] A communication connection can be established between the server 130 and the electronic device 110. The communication connection can be established by wired or wireless means. The communication connection can include, but is not limited to, a Bluetooth connection, a mobile network connection, a Universal Serial Bus (USB) connection, a Wireless Fidelity (WiFi) connection, and the like, and embodiments of the present disclosure are not limited in this regard. In embodiments of the present disclosure, the server 130 and the electronic device 110 can implement signaling interaction through the communication connection therebetween.

[0035] It should be understood that the structure and function of the various elements in the environment 100 are described for illustrative purposes only, and do not imply any limitation on the scope of the present disclosure.

[0036] Some example embodiments of the present disclosure will be described below with continued reference to the accompanying drawings.

[0037] Example Interaction

[0038] FIGS. 2A-2C Example interfaces 200A-200C according to some embodiments of the present disclosure are shown. The interfaces 200A-200C may, for example, be provided by the electronic device 110 shown. FIG. 1

[0039] In some embodiments, as shown in FIG. 2A The electronic device 110 can present an interaction interface 200A with a virtual object. The virtual object can refer to “virtual object 1” shown. FIG. 2A The electronic device 110 can provide an image input control 211 in the interaction interface 200A. Through the image input control 211, a user can input an image to interact with the virtual object.

[0040] For example, the electronic device 110 can provide an information input component 210 in the interaction interface 200A, which includes the image input control 211.

[0041] A user can trigger the image input control 211 in the interaction interface 200A to input an image to achieve information interaction with the virtual object.

[0042] For example, the image input by the user can be a picture, such as a photo of any content. For example, in a health consultation scenario, the image input by the user can be a picture of a symptom on the body surface, a text picture of a test report, an image of a drug appearance or drug packaging, etc., which is not limited here.

[0043] In some implementations, the image input by the user can also be an audio and video image, such as an audio and video clip carrying voice description for a symptom, or an audio and video clip carrying voice information for a drug, etc., which is not limited here.

[0044] According to embodiments of the present disclosure, the user can input an image stored locally in the electronic device 110 via the image input control 211, or take a new image through the image input control 211, or generate a new image through the image input control 211 by invoking an image generation application (such as a camera) in the electronic device 110, or find and obtain an image through the image input control 211 by invoking a storage application (such as a cloud disk) installed in the electronic device 110, which is not limited here.

[0045] Referring to FIG. 2A ​As shown, in embodiments of the present disclosure, the information input component 210 can further include at least one of a voice input control 212, a text input control 213, and other input controls. Among them, the other input controls can be input controls triggered by the "+" as shown in the middle, for example, an expression input control. FIG. 2A

[0046] In some embodiments, the user can input voice information in the interactive interface 200A through the voice input control 212 in the information input component 210. The voice information input by the user can be independent interactive information such as a question or a reply, or can be associated description information of the input image input after or before the input image, for example, a question content about any information shown in the input image.

[0047] In some embodiments, the user can input text information in the interactive interface 200A through the text input control 213 in the information input component 210. The text information input by the user can be independent interactive information such as a question or a reply, or can be associated description information of the input image input after or before the input image, for example, a question content about any information shown in the input image.

[0048] In some embodiments, the user can also input other information by clicking the "+" to trigger other input controls. For example, by clicking "+" to trigger an expression input control, and inputting corresponding emoticons or expression images as a reply or response to the information provided by the virtual object.

[0049] Referring to FIG. 2B As shown, the electronic device 110 can present an interactive interface 200B with a virtual object. The interactive interface 200B includes an image input control, and can also include a set of interactive entrances 230 corresponding to a set of preset interactive scenarios. For example, the preset interactive scenarios can include health encyclopedia, medical consultation, symptom identification, report interpretation, medication consultation, and other health consultation scenarios, and can also include AI image generation, document generation, and other interactive scenarios, which are not limited here.

[0050] In some embodiments, the user can input an image 201 through the image input control in the interactive interface 200B, and the electronic device 110 can identify the content of the image 201 after receiving the image 201 input by the user, then determine a target interactive scenario matched with the image according to the identification result, and call a processing module corresponding to the target interactive scenario to process the image 201 input by the user, so as to obtain a response content 220 for the image and display it in the interactive interface 200B.

[0051] ​For example, after receiving the image 201 input by the user, the electronic device 110 provides the image 201 to the visual language model to determine a target interaction scenario from a plurality of preset interaction scenarios, so as to determine a target processing module to process the image 201, so as to ensure the adaptability of the response content 220 to the image 201 and the target interaction scenario thereof.

[0052] In some embodiments, the user can also select a target interaction entry from a set of interaction entries 230 and trigger, for example, select the interaction entry corresponding to “report interpretation” and click, and then input an interaction image or other interaction information via the target interaction entry. After the electronic device 110 receives the interaction image or other interaction information, the corresponding response content is presented in the interaction interface 200B.

[0053] For example, the user inputs the image 201 by triggering the target interaction entry, and the electronic device 110 responds to the triggering of the target interaction entry, acquires the image 201, and then based on the target interaction scenario corresponding to the target interaction entry, mobilizes the processing module corresponding to the target interaction scenario to process the image 201, generates the response content 220 for the image 201, and displays in the interaction interface 200B.

[0054] In some embodiments, the generation process of the response content 220 includes: after the electronic device 110 receives the image 201, determining a target response mode for the image 201 based on the context information associated with the interaction interface, the target response mode indicating whether to provide follow-up question content; and acquiring the response content 220 generated based on the target response mode.

[0055] For example, after determining the target response mode for the image 201, the electronic device calls the processing module corresponding to the target response mode to process the image 201, and then acquires the response content 220 for the image 201 from the processing module.

[0056] In some embodiments, the electronic device 110 determines that the image 201 corresponds to a first response mode in response to the context information associated with the interaction interface indicating that the image 201 corresponds to a first round of interaction with a virtual object, the first response mode indicating that no follow-up question content is provided.

[0057] In some embodiments, the electronic device 110 determines that the image 201 corresponds to the first response mode in response to the context information associated with the interaction interface indicating that the image 201 is independent of the health consultation process, for example, the content of the image 201 is completely different from the context information or belongs to different content, the first response mode indicating that no follow-up question content is provided. For example, the context is about a knee joint examination report, and the image 201 is a hand symptom diagram.

[0058] In some embodiments, the electronic device 110 determines that the image 201 corresponds to the first response mode, i.e., no follow-up content is provided, in response to the context information associated with the interactive interface indicating that the historical messages of the interactive interface 200B have a relevance to the image 201 that is lower than a threshold.

[0059] The relevance between the context information and the image 201 can be evaluated by a relevance level or a relevance score, for example.

[0060] For example, the relevance level between the context information and the image 201 can be divided according to a preset level division standard, and the relevance levels include, in sequence, completely different, a small amount of relevance, medium relevance, close relevance, and almost the same. Correspondingly, the relevance threshold between the two can be any one of the medium relevance, close relevance, or almost the same in the above relevance levels.

[0061] For example, the context information associated with the interactive interface and the image 201 can be respectively identified and analyzed by using a deep learning model or a large language model, and a relevance score between the two can be calculated. For example, any one of a ten-point system, a percentage system, or a percentage can be selected to determine the relevance score. Correspondingly, the relevance threshold between the two is a specific score value, such as 7.5 points in a ten-point system, 80 points in a percentage system, or 80%, etc.

[0062] In some embodiments, the electronic device 110 determines that the first input image corresponds to the second response mode in response to the context information associated with the interactive interface indicating that the image 201 is associated with a health consultation process, or the historical messages of the interactive interface 200B have a relevance to the image 201 that reaches a threshold. The second response mode indicates that follow-up content is provided. The follow-up content can include follow-up content for the content shown in the image 201. For example, for a symptom image, the follow-up content can be “how long has this condition lasted?”

[0063] Referring to FIG. 2C, FIG. 2C As shown in FIG. 2C, the electronic device 110 presents an interactive interface 200C with a virtual object. The interactive interface 200C includes an image input control, and can also include a set of interactive entrances 230 corresponding to a set of preset interactive scenarios. For example, the preset interactive scenarios can include health encyclopedia, medical consultation, symptom identification, report interpretation, medication consultation, and other health consultation scenarios, and can also include AI image generation, document generation, and other interactive scenarios, without limitation.

[0064] In some embodiments, the user can input the image 201 through an image input control in the interaction interface 200C, or through triggering a target interaction entry in the set of interaction entries 230, after receiving the image 201, the electronic device 110 determines that the image 201 corresponds to the second response mode, i.e., providing follow-up content, in response to the context information associated with the interaction interface 200C indicating that the image 201 is associated with the health consultation process, or in response to the relevance between the historical message of the interaction interface 200C and the image 201 reaching a threshold. The electronic device 110 invokes a corresponding follow-up module to generate the follow-up content 221 and present it in the interaction interface 200C; then receives the reply content 202 given by the user for the follow-up content, and invokes a corresponding processing module to generate the response content 220 based on the reply content 202, the image 201, and the historical message, and display it in the interaction interface 200C.

[0065] In some embodiments, the electronic device 110 can generate the follow-up content 221 only once, or can generate the follow-up content again for the reply content 202 until a preset number of times or a preset follow-up condition is met, stop follow-up, and generate the corresponding response content 220.

[0066] In this way, the embodiments of the present disclosure can accurately identify the target interaction scene that the user wants according to the image input by the user, and generate and provide the response content accordingly, ensuring that the response content is adapted to the image input by the user and its target interaction scene, and improving the user experience.

[0067] In addition, the embodiments of the present disclosure also provide a set of interaction entries corresponding to a set of preset interaction scenes, which facilitates the user to manually select a target interaction entry to trigger a corresponding interaction scene, provides more interaction options for the user, and improves the user experience; and in this mode, the interaction scene and its corresponding interaction entry can be expanded at any time, improving the scalability of the interaction scene.

[0068] Example Scenario

[0069] FIGS. 3A-3C Example scenarios 300A to 300C according to some embodiments of the present disclosure are shown. The scenarios 300A to 300C are respectively the process of the electronic device 110 generating and presenting response content according to the image input by the user in different scenarios.

[0070] In some embodiments, as shown in FIG. 3A The scheme provided by the present disclosure can be applied to the symptom interpretation scenario 300A in the health consultation process.

[0071] Referring to FIG. 3AAs shown, after the electronic device 110 processes the symptom image input by the user through the health interception process 310, it is determined that the symptom image matches the symptom interpretation scene, and the symptom interpretation module 301 is called to process and respond to the symptom image.

[0072] For example, the symptom interpretation module 301 processes the symptom image as follows:

[0073] In block 321, the symptom interpretation module 301 determines whether the symptom image corresponds to the first round of interaction with the virtual object, and if so, block 322 is executed.

[0074] In block 322, the symptom interpretation module 301 calls the visual language model symptom interpretation bot A to process the symptom image, and executes block 323; wherein the visual language model symptom interpretation bot A only processes the image content, does not ask questions about the image content, for example, identifies the symptoms in the image and converts them into text descriptions, and can also analyze and suggest the identified symptoms;

[0075] In block 323, the symptom interpretation module 301 obtains response content for the symptom image based on at least one of the output symptom description, reason, and suggestion, and then executes block 330.

[0076] In block 330, the symptom interpretation module 301 streams the response content to the electronic device 110; and finally, the electronic device 110 executes the output process 340 to present the response content in the interaction interface.

[0077] Returning to block 321, if the symptom interpretation module 301 determines that the symptom image does not correspond to the first round of interaction with the virtual object, block 324 is executed.

[0078] In block 324, the symptom interpretation module 301 determines whether it is currently in the consultation process, and if not, block 322 is executed; if so, blocks 325 and 327 are executed; wherein the determination of whether it is currently in the consultation process can include: based on whether the context information associated with the interaction interface is relevant to the symptom image, if it is relevant, it is in the consultation process, if it is not relevant, it is not in the consultation process.

[0079] In block 325, the symptom interpretation module 301 calls the visual language model to determine the relevance score of the current message and the historical message, and executes block 326 based on the relevance score; wherein the current message includes the current symptom image and the context information, and the historical message can be the historical message of the current interaction interface.

[0080] Exemplarily, the determination of the association score of the current image-text message and the historical message can be implemented through image recognition, text conversion, semantic analysis, similarity comparison, etc. processes, and can refer to related public technologies, which will not be described here.

[0081] At block 326, the symptom interpretation module 301 compares the association score determined in block 325 with a preset threshold to determine whether the association score is greater than the threshold; if not, return to block 322; if yes, execute block 328.

[0082] At block 328, the symptom interpretation module 301 calls the visual language model symptom recognition botB to identify the symptoms in the symptom image, obtains the symptom description, and generates the corresponding follow-up content, and then executes block 329; wherein the follow-up content is the follow-up information generated for the symptom image in combination with the historical message, the context information, the symptom description, etc.

[0083] At block 329, the symptom interpretation module 301 collates and outputs the symptom description and the follow-up content obtained in block 328 to obtain the response content, and executes block 330.

[0084] Returning to block 324, the symptom interpretation module 301, in response to determining that it is currently in the consultation process, can also call the visual language model to rewrite the consultation prompt information of the symptom image in combination with the context information, and execute block 328.

[0085] In some embodiments, as shown in FIG. 3B The scheme provided by the present disclosure can be applied to the report interpretation scene 300B in the health consultation process.

[0086] Referring to FIG. 3B After the electronic device 110 inputs the report image through the health interception process 310, it is determined that the report image matches the report interpretation scene, and the report interpretation module 302 is called to process and respond to the report image.

[0087] Exemplarily, the processing flow of the report interpretation module 302 on the report image is as follows:

[0088] At block 351, the report interpretation module 302 first calls the OCR (Optical Character Recognition) model to extract the text information in the report image, and then executes block 352;

[0089] At block 352, the report interpretation module 302 determines whether the virtual object corresponding to the report image is the first round of interaction, if not, execute block 353; if yes, execute block 356;

[0090] At block 353, the report interpretation module 302 invokes the main bot to find the context information associated with the interaction interface, and performs block 354.

[0091] At block 354, the report interpretation module 302 determines whether the current is in the consultation process, if yes, performs block 355; if no, performs block 356; wherein the way of determining whether the current is in the consultation process can include: based on whether the report image has relevance with the context information associated with the interaction interface, if having relevance, then in the consultation process, if not having relevance, then not in the consultation process;

[0092] At block 355, the report interpretation module 302 invokes the consultation bot to generate response content for the report image in combination with the text information and the context information, and performs block 330;

[0093] At block 330, the report interpretation module 302 streams the response content to the electronic device 110; finally, the electronic device 110 performs the output process 340 to present the response content in the interaction interface;

[0094] Returning to block 352, if the report interpretation module 302 determines that the report image corresponds to the first round of interaction of the virtual object, block 356 is performed;

[0095] At block 356, the report interpretation module 302 invokes the main bot to process the text information to obtain the response content for the report image, and performs block 330.

[0096] In some embodiments, as shown in FIG. 3A The scheme provided by the present disclosure can be applied to a drug identification scene 300C in the health consultation process.

[0097] Referring to FIG. 3C After the electronic device 110 inputs the drug image through the health interception process 310, it is determined that the drug image matches the drug identification scene, and the drug identification module 303 is invoked to process and respond to the drug image.

[0098] For example, the processing process of the drug identification module 303 on the drug image is as follows:

[0099] At block 361, the drug identification module 303 determines that the drug image is adapted to the drug identification module, and performs blocks 362 and 367;

[0100] At block 362, an OCR model is first invoked to identify the drug object in the drug image, and then block 363 is performed;

[0101] At block 363, the medicine recognition module 303 retrieves the associated information of the medicine object from the medicine knowledge base according to the medicine object identified at block 362, and performs block 364;

[0102] At block 364, the medicine recognition module 303 outputs the medicine specification corresponding to the medicine object, and performs block 365 or block 371;

[0103] At block 365, the medicine recognition module 303 calls a model (e.g., a generative model) to analyze and summarize the medicine object and the medicine specification, to obtain comprehensive information, and then performs block 366;

[0104] At block 366, the medicine recognition module 303 generates response content for the medicine image based on the comprehensive information obtained at block 365, so as to correctly reply to the user, and performs block 330;

[0105] At block 330, the medicine recognition module 303 streams the response content to the electronic device 110; and finally, the electronic device 110 performs the output process 340 to present the response content in the interactive interface;

[0106] At block 367, the medicine recognition module 303 determines whether the medicine image corresponds to the first round of interaction of the virtual object, and if yes, performs block 365; and if not, performs block 368;

[0107] At block 368, the medicine recognition module 303 determines whether it is currently in the consultation process, and if not, performs block 365; and if yes, performs block 369; wherein the determination of whether it is currently in the consultation process can include: determining whether the context information associated with the interactive interface is relevant to the medicine image, and if yes, it is in the consultation process; and if not, it is not in the consultation process;

[0108] At block 369, the medicine recognition module 303 calls a model (e.g., a visual language model or other generative model) to determine the relevance score of the current medicine information and the context information, and performs block 370 according to the relevance score;

[0109] For example, the determination of the relevance score of the current medicine information and the context information can be achieved through image recognition, text conversion, semantic analysis, similarity comparison, etc., which can refer to related public technologies, and will not be described here;

[0110] At block 370, the medicine recognition module 303 compares the relevance score determined at block 369 with a preset threshold, to determine whether the relevance score is greater than the threshold; if not, it returns to perform block 365; and if yes, it performs block 371;

[0111] At block 371, the medicine identification module 303 invokes a model (e.g., a generative model) to process the medicine object, the medicine instruction in combination with the context information to obtain a follow-up question content, and then perform block 372.

[0112] At block 372, the medicine identification module 303 gives a reply and a follow-up question to the user based on the follow-up question content obtained at block 371 in combination with the information of the medicine object and its instruction, generates response content for the medicine image, and performs block 330.

[0113] Example Process

[0114] FIG. 4 A flowchart illustrating an example process 400 of information processing according to some embodiments of the present disclosure is shown. The process 400 can be implemented at the electronic device 110. The process 400 is described below with reference to FIG. 1 .

[0115] As shown in FIG. 4 , at block 410, the electronic device 110 presents an interaction interface with a virtual object, the interaction interface including an image input control.

[0116] At block 420, the electronic device 110 receives a first input image via the image input control.

[0117] At block 430, the electronic device 110 provides, in the interaction interface, first response content for the first input image in response to the first input image being determined to match a first interaction scenario, the first response content being generated by a first processing module corresponding to the first interaction scenario.

[0118] According to the information processing method of the embodiments of the present disclosure, the electronic device can automatically determine the first interaction scenario matching the first input image, and mobilize the corresponding first processing module to generate the first response content, and then provide the first response content in the interaction interface. By accurately matching the first interaction scenario, the adaptation accuracy of the interaction scenario is improved, and then the first processing module corresponding to the first interaction scenario is mobilized to generate the first response content, the accuracy of information processing is improved, so as to ensure the adaptation of the obtained first response content to the interaction scenario of the first input image, and improve the user experience.

[0119] In some embodiments, the interaction interface presented by the electronic device 110 further includes a set of interaction entries corresponding to a set of preset interaction scenarios, the set of preset interaction scenarios including the first interaction scenario. Correspondingly, the information processing method of the present disclosure further includes: based on triggering of a target interaction entry in the set of interaction entries, obtaining a second input image; and providing, in the interaction interface, second response content for the second input image, the second response content being generated by a second processing module corresponding to a second interaction scenario, the second interaction scenario corresponding to the target interaction entry.

[0120] In some embodiments, the information processing method of the present disclosure further includes: providing the first input image to a visual language model to determine, from a plurality of preset interaction scenarios, a first interaction scenario matching the first input image.

[0121] In some embodiments, the first response content is generated based on the following process: determining, based on context information associated with the interaction interface, a target response mode for the first input image, the target response mode indicating whether to provide follow-up question content; and obtaining the first response content generated based on the target response mode.

[0122] In some embodiments, determining, based on the context information associated with the interaction interface, the target response mode for the first input image includes any one of: in response to the context information indicating that the first input image corresponds to a first round of interaction with the virtual object, determining that the first input image corresponds to a first response mode, the first response mode indicating not to provide follow-up question content; in response to the context information indicating that the interaction interface is independent of a health consultation process, determining that the first input image corresponds to the first response mode; in response to the context information indicating that a relevance of a historical message of the interaction interface to the first input image is below a threshold value, determining that the first input image corresponds to the first response mode.

[0123] In some embodiments, determining, based on the context information associated with the interaction interface, the target response mode for the first input image includes: in response to the context information indicating that the first input image is associated with the health consultation process or that a relevance of a historical message of the interaction interface to the first input image reaches a threshold value, determining that the first input image corresponds to a second response mode, the second response mode indicating to provide follow-up question content.

[0124] In some embodiments, different target response modes correspond to different processing entities for generating the first response content.

[0125] In some embodiments, the first interaction scenario indicates identifying a symptom associated with the first input image, and the first response content includes at least one of: a symptom description associated with the symptom in the first input image; cause information associated with the symptom in the first input image; and suggestion information associated with the symptom in the first input image.

[0126] In some embodiments, the first interaction scenario indicates to obtain an interpretation corresponding to the report in the first input image, and the first response content includes interpretation text about the report in the first input image.

[0127] In some embodiments, the first interaction scenario indicates to identify a drug object in the first input image, and the first response content includes description text of the drug object.

[0128] In some embodiments, the first response content is determined based on a drug instruction associated with the drug object.

[0129] Example Devices and Apparatus

[0130] Embodiments of the present disclosure also provide a corresponding device for implementing the above method or process. FIG. 5 A schematic structural block diagram of an example device 500 for xxx is shown according to certain embodiments of the present disclosure. The device 500 can be implemented as or included in the electronic device 110. Various modules / components in the device 500 can be implemented by hardware, software, firmware, or any combination thereof.

[0131] As shown in FIG. 5 The device 500 includes an interface presentation module 510, an image receiving module 520, and a first content providing module 530. The presentation module 510 is configured to present an interaction interface with a virtual object, the interaction interface including an image input control; the image receiving module 520 is configured to receive a first input image via the image input control; and the first content providing module 530 is configured to, in response to the first input image being determined to match a first interaction scenario, provide first response content for the first input image in the interaction interface, the first response content being generated by a first processing module corresponding to the first interaction scenario.

[0132] In some embodiments, the interaction interface further includes a set of interaction portals corresponding to a set of preset interaction scenarios, the set of preset interaction scenarios including the first interaction scenario; and the information processing device provided by the present disclosure can further include an image obtaining module and a second content providing module, wherein the image obtaining module is configured to, based on triggering of a target interaction portal in the set of interaction portals, obtain a second input image; and the second content providing module is configured to provide second response content for the second input image in the interaction interface, the second response content being generated by a second processing module corresponding to a second interaction scenario, the second interaction scenario corresponding to the target interaction portal.

[0133] In some embodiments, the information processing apparatus provided by the present disclosure further includes a scene determination module configured to provide the first input image to a visual language model to determine a first interactive scene matching the first input image from a plurality of preset interactive scenes.

[0134] In some embodiments, the first response content is generated based on a process including: determining, based on the context information associated with the interactive interface, a target response mode for the first input image, the target response mode indicating whether to provide follow-up content; and obtaining the first response content generated based on the target response mode.

[0135] In some embodiments, determining, based on the context information associated with the interactive interface, the target response mode for the first input image includes any one of: in response to the context information indicating that the first input image corresponds to a first round of interaction with the virtual object, determining that the first input image corresponds to a first response mode, the first response mode indicating not to provide follow-up content; in response to the context information indicating that the interactive interface is independent of the health consultation process, determining that the first input image corresponds to the first response mode; in response to the context information indicating that a relevance of a historical message of the interactive interface to the first input image is below a threshold, determining that the first input image corresponds to the first response mode.

[0136] In some embodiments, determining, based on the context information associated with the interactive interface, the target response mode for the first input image includes: in response to the context information indicating that the first input image is associated with the health consultation process or a relevance of a historical message to the first input image reaches a threshold, determining that the first input image corresponds to a second response mode, the second response mode indicating to provide follow-up content.

[0137] In some embodiments, different target response modes correspond to different processing entities for generating the first response content.

[0138] In some embodiments, the first interactive scene indicates to identify a symptom associated with the first input image, and the first response content includes at least one of: a symptom description associated with the symptom in the first input image; cause information associated with the symptom; suggestion information associated with the symptom.

[0139] In some embodiments, the first interactive scene indicates to obtain an explanation corresponding to a report in the first input image, and the first response content includes explanation text about the report.

[0140] In some embodiments, the first interactive scene indicates to identify a drug object in the first input image, and the first response content includes description text of the drug object.

[0141] In some embodiments, the first response content is determined based on a drug instruction associated with the drug object in the first input image.

[0142] FIG. 6 A block diagram illustrating an electronic device 600 in which one or more embodiments of the disclosure can be implemented is shown. It should be understood that FIG. 6 The electronic device 600 shown is merely exemplary and should not be construed as limiting the scope of the embodiments described herein. FIG. 6 The electronic device 600 shown can be used to implement FIG. 1 a terminal device 110.

[0143] As FIG. 6 shown, the electronic device 600 is in the form of a general electronic device. Components of the electronic device 600 can include, but are not limited to, one or more processors or processing units 610, a memory 620, a storage device 630, one or more communication units 640, one or more input devices 650, and one or more output devices 660. The processing unit 610 can be a real or virtual processor and is capable of executing various processing in accordance with programs stored in the memory 620. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve parallel processing capability of the electronic device 600.

[0144] The electronic device 600 typically includes a number of computer storage media. Such media can be any available media that is accessible by the electronic device 600 and includes both volatile and non-volatile media, removable and non-removable media. The memory 620 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 630 can be a removable or non-removable media and can include machine-readable media such as a flash drive, a magnetic disk drive, or any other media that can be used to store information and / or data and that can be accessed by the electronic device 600.

[0145] The electronic device 600 can further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown, the electronic device 600 can further include a network-based storage device. The network-based storage device can include a cloud storage device, and can be accessed via the communication unit 640. FIG. 6As shown in FIG. 12, a disk drive 621 or CD drive 622 can be provided that enable reading from or writing to a removable, non- volatile disk, such as a "floppy drive" or a "click-together" CD or DVD drive. In these instances, each drive can be connected to the system bus 620 by one or more data media interfaces. The memory 620 can include a computer program product 625 having one or more program modules configured to carry out the various methods or actions of the various embodiments of the present disclosure.

[0146] The communication unit 640 enables communication with other electronic devices over a communication medium. Additionally, the functionality of the components of the electronic device 600 can be implemented in a single computing cluster or a plurality of computer machines that are capable of communicating with one another over a communication connection. As such, the electronic device 600 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network nodes in the networking environment.

[0147] The input device 650 can be one or more input devices, such as a mouse, a keyboard, a trackball, etc. The output device 660 can be one or more output devices, such as a display, a speaker, a printer, etc. The electronic device 600 can also communicate with one or more external devices (not shown), such as a storage device, a display device, etc., through the communication unit 640, as needed, with one or more devices that enable a user to interact with the electronic device 600, or with any devices (e.g., a network card, a modem, etc.) that enable the electronic device 600 to communicate with one or more other electronic devices. Such communication can be carried out via an input / output (I / O) interface (not shown).

[0148] According to an example implementation of the present disclosure, a computer readable storage medium is provided having computer executable instructions stored thereon, where the computer executable instructions are executed by a processor to implement the method described above. According to an example implementation of the present disclosure, a computer program product is also provided that is tangibly stored on a non-transitory computer readable medium and includes computer executable instructions, where the computer executable instructions are executed by a processor to implement the method described above.

[0149] Various aspects of the disclosure can be described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, systems, and computer program products according to this disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer readable program instructions.

[0150] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0151] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0152] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple implementations of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part for a module, program segment or instruction, and a part for a module, program segment or instruction comprises one or more executable instructions for realizing the logical function of the specification. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the function or action of the specification, or can be realized by a combination of special hardware and computer instructions.

[0153] While various implementations of the present disclosure have been described above, the foregoing description is intended to be illustrative, not exhaustive, and not limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is selected to best explain the principles of the implementations, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. An information processing method, comprising: Presenting an interactive interface with the virtual object, the interactive interface including an image input control; receiving a first input image via the image input control; as well as In response to the first input image being determined to match the first interactive scene, first response content for the first input image is provided in the interactive interface, where the first response content is generated by a first processing module corresponding to the first interactive scene.

2. The method according to claim 1, wherein The interactive interface further includes a set of interactive entrances corresponding to a set of preset interactive scenarios, the set of preset interactive scenarios including the first interactive scenario, and the method further includes: Based on triggering a target interactive portal in the set of interactive portals, acquiring a second input image; and Second response content for the second input image is provided in the interactive interface, where the second response content is generated by a second processing module corresponding to a second interactive scene, and the second interactive scene corresponds to the target interactive entrance.

3. The method according to claim 1, further comprising: The first input image is provided to a visual language model to determine the first interaction scene matching the first input image from a plurality of preset interaction scenes.

4. The method according to claim 1, wherein The first response content is generated based on the following process: determining, based on context information associated with the interactive interface, a target response mode for the first input image, the target response mode indicating whether to provide follow-up content; as well as The first response content generated based on the target response mode is obtained.

5. The method according to claim 4, wherein The determining, based on the context information associated with the interactive interface, a target response mode for the first input image includes any one of the following: In response to the context information indicating that the first input image corresponds to a first round of interaction with the virtual object, determining that the first input image corresponds to a first response mode, the first response mode indicating that no follow-up content is provided; In response to the context information indicating that the interactive interface is independent of a health consultation process, determining that the first input image corresponds to the first response mode; In response to the context information indicating that the correlation between the historical messages of the interactive interface and the first input image is lower than a threshold, it is determined that the first input image corresponds to the first response mode.

6. The method according to claim 5, wherein: The determining, based on context information associated with the interactive interface, a target response mode for the first input image includes: In response to the context information indicating that the first input image is associated with the health consultation process or the correlation between the historical message and the first input image reaches the threshold, it is determined that the first input image corresponds to a second response mode, and the second response mode indicates providing follow-up content.

7. The method according to claim 4, wherein: Different target response modes correspond to different processing entities for generating the first response content.

8. The method according to claim 1, wherein The first interaction scenario indicates identifying a symptom associated with the first input image, and the first response content includes at least one of the following: a symptom description associated with the symptom in the first input image; causal information associated with the described symptoms; Advisory information associated with the described symptoms.

9. The method according to claim 1, wherein The first interactive scenario indicates obtaining an explanation corresponding to a report in the first input image, and the first response content includes an explanation text about the report.

10. The method according to claim 1, wherein The first interactive scenario indicates identifying a medicine object in the first input image, and the first response content includes a description text of the medicine object.

11. The method according to claim 10, wherein: The first response content is determined based on a drug instruction sheet associated with the drug object.

12. An information processing device comprising: An interface presentation module is configured to present an interaction interface with the virtual object, wherein the interaction interface includes an image input control; an image receiving module, configured to receive a first input image via the image input control; as well as The first content providing module is configured to provide first response content for the first input image in the interactive interface in response to the first input image being determined to match the first interactive scene, wherein the first response content is generated by a first processing module corresponding to the first interactive scene.

13. An electronic device comprising: at least one processing unit; as well as At least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to any one of claims 1 to 11 when executed by the at least one processing unit.

14. A computer-readable storage medium having a computer program stored thereon, wherein the computer program can be executed by a processor to implement the method according to any one of claims 1 to 11.

15. A computer program product tangibly stored in a computer storage medium and comprising computer executable instructions which, when executed by a device, cause the device to perform the method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Knowledge question and answer method, device and equipment and storage medium

    CN116561276A

  • Medical visual question and answer method and system based on multi-task modeling

    CN119202334A