Interaction processing method and device and electronic equipment
By analyzing images on electronic devices and providing editable recommendations, the challenge of accurately locating and focusing on issues in intelligent customer service has been solved, improving the accuracy of responses and user satisfaction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- LENOVO (BEIJING) LTD
- Filing Date
- 2026-01-25
- Publication Date
- 2026-05-08
AI Technical Summary
Existing intelligent customer service systems struggle to accurately pinpoint and focus on user issues, resulting in inaccurate responses.
By acquiring images through the interface of electronic devices, analyzing the images using machine learning models, and providing recommendations, users can edit the images and recommendations to select the response that best suits their needs.
This improved the accuracy of the responses, making the recommended information more aligned with user needs and enhancing the user experience.
Smart Images

Figure CN121996128A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to an interactive processing method, apparatus and electronic device. Background Technology
[0002] Currently, when users encounter problems, they usually provide text descriptions or take pictures to the intelligent customer service, which then provides answers based on the text or images, such as how to solve system malfunctions or how to handle injuries.
[0003] Since intelligent solutions are based on the accurate positioning of the problem, they require a high level of expertise in accurately identifying, focusing on, and updating the problem in a timely manner. Summary of the Invention
[0004] In view of the above, this application provides an interactive processing method, apparatus, and electronic device, as follows:
[0005] An interactive processing method, comprising:
[0006] On the first interface, obtain the first image;
[0007] Based on the first model, the first image is analyzed to obtain at least one first recommendation; the first recommendation is related to at least one image object determined by analyzing the first image.
[0008] On the second interface, output the at least one first recommendation message;
[0009] Based on the target recommendation information determined by the user from the at least one first recommendation information, the response result corresponding to the target recommendation information is output.
[0010] Optionally, before outputting the response result corresponding to the target recommendation information based on the target recommendation information determined by the user from the at least one first recommendation information, the method further includes:
[0011] In response to the triggering operation of the editing control in the second interface, the editing data of the first image is obtained in the third interface;
[0012] Based on the first model, the edited data is analyzed to obtain at least one second recommendation message;
[0013] On the second interface, output the at least one piece of second recommendation information;
[0014] In response to the user's second selection operation on the at least one second recommendation, the target recommendation is determined.
[0015] Optionally, before outputting the response result corresponding to the target recommendation information based on the target recommendation information determined by the user from the at least one first recommendation information, the method further includes:
[0016] In response to a user's first selection operation on the at least one first recommendation information, the target recommendation information is determined;
[0017] The method further includes:
[0018] From the second interface, a first selection operation corresponding to the target recommendation information is received;
[0019] or,
[0020] In response to the triggering operation of the confirmation control output in the second interface, the at least one second recommendation message is output in the fourth interface;
[0021] On the fourth interface, a first selection operation corresponding to the target recommendation information is received.
[0022] Optionally, in the above method, obtaining editing data for the first image on the third interface includes:
[0023] On the third interface, output at least one of the following: a selection control, a text editing control, and an image cropping control;
[0024] In response to the line input via the selection control, the selected portion of the image region in the first image is obtained;
[0025] In response to text input through the text editing control, the text content added to the first image is obtained;
[0026] In response to a cropping box input via the image cropping control, a cropped portion of the image region in the first image is obtained.
[0027] Optionally, the above method analyzes the first image based on the first model to obtain at least one first recommendation message, including:
[0028] Based on the first model, the first image is analyzed according to the image recognition algorithm corresponding to the target scene identifier to obtain image features;
[0029] Based on the image features, at least one first recommendation message is selected from the information database corresponding to the target scene identifier;
[0030] The information database includes multiple recommendation entries; the target scene identifier represents the business scene type corresponding to the first image.
[0031] Optionally, the above method may involve filtering at least one first recommendation from the information database corresponding to the target scene identifier based on the image features, including:
[0032] From the information database corresponding to the target scene identifier, determine the target database corresponding to the target recommendation identifier; the target database includes at least one recommendation message.
[0033] Based on the image features, at least one first recommendation message is selected from the target library;
[0034] The target recommendation identifier is obtained through the first interface, and the target recommendation identifier represents the information type.
[0035] Optionally, in the above method, the output position of the first recommendation information on the second interface is: the display position of the image object related to the first recommendation information in the first image.
[0036] Optionally, in the above method, the output position of the second recommendation information on the second interface is: the display position of the image object related to the second recommendation information in the first image.
[0037] An interactive processing device, comprising:
[0038] An image acquisition unit is used to acquire a first image on a first interface;
[0039] An information processing unit is configured to analyze the first image based on a first model to obtain at least one first recommendation message; the first recommendation message is related to at least one image object determined by analyzing the first image.
[0040] An information recommendation unit is used to output the at least one first recommendation message on the second interface;
[0041] The response output unit is used to output the response result corresponding to the target recommendation information based on the target recommendation information determined by the user from the at least one first recommendation information.
[0042] An electronic device, comprising:
[0043] A display, used to output the first and second interfaces;
[0044] The processor is configured to: acquire a first image on the first interface; analyze the first image based on a first model to obtain at least one first recommendation; the first recommendation is related to at least one image object determined by analyzing the first image; output the at least one first recommendation on the second interface; and output a response result corresponding to the target recommendation based on the target recommendation determined by the user from the at least one first recommendation.
[0045] As can be seen from the above technical solutions, in the interactive processing method, apparatus, and electronic device disclosed in this application, after obtaining a first image on a first interface, the first image is analyzed based on a first model, and the obtained recommendation information is output through a second interface. Then, based on the target recommendation information determined by the user from this recommendation information, the corresponding response result is output on a third interface. It is evident that this application provides recommendation information to the user to facilitate the selection of recommendation information, thereby making the determined target recommendation information more in line with the user's needs. This improves the accuracy of the output response result by enhancing the accuracy of the recommendation information. Attached Figure Description
[0046] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0047] Figure 1 A flowchart illustrating an interactive processing method provided in an embodiment of this application;
[0048] Figure 2 This is an example diagram of obtaining the first image in an embodiment of this application;
[0049] Figure 3 This is an example diagram showing the output of the first recommendation information in an embodiment of this application;
[0050] Figure 4 This is an example diagram showing the output response results in an embodiment of this application;
[0051] Figure 5 A partial flowchart of an interactive processing method provided in an embodiment of this application;
[0052] Figure 6 This is an example diagram of editing the first image in an embodiment of this application;
[0053] Figure 7 This is an example diagram showing the output of the second recommendation information in an embodiment of this application;
[0054] Figure 8 This is an example diagram illustrating the output of a response result based on the second recommendation information output from the edited first image in an embodiment of this application.
[0055] Figure 9 This is another example diagram illustrating the output of a response result based on the second recommendation information output from the edited first image in this application embodiment;
[0056] Figure 10 This is an example diagram illustrating the output of response results based on the first recommendation information in an embodiment of this application;
[0057] Figure 11 This is another example diagram illustrating the output of response results based on the first recommendation information in this application embodiment;
[0058] Figure 12 This is another part of the flowchart of an interactive processing method provided in an embodiment of this application;
[0059] Figure 13 This is an example diagram illustrating the selection of business scenario types in the embodiments of this application;
[0060] Figure 14 This is an example diagram illustrating the selection of information type after selecting a business scenario type in this embodiment of the application;
[0061] Figure 15 This is a schematic diagram of the structure of an interactive processing device provided in an embodiment of this application;
[0062] Figure 16 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;
[0063] Figure 17 , Figure 18 and Figure 19 These are example diagrams illustrating how this application applies to interactions on mobile phones within IT service business scenarios. Detailed Implementation
[0064] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0065] refer to Figure 1 The diagram shown illustrates an implementation flowchart of an interactive processing method provided in this application. This method is applicable to electronic devices capable of data processing, such as mobile phones, tablets, or laptops. The technical solution in this embodiment primarily aims to improve the accuracy of the output response results.
[0066] Specifically, the method in this embodiment may include the following steps:
[0067] Step 101: Obtain the first image on the first interface.
[0068] The first image can be an image obtained by the user triggering the image acquisition device to perform real-time image acquisition on the first interface, or it can be an image uploaded by the user through image selection on the first interface. Accordingly, the first interface is the interface that can obtain the first image.
[0069] Specifically, the first interface displays a capture control for the image acquisition device and an upload control for image uploading. The user can click the capture control, and the electronic device, in response, captures the first image through the image acquisition device. Alternatively, the user can click the upload control, and the electronic device, in response, displays a photo album browsing page. The user can select an image on the browsing page to upload, thus obtaining the uploaded first image. After obtaining the first image, it is displayed on the first interface to present the image to be analyzed.
[0070] It should be noted that in this embodiment, the first interface can be output in response to the user's startup operation. The startup operation refers to the user clicking on the function buttons in the interactive interface of the interactive intelligent agent implemented in this embodiment.
[0071] For example, taking a mobile phone as an electronic device, such as Figure 2 As shown, after the interactive agent is activated on the mobile phone, an interactive interface is output. The interactive interface has an input box and an image input control. After the user clicks the image input control, the mobile phone outputs a first interface, which deploys a capture control and an upload control. The user can click the capture control, and the mobile phone captures the first image through the camera; or, the user can click the upload control, and the mobile phone outputs a photo album browsing page. The user can select one of the images on the browsing page to upload, and the mobile phone can obtain the uploaded first image, which is displayed on the first interface.
[0072] In addition, the interactive interface can include other controls, such as speech-to-text controls and send controls. When the user clicks the speech-to-text control, the electronic device uses its microphone to capture speech and converts it into text as input information, which is then provided to the backend model for inference. When the user clicks the send control, the electronic device can send the text and / or image from the text input box to the backend model for inference.
[0073] Step 102: Based on the first model, analyze the first image to obtain at least one first recommendation message.
[0074] The first model is capable of image analysis and outputting recommendation information. This first model can be a machine learning model that recognizes natural language and / or other inputs (such as audio, video, images, tables, etc.) and performs comprehensive language processing tasks such as semantic analysis and question answering, thereby generating input-related outputs. The first model can be a generative model or a generative language model (GLM). For example, it can include large language models (LLMs) and GPT (Generative Pre-trained Transformer).
[0075] For example, the first image is a screenshot of a phone malfunctioning. By analyzing the first image using the first model, at least one first recommendation can be obtained. The first recommendation can be the question that the user needs to express based on the first image, such as "How to solve the problem?" or "How to enable the function?"
[0076] It should be noted that the first recommendation information is related to at least one image object identified by analyzing the first image. For example, the first image, after being analyzed by the first model, can identify at least one image object. An image object can be understood as a content object contained in the first image, such as a keyboard, indicator lights, or a touchpad. The first recommendation information output by the first model is recommendation information specific to the image objects contained in the first image, such as recommendations on how to clean the keyboard or the type of fault indicated by the indicator lights.
[0077] Step 103: On the second interface, output at least one first recommendation message.
[0078] The second interface can be an interface that displays the first image. Specifically, the second interface can be an interface that displays the first image and outputs the first recommendation information after the first image has been obtained through the first interface. For example, ... Figure 3 As shown, after the camera is triggered to capture the first image through the first interface, the first image is displayed on the first interface. After analysis by the first model, at least one first recommendation information obtained by the first model is output on the first interface with the first image through layer overlay, such as "How to reinstall the J keycap?", "Do the keycaps on the keyboard need to be replaced due to wear?", "How to clean the hair in the gaps of the keyboard?", etc., thus forming the second interface.
[0079] In addition, before obtaining the first recommendation information, a prompt message such as "Recognizing" is output on the second interface, indicating that the electronic device is analyzing the first image.
[0080] It should be noted that there is also a retake control on the second interface. After the user clicks the retake control, the electronic device can return to the first interface to re-capture or upload the image.
[0081] Step 104: Based on the target recommendation information determined by the user from at least one first recommendation information, output the response result corresponding to the target recommendation information.
[0082] In one implementation, step 104 can process the target recommendation information through a second model to output the corresponding response result.
[0083] The second model can be the same as the first model, or the second model can be a different model. The agent implemented in this embodiment can call the first model and the second model to perform image analysis and output a response.
[0084] Specifically, in this embodiment, the response result corresponding to the target recommendation information can be output in the interactive interface. The interactive interface can be the interface output after the interactive intelligent agent implemented in this embodiment is started. For example, such as Figure 4 As shown, after the user selects the target recommendation information, the first image and the target recommendation information are displayed in the input box of the interactive interface. After the user clicks the send control, the agent processes the target recommendation information and the first image through the second model and outputs a response. For example, in response to the target recommendation information "Do the keyboard keycaps need to be replaced due to wear?", the output is "Based on the degree of wear of the keycaps, it is recommended to replace the keyboard".
[0085] It should be noted that the second interface can be displayed as a layer over the first interface, and the first interface can be overlaid as a layer over the interactive interface. Based on this, after the user confirms the target recommendation information, they can jump from the original interface to the interactive interface, or the interactive interface can be overlaid on the original interface using layers, thereby outputting the response result through the interactive interface.
[0086] Alternatively, the first interface, the second interface, and the interactive interface can be tiled on the screen to output corresponding content to the user.
[0087] In one implementation, after determining the target recommendation information, the target recommendation information and the first image are displayed in the input box of the interactive interface. The user can edit the target recommendation information again in the input box so that the target recommendation information in the input box can more accurately express the user's needs. Thus, after the user clicks the send control, the electronic device can process the target recommendation information and the first image through the second model to output the response result.
[0088] As can be seen from the above technical solution, in the interactive processing method provided by this application embodiment, after obtaining the first image on the first interface, the first image is analyzed based on the first model, and the obtained recommendation information is output through the second interface. Then, based on the target recommendation information determined by the user from this recommendation information, the corresponding response result is output on the third interface. It is evident that this embodiment provides recommendation information to the user to facilitate the selection of recommendation information, thereby making the determined target recommendation information more in line with the user's needs. This improves the accuracy of the output response result by enhancing the accuracy of the recommendation information.
[0089] In one implementation, before outputting the response result in step 104, the target recommendation information can be determined in the following ways: Figure 5 As shown:
[0090] Step 501: In response to the triggering operation of the editing control in the second interface, obtain the editing data of the first image in the third interface.
[0091] In addition to outputting the first recommended information, the second interface also outputs an editing control. If the user feels that the first recommended information output by the second interface does not meet their needs, they can click the editing control to trigger it. Based on this, the electronic device responds to the trigger operation and outputs a third interface. The third interface is an interface that allows the user to edit the first image. The third interface can be overlaid on the first interface as a layer. After the user performs editing operations on the third interface, the electronic device can obtain the corresponding editing data, which includes a portion of the image area obtained by editing the first image, and / or, the text content added to the first image.
[0092] For example, the first image is output on the third interface, and the user can perform at least one editing operation on the third interface, such as selecting the image area, cropping the image area, or adding text, thereby obtaining the corresponding editing data.
[0093] In one implementation, step 501 can be achieved in the following way:
[0094] First, on the third interface, output at least one of the following: a selection control, a text editing control, and an image cropping control;
[0095] For example, such as Figure 6 As shown, the selection control is an operation space for selecting a portion of the image area in the first image by circling lines, the text editing control is a control for adding text to the first image, and the image cropping control is a control for cropping the image area of the first image by using a cropping box;
[0096] Then, at least one of the following processes can be performed:
[0097] In response to the lines input via the selection control, the selected portion of the image region in the first image is obtained;
[0098] In response to text entered via the text editing control, obtain the text content added to the first image;
[0099] In response to a cropping box entered via an image cropping control, a cropped portion of the image region in the first image is obtained.
[0100] For example, such as Figure 6 As shown, the user selects the image area where the button "J" is located in the first image using the selection control. Based on this, the corresponding editing data can be obtained in this embodiment.
[0101] Step 502: Based on the first model, analyze and edit the data to obtain at least one second recommendation.
[0102] Specifically, in this embodiment, the first model analyzes the selected or cropped image regions and / or added text content in the edited data to obtain the second recommendation information.
[0103] The second recommendation may contain the same information as the first recommendation, or it may contain different information. Furthermore, the number of second recommendations may be the same as or different from the number of first recommendations.
[0104] As can be seen, in this embodiment, the user can edit the first image, causing the first model to re-analyze and thereby update the recommendation information, i.e., the second recommendation information.
[0105] Step 503: On the second interface, output at least one second recommendation message.
[0106] For example, such as Figure 7 As shown, after the user selects a portion of the first image area in the third interface (image editing interface), the user returns to the second interface (image display interface) where the first image is output. On the second interface, the second recommended information replaces the original first recommended information, such as "How to fix the J key on the keyboard?", "Does the J key need to be replaced if it is faulty?", "Why can't I press the J key on the keyboard?", etc.
[0107] As can be seen, under this implementation method, users can make the specific image content to be analyzed more focused by selecting, cropping or adding text content to the image area. This makes the analyzed second recommendation information more in line with the user's questioning needs in the current business scenario. Based on this, the subsequent model can also provide users with more accurate analysis and response results when performing response analysis.
[0108] Step 504: In response to the user's second selection action on at least one second recommendation, determine the target recommendation.
[0109] In one implementation, the user can directly click on one of the second recommended information on the second interface. The selected second recommended information is the target recommended information. Based on this, in this embodiment, a click operation (i.e., a second selection operation) in response to the target recommended information can be received from the second interface. The target recommended information is then determined in response to this second selection operation. Furthermore, in response to the second selection operation, the target recommended information is displayed in the input box of the fourth interface. After the user clicks the send control on the fourth interface, the electronic device, in response to the trigger operation of the send control, processes the first image and the target recommended information through the second model and outputs the response result corresponding to the target recommended information.
[0110] For example, such as Figure 8 As shown, the user directly clicks on one of the second recommended messages, "Why can't I press the J key on the keyboard?", on the second interface. The clicked second recommended message is identified as the target recommended message and is displayed in the input box of the fourth interface (interactive interface). At the same time, the edited first image (such as lines with a circled image, a cropping box, or added text content) is also displayed in the editing area of the fourth interface. After the user clicks the "send" control on the fourth interface, the interactive agent processes the first image and the target recommended message through the second model and outputs a response, such as "The keycaps may be worn out; it is recommended to replace the keyboard."
[0111] In another implementation, the user can first click a confirmation control on the second interface. Based on this, the electronic device can respond to the triggering operation of the confirmation control output on the second interface and output the second recommendation information output on the second interface on the fourth interface. Then, the user can click on one or more of the second recommendation information on the fourth interface; the selected second recommendation information is the target recommendation information. There can be one or more target recommendation information. Therefore, in this embodiment, the click operation (i.e., the second selection operation) corresponding to the target recommendation information can be received from the fourth interface. In response to this second selection operation, the target recommendation information is determined. Furthermore, in response to the second selection operation, the target recommendation information is displayed in the input box on the fourth interface. After the user clicks the send control on the fourth interface, the electronic device responds to the triggering operation of the send control, processes the first image according to the target recommendation information using the second model, and outputs the response result corresponding to the target recommendation information.
[0112] For example, such as Figure 9As shown, after the user clicks the confirmation control "Still want to use" on the second interface (image display interface), a thumbnail of the first image and the second recommended information output by the second interface are output in the editing area of the fourth interface (interactive interface), such as "How to fix the J key on the keyboard?", "Does the J key need to be replaced if it is faulty?", "Why can't I press the J key on the keyboard?". The user can select one or more of the second recommended information, such as "Why can't I press the J key on the keyboard?", and the clicked second recommended information is determined as the target recommended information. The target recommended information is displayed in the input box. After the user clicks the send control "Send", the interactive agent processes the first image and the target recommended information through the second model and outputs a response result, such as "The keycaps may be worn out. It is recommended to replace the keyboard".
[0113] The fourth interface can be the interactive interface output after the interactive agent is started. On the fourth interface, the second recommendation information and the first image can be displayed in the editing area corresponding to the input box, such as... Figure 9 As shown in the diagram, the user can select any one or more second recommended information items in the editing area by clicking. The second recommended information clicked by the user is the target recommended information. In response to the user's click, the electronic device outputs the determined target recommended information in the input box of the fourth interface. Based on this, after the user clicks the send control in the fourth interface, the interactive agent can use the second model to reason about the first image according to the target recommended information to output a response result.
[0114] As can be seen, in this embodiment, after obtaining the first image, the first model analyzes the first image to provide the user with first recommendation information. If the user is not satisfied with the first recommendation information, they can edit the first image. This allows the first model to focus its analysis on the image content edited by the user, such as the circled image area or the added text content. By analyzing the edited data obtained after editing, the first model can provide the user with new recommendation information that better meets their needs, namely, second recommendation information. Thus, the user's editing makes the new recommendation information output by the first model more in line with the user's intention, thereby enabling more accurate positioning of recommendation information and making the output response result more accurate after processing the recommendation information.
[0115] In one implementation, before outputting the response result in step 104, the target recommendation information can be determined in the following way:
[0116] In response to the user's first selection of at least one primary recommendation, the target recommendation is determined.
[0117] In one implementation, the user can directly click on one of the first recommended information on the second interface. The selected first recommended information is the target recommended information. Based on this, in this embodiment, the click operation (i.e., the first selection operation) in response to the target recommended information can be received from the second interface. In response to this first selection operation, the target recommended information is determined. Furthermore, in response to the first selection operation, the target recommended information is displayed in the input box of the fourth interface. After the user clicks the send control on the fourth interface, the electronic device, in response to the trigger operation of the send control, processes the first image and the target recommended information through the second model and outputs the response result corresponding to the target recommended information.
[0118] For example, such as Figure 10 As shown, the user directly clicks on one of the first recommended messages on the second interface. The clicked first recommended message is identified as the target recommended message, such as "Do the keyboard keycaps need to be replaced due to wear?". The target recommended message is displayed in the input box of the fourth interface (interactive interface), and the first image is also displayed in the editing area. After the user clicks the "send" control on the fourth interface, the interactive agent processes the first image and the target recommended message through the second model and outputs a response result, such as "It is recommended to replace the keyboard".
[0119] In another implementation, the user can first click a confirmation control on the second interface. Based on this, the electronic device can respond to the trigger operation of the confirmation control output on the second interface and output the first recommended information output on the second interface on the fourth interface. Then, the user can click on one or more of the first recommended information on the fourth interface; the selected first recommended information is the target recommended information. There can be one or more target recommended information. Based on this, in this embodiment, the click operation corresponding to the target recommended information (i.e., the first selection operation) can be received from the fourth interface, and the target recommended information is determined in response to the first selection operation. Furthermore, in response to the first selection operation, the target recommended information is displayed in the input box on the fourth interface. After the user clicks the send control on the fourth interface, the electronic device responds to the trigger operation of the send control, processes the first image according to the target recommended information using the second model, and outputs the response result corresponding to the target recommended information.
[0120] For example, such as Figure 11As shown, after the user clicks the confirmation control "Still want to use" on the second interface (image display interface), a thumbnail of the first image and the first recommended information output by the second interface are output in the editing area of the fourth interface (interactive interface), such as "How to clean hair from the gaps in the keyboard?", "Do the keycaps on the keyboard need to be replaced due to wear?", "How to reinstall the J keycaps?" etc. The user can select one or more of the first recommended information, such as "Do the keycaps on the keyboard need to be replaced due to wear?" to click. The clicked first recommended information is determined as the target recommended information and is displayed in the input box. After the user clicks the send control "Send", the interactive agent processes the first image and the target recommended information through the second model and outputs a response result, such as "It is recommended to replace the keyboard".
[0121] As can be seen, in this embodiment, recommended information can be selected directly on the second interface, or, if the recommended information on the second interface is confirmed to be available, recommended information can be selected on the fourth interface. This provides users with multiple ways to select recommended information, thereby enriching the user experience.
[0122] In one implementation, step 102 can be achieved in the following way, such as... Figure 12 As shown:
[0123] Step 1201: Based on the first model, analyze the first image according to the image recognition algorithm corresponding to the target scene identifier to obtain image features.
[0124] Image features can include characteristics such as color, texture, shape, and spatial relationships. In this embodiment, an appropriate image recognition algorithm can be selected according to the target scene identifier. Based on this, the first model uses the image recognition algorithm to analyze the first image, thereby obtaining the image features.
[0125] It should be noted that the target scene identifier represents the business scene type corresponding to the first image, and different business scene types correspond to different image recognition algorithms. For example, business scene types can include IT service type, reading analysis type, poster design type, etc., and correspondingly, the image recognition algorithms corresponding to each business scene type are different.
[0126] Specifically, in this embodiment, multiple business scenario identifiers can be provided to the user through the fourth interface. Each business scenario identifier corresponds to a different business scenario type, for example, such as Figure 13 As shown, multiple business scenario types are presented to the user through a drop-down menu in the fourth interface (interactive interface). The user can select one of the business scenario types from the drop-down menu. Based on this, in this embodiment, in response to the user's selection operation of the drop-down menu, the business scenario identifier corresponding to the business scenario type selected is determined as the target scenario identifier.
[0127] Furthermore, in this embodiment, after determining the target scene identifier, it can be provided to the second model. When the response result is obtained in step 104, the second model can output the response result based on the target scene identifier, target recommendation information and the first image.
[0128] Step 1202: Select at least one first recommendation from the information database corresponding to the target scene identifier according to the image features.
[0129] The information database contains multiple recommended information entries, with different databases corresponding to different business scenario types. For example, there is a question database for IT services, a translation database for reading analysis, and an image database for poster design. Based on this, after obtaining image features, the first model selects the first recommended information from the database corresponding to the target scenario identifier.
[0130] It should be noted that the correlation between the selected first recommendation information and the image features meets the selection criteria. These criteria can be: the correlation between the first recommendation information and the image features is greater than or equal to a correlation threshold; or, the correlation between the first recommendation information and the image features is ranked in descending order and is among the top N, where N is a positive integer greater than or equal to 1.
[0131] The relevance value can be the correlation coefficient between the first recommendation information and the first image, such as the Pearson correlation coefficient. Alternatively, the relevance value can be the similarity value between the first recommendation information and the first image.
[0132] Based on the above implementation, step 1202, when filtering the first recommended information, can be achieved in the following way:
[0133] First, determine the target library corresponding to the target recommendation identifier from the information library corresponding to the target scene identifier; the target library includes at least one recommendation information; then, according to image features, filter at least one first recommendation information from the target library.
[0134] Among them, the target recommendation identifier represents the information type, and different information types correspond to different target libraries within their respective business scenario types. Furthermore, different business scenario types can correspond to different information types. Taking the IT service scenario type as an example, the recommended information is divided into multiple information types, such as general types, startup error types, device failure types, and software exception types. Correspondingly, the information library includes target libraries for each of these types.
[0135] Specifically, the target recommendation identifier can be obtained through the first interface, which is the interface for obtaining the first image. In addition to the data collection and upload controls, the first interface also outputs a list of information types, such as... Figure 14 As shown, taking the IT service scenario type as an example, after selecting the IT service type through the fourth interface (interactive interface) and clicking "image input control", the user can select one of the information types from the general type, startup error type, device failure type and software exception type output in the first interface. In response to the type selection operation, the electronic device can determine the recommendation identifier corresponding to the information type selected by the user as the target recommendation identifier, such as the target recommendation identifier representing the general type.
[0136] In another implementation, in step 102, the first image can be analyzed by the first model according to the image recognition algorithm corresponding to the target scene identifier to obtain image features. Combined with the recommendation information contained in the information database corresponding to the target scene identifier, data reasoning is performed to generate multiple candidate recommendation information that match the first image, and at least one first recommendation information with a reasoning probability higher than the probability threshold is selected.
[0137] For example, the first model can be a generative model. Based on this, the first model can use the recommended information in the information database as training samples and the image features of the first image as input to perform data reasoning, so as to infer the question information that may appear in the first image, thereby generating candidate recommended information, and determining the candidate recommended information with a higher reasoning probability as the first recommended information.
[0138] In one implementation, obtaining the second recommendation information in step 502 can be achieved in the following way:
[0139] Based on the first model, the image recognition algorithm corresponding to the target scene identifier is used to analyze the edited data, such as the selected image region and the added text content, to obtain image features; then, at least one second recommendation is selected from the information database corresponding to the target scene identifier according to the image features.
[0140] It should be noted that the specific implementation method of step 502 can refer to the specific implementation method of step 102 in the previous text, and will not be described in detail here.
[0141] In one implementation, when the first recommendation information is output on the second interface, the output position of the first recommendation information on the second interface is: the display position of the image object related to the first recommendation information in the first image. That is, the output position of the first recommendation information on the second interface matches the output position of the first object on the second interface, where the first object is the image object in the first image corresponding to the first recommendation information.
[0142] For example, such as Figure 10 As shown, the first recommendation message, "How to clean hair from keyboard crevices?", is displayed in the image object "hair" in the first image, and in the display position on the second interface.
[0143] In one implementation, the output position of the second recommendation information on the second interface is: the display position of the image object related to the second recommendation information in the first image. That is, the output position of the second recommendation information on the second interface matches the output position of the second object on the second interface, where the second object is the image object in the first image corresponding to the second recommendation information.
[0144] For example, such as Figure 8 As shown, the second recommendation message "Does the J key malfunction require a keycap replacement?" generated because "J" is circled on the first image is displayed around the position where "J" is displayed on the second interface.
[0145] refer to Figure 15 This is a schematic diagram of the structure of an interactive processing device provided in an embodiment of this application. The device can be an interactive intelligent agent deployed in an electronic device. The device in this embodiment may include the following structure:
[0146] The image acquisition unit 1501 is used to acquire a first image on the first interface;
[0147] The information processing unit 1502 is configured to analyze the first image based on a first model to obtain at least one first recommendation message; the first recommendation message is related to at least one image object determined by analyzing the first image.
[0148] Information recommendation unit 1503 is used to output the at least one first recommendation message on the second interface;
[0149] The response output unit 1504 is used to output the response result corresponding to the target recommendation information based on the target recommendation information determined by the user from the at least one first recommendation information.
[0150] In this embodiment, the target recommendation information and the first image can be processed by the second model to output a response result. Specifically, the first model and the second model can be models deployed in the agent, or the first model and the second model can be independently deployed models, and the agent can call the first model and the second model to achieve the corresponding data processing; or the first model can be a model deployed in the agent, and the second model can be an independently deployed model, such as a large language model.
[0151] As can be seen from the above technical solution, in the interactive processing device provided in this application embodiment, after obtaining a first image on a first interface, the first image can be analyzed based on a first model, and the obtained recommendation information can be output through a second interface. Then, based on the target recommendation information determined by the user from this recommendation information, the corresponding response result is output on a third interface. It is evident that this embodiment provides recommendation information to the user to facilitate the selection of recommendation information, thereby making the determined target recommendation information more in line with the user's needs. This improves the accuracy of the output response result by enhancing the accuracy of the recommendation information.
[0152] In one implementation, before the response output unit 1504 outputs the response result corresponding to the target recommendation information based on the target recommendation information determined by the user from the at least one first recommendation information, it is further configured to: in response to a trigger operation on the editing control in the second interface, obtain editing data of the first image in the third interface; analyze the editing data based on the first model to obtain at least one second recommendation information; output the at least one second recommendation information in the second interface; and determine the target recommendation information in response to a second selection operation by the user on the at least one second recommendation information.
[0153] In one implementation, before the response output unit 1504 outputs the response result corresponding to the target recommendation information based on the target recommendation information determined by the user from the at least one first recommendation information, it is further configured to: determine the target recommendation information in response to the user's first selection operation on the at least one first recommendation information;
[0154] The response output unit 1504 is further configured to: receive a first selection operation corresponding to the target recommendation information from the second interface; or, in response to a trigger operation of a confirmation control output in the second interface, output the at least one piece of second recommendation information on the fourth interface; and receive the first selection operation corresponding to the target recommendation information on the fourth interface.
[0155] In one implementation, when the response output unit 1504 obtains editing data for the first image on the third interface, it is specifically used to: output at least one of a selection control, a text editing control, and an image cropping control on the third interface; obtain a selected portion of the image area in the first image in response to a line input through the selection control; obtain text content added to the first image in response to text input through the text editing control; and obtain a cropped portion of the image area in the first image in response to a cropping box input through the image cropping control.
[0156] In one implementation, the information processing unit 1502 is specifically used to: analyze the first image based on a first model and according to an image recognition algorithm corresponding to the target scene identifier to obtain image features; and filter at least one first recommendation information from the information database corresponding to the target scene identifier according to the image features; wherein the information database includes multiple recommendation information; and the target scene identifier represents the business scene type corresponding to the first image.
[0157] Specifically, when the information processing unit 1502 filters at least one first recommendation information from the information library corresponding to the target scene identifier according to the image features, it is used to: determine the target library corresponding to the target recommendation identifier from the information library corresponding to the target scene identifier; the target library includes at least one recommendation information; filter at least one first recommendation information from the target library according to the image features; wherein the target recommendation identifier is obtained through the first interface, and the target recommendation identifier represents the information type.
[0158] In one implementation, the output position of the first recommendation information on the second interface is: the display position of the image object related to the first recommendation information in the first image.
[0159] In one implementation, the output position of the second recommendation information on the second interface is: the display position of the image object related to the second recommendation information in the first image.
[0160] It should be noted that the specific implementation of each unit in this embodiment can be referred to the corresponding content above, and will not be described in detail here.
[0161] refer to Figure 16 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device can be a mobile phone, tablet, or other device, and may include the following structure:
[0162] Display 1601 is used to output at least the first interface and the second interface;
[0163] The processor 1602 is configured to: obtain a first image on the first interface; analyze the first image based on a first model to obtain at least one first recommendation; output the at least one first recommendation on the second interface; and output a response result corresponding to the target recommendation based on the target recommendation determined by the user from the at least one first recommendation.
[0164] As can be seen from the above technical solution, in the electronic device provided in this application embodiment, after obtaining a first image on a first interface, the first image can be analyzed based on a first model, and the obtained recommendation information can be output through a second interface. Then, based on the target recommendation information determined by the user from this recommendation information, the corresponding response result is output on a third interface. It is evident that this embodiment provides recommendation information to the user to facilitate the selection of recommendation information, thereby making the determined target recommendation information more in line with the user's needs. This improves the accuracy of the output response result by enhancing the accuracy of the recommendation information.
[0165] Taking an IT service scenario as an example, the technical solution of this application is illustrated below:
[0166] Firstly, in IT service scenarios, many ordinary users lack the technical expertise to accurately and clearly describe system malfunctions, interface anomalies, and operational errors in writing. These users typically only perceive superficial phenomena such as unusable functions or interface problems, struggling to articulate key information like fault characteristics, error message details, and operational context. This prevents backend customer service or intelligent diagnostic systems from quickly pinpointing the core issue, prolonging problem-solving cycles and degrading the user experience. Simultaneously, users have a need for convenient problem description and adjustment, hoping to accurately convey their needs without complex input, enabling efficient diagnosis and resolution.
[0167] In view of this, this application addresses the core needs of novice users in IT service scenarios by constructing a closed-loop interactive mechanism of "image acquisition and recognition - intelligent question recommendation - selection and adjustment of key areas - interactive confirmation - accurate diagnosis". The key points are as follows:
[0168] 1. Image acquisition and business scenario binding: Supports users to take photos or upload images. The backend model (i.e., the first model) selects suitable recognition algorithms based on IT service business scenarios (such as desktop maintenance, application scenario failure, network anomalies, etc.) to improve the accuracy of image feature extraction.
[0169] 2. Intelligent Problem Point Annotation and Recommendation: After the backend model parses the image, it combines the service's common problem library (i.e., information library) and accurately annotates three high-probability problems based on image features (such as "abnormal error pop-up content", "unresponsive interface buttons", and "abnormal hardware connection icons"), and provides concise text descriptions for users to view intuitively.
[0170] 3. Point-and-click problem confirmation interaction: Users can quickly confirm their core needs by clicking on the problems marked in the image. The backend model (i.e., the second model) generates targeted diagnostic and solution plans based on the problem points selected by the user, combined with image features and business scenarios, realizing the linkage of "image-text-diagnosis".
[0171] 4. Selection-based problem adjustment and supplementation: If the recommended problems do not meet the user's actual needs, the user can select the target area in the image by drawing circles on the screen or add text descriptions. The backend model analyzes the features of the selected area in real time and regenerates the problem points and diagnostic solutions that are suitable for that area. Multiple selections and adjustments are supported until the user's needs are accurately matched.
[0172] Taking a mobile phone as an example, the following is the interaction process of a user in an IT service business scenario through an interactive intelligent agent:
[0173] A. The user opens the interactive smart agent and switches the top tab (drop-down menu) to the "IT Services" category, such as... Figure 17 As shown in 'a'.
[0174] B. The user clicks the image recognition and diagnostic button in the lower right corner of the screen to activate the corresponding function, such as... Figure 17 As shown in b in the figure.
[0175] C. Users can take pictures of device malfunctions on the photo-taking page and select different malfunction categories to help describe the problem, such as... Figure 17 The error message may be displayed as "General" or "Startup Error" in the C++ version.
[0176] D. After the user selects the wrong category, click the white camera button to take a picture, such as... Figure 17 As shown in d.
[0177] E. After the user takes a photo, the intelligent agent automatically enters the recognition page, such as... Figure 17 As shown in 'e', at this point, the backend model analyzes the images taken by the user and outputs three recommendation questions.
[0178] F. Recognition successful. Three recommended questions will automatically appear on the screen, such as... Figure 17 As shown in f in the figure.
[0179] G. When a user believes the recommended question is incorrect or needs more precise location marking, the user can click the pencil button in the lower left corner to make modifications, as shown in a of 18.
[0180] H. Enter the editing page, such as Figure 18 As shown in b in the figure.
[0181] I. Users can click the first selection button in the bottom function area and draw a circle on the screen with their finger to complete the selection, such as... Figure 18 As shown in c, at this point, the backend model analyzes the image with the circled lines and outputs a new recommendation question.
[0182] J. After the user clicks, the screen will output a new recommendation question based on the selected area, such as... Figure 18As shown in d.
[0183] K. Among these more precise recommended questions, users click on a specific recommended question, which then loads a more precise question description and image (with a selection line) into the IT service input box. For example... Figure 18 As shown in e.
[0184] L. Users can click the send button in the lower right corner to send the question and image to the backend processing system (i.e., the second model, such as the large language model). Figure 18 As shown in f in the figure.
[0185] M. Based on the precise question and image submitted by the user, output the corresponding solution, such as... Figure 19 As shown in the image, one round of interaction is complete.
[0186] It is evident that this application has the following advantages:
[0187] 1. This application can lower the user's operating threshold and adapt to the needs of novice users: users do not need to have professional knowledge or input cumbersome text. They can accurately convey their needs through the extremely simple operation of "viewing and selecting" and "circling and focusing", which solves the core pain point of novice users having difficulty in describing their problems and greatly improves the user service experience.
[0188] 2. This application can improve diagnostic accuracy and shorten the problem-solving cycle: By using "intelligent recommendation of problem points + user interaction confirmation", the core of the problem can be accurately identified, avoiding diagnostic bias caused by redundant image information or vague user descriptions in the background model, thus improving the pertinence of the diagnostic solution and shortening the average problem-solving time by more than 40%.
[0189] 3. This application can dynamically adapt to diverse needs and has strong interactive flexibility: the selection and adjustment function supports users to actively correct the problem location, adapt to complex scenarios such as mismatch between recommended problem points and multiple concurrent problems, and can also be compatible with problem identification in all business scenarios of IT services. Its generalization ability far exceeds that of existing single parsing solutions.
[0190] 4. This application can construct a two-way interactive closed loop to optimize model iteration: the interactive data of user clicks and selections can be fed back to the background model to continuously optimize the problem point recommendation algorithm and recognition accuracy, forming a virtuous cycle of "interaction-diagnosis-iteration" and improving the system's adaptability in the long term.
[0191] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.
[0192] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0193] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0194] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An interactive processing method, comprising: On the first interface, obtain the first image; Based on the first model, the first image is analyzed to obtain at least one first recommendation message; The first recommendation information is related to at least one image object determined by analyzing the first image; On the second interface, output the at least one first recommendation message; Based on the target recommendation information determined by the user from the at least one first recommendation information, the response result corresponding to the target recommendation information is output.
2. The method according to claim 1, before outputting the response result corresponding to the target recommendation information based on the target recommendation information determined by the user from the at least one first recommendation information, the method further includes: In response to the triggering operation of the editing control in the second interface, the editing data of the first image is obtained in the third interface; Based on the first model, the edited data is analyzed to obtain at least one second recommendation message; On the second interface, output the at least one piece of second recommendation information; In response to the user's second selection operation on the at least one second recommendation, the target recommendation is determined.
3. The method according to claim 1 or 2, before outputting the response result corresponding to the target recommendation information based on the target recommendation information determined by the user from the at least one first recommendation information, the method further includes: In response to a user's first selection operation on the at least one first recommendation information, the target recommendation information is determined; The method further includes: From the second interface, a first selection operation corresponding to the target recommendation information is received; or, In response to the triggering operation of the confirmation control output in the second interface, the at least one second recommendation message is output in the fourth interface; On the fourth interface, a first selection operation corresponding to the target recommendation information is received.
4. The method according to claim 2, wherein obtaining editing data for the first image on the third interface includes: On the third interface, output at least one of the following: a selection control, a text editing control, and an image cropping control; In response to the line input via the selection control, the selected portion of the image region in the first image is obtained; In response to text input through the text editing control, the text content added to the first image is obtained; In response to a cropping box input via the image cropping control, a cropped portion of the image region in the first image is obtained.
5. The method according to claim 1, wherein based on the first model, the first image is analyzed to obtain at least one first recommendation message, including: Based on the first model, the first image is analyzed according to the image recognition algorithm corresponding to the target scene identifier to obtain image features; Based on the image features, at least one first recommendation message is selected from the information database corresponding to the target scene identifier; The information database includes multiple recommendation entries; The target scene identifier represents the business scene type corresponding to the first image.
6. The method according to claim 5, wherein at least one first recommendation information is selected from the information database corresponding to the target scene identifier according to the image features, comprising: From the information database corresponding to the target scene identifier, determine the target database corresponding to the target recommendation identifier; The target database includes at least one recommendation message; Based on the image features, at least one first recommendation message is selected from the target library; The target recommendation identifier is obtained through the first interface, and the target recommendation identifier represents the information type.
7. The method according to claim 1, wherein the output position of the first recommendation information on the second interface is: the display position of the image object related to the first recommendation information in the first image.
8. The method according to claim 2, wherein the output position of the second recommendation information on the second interface is: the display position of the image object related to the second recommendation information in the first image.
9. An interactive processing device, comprising: An image acquisition unit is used to acquire a first image on a first interface; An information processing unit is configured to analyze the first image based on a first model to obtain at least one first recommendation message. The first recommendation information is related to at least one image object determined by analyzing the first image; An information recommendation unit is used to output the at least one first recommendation message on the second interface; The response output unit is used to output the response result corresponding to the target recommendation information based on the target recommendation information determined by the user from the at least one first recommendation information.
10. An electronic device, comprising: A display, used to output the first and second interfaces; A processor for obtaining a first image on the first interface; Based on the first model, the first image is analyzed to obtain at least one first recommendation message; The first recommendation information is related to at least one image object determined by analyzing the first image; on the second interface, the at least one piece of the first recommendation information is output; Based on the target recommendation information determined by the user from the at least one first recommendation information, the response result corresponding to the target recommendation information is output.