Information processing method
By selecting an object during the display output process of an electronic device to generate a second object and calling the second model to perform information processing tasks, the problem of complex interaction in the prior art is solved, and more efficient information processing and convenient user interaction are achieved.
Patent Information
- Application Number
- CN202511073201.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-11-07
AI Technical Summary
In existing technologies, electronic devices require users to input questions before they can call data processing models to process information, resulting in a complex and inconvenient interaction process.
During the display output process of electronic devices, an object is selected through a target operation, a second object is generated using a first model, and these objects are displayed in the display window to call the second model to perform information processing tasks, thus simplifying the interaction process between the user and the device.
It improves the ease of interaction between users and electronic devices, reduces operational complexity, enhances information processing efficiency and response speed, and strengthens the intuitiveness and convenience of information processing.
Smart Images

Figure CN120909477A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of information processing, and in particular to an information processing method. BACKGROUND
[0002] At present, many electronic devices are configured with data processing models with inference capability. After a user inputs information to be processed, the electronic device can call the data processing model to perform a corresponding inference task based on the information. For example, the user inputs an image of a kettle and inputs a question "What is the price of the kettle?", and the electronic device can call the data processing model to process the image and the question to obtain a corresponding reply.
[0003] In the related technology described above, the electronic device needs to call the data processing model for processing after the user inputs the question, and the interaction process between the user and the electronic device is relatively complex, and the convenience is relatively low. SUMMARY
[0004] Therefore, the present application discloses the following technical solutions:
[0005] The first aspect of the present application provides an information processing method, comprising:
[0006] In the display output process of the electronic device, a target operation is obtained; the target operation is used to select at least one first object displayed by the electronic device;
[0007] At least one second object is generated based on the first model and the selected first object;
[0008] In the display window of the at least one first object, the at least one second object is displayed; the at least one second object can be triggered to call a second model to perform a corresponding information processing task, and the first model and the second model are one model or different models.
[0009] Optionally, the obtaining of the target operation in the display output process of the electronic device comprises:
[0010] In the process of displaying and outputting the image collected by the collection module in real time, the target operation is obtained, and the first object is an object in the image displayed by the electronic device in real time.
[0011] Optionally, the generating of the at least one second object based on the first model and the selected first object comprises:
[0012] At least two second objects are generated based on the first model and the selected at least two first objects, and the at least two second objects include a first type of second object and a second type of second object;
[0013] The method further comprises:
[0014] in response to the first type of second object being triggered, invoking the second model to perform a non-associated information processing task related to only one of the first objects;
[0015] in response to the second type of second object being triggered, invoking the second model to perform an associated information processing task related to at least two of the first objects.
[0016] Optionally, the invoking the second model to perform an associated information processing task related to at least two of the first objects comprises:
[0017] displaying at least one sub-second object, the at least one sub-second object being able to be triggered to invoke the second model to perform a corresponding associated task, different sub-second objects corresponding to different associated information processing tasks;
[0018] in response to the sub-second object being triggered, invoking the second model to perform an associated information processing task corresponding to the triggered sub-second object;
[0019] the sub-second object being generated before the second type of second object is triggered or being generated after the second type of second object is triggered.
[0020] Optionally, the generating at least one sub-second object comprises:
[0021] obtaining two candidate information sets matched with two of the first objects, each of the first objects matching one candidate information set;
[0022] extracting target candidate information between the two candidate information sets with a similarity greater than a threshold value, to generate at least one sub-second object according to the target candidate information.
[0023] Optionally, the generating at least one second object based on the first model and the selected first object comprises:
[0024] obtaining at least one candidate information matched with the selected first object;
[0025] extracting at least one keyword corresponding to the at least one candidate information according to the first model;
[0026] generating at least one second object corresponding to the at least one keyword based on the at least one keyword, each of the keywords being used to generate one corresponding second object.
[0027] Optionally, the obtaining at least one candidate information matched with the selected first object comprises:
[0028] identifying the selected first object to obtain corresponding object information;
[0029] The object information is compared with a plurality of pre-stored information to determine at least one candidate information matching the selected first object from the plurality of pre-stored information.
[0030] Optionally, the comparing the object information with the plurality of pre-stored information to determine at least one candidate information matching the selected first object from the plurality of pre-stored information comprises:
[0031] determining a plurality of initial information according to the similarity between the object information and the plurality of pre-stored information;
[0032] determining at least one candidate information matching the selected first object from the plurality of initial information according to historical data, the historical data representing the frequency of performing an information processing task corresponding to different initial information based on the second model.
[0033] Optionally, the at least one second object comprises three types of second objects, each of the three types of second objects being associated with a candidate information.
[0034] The method further comprises:
[0035] in response to the three types of second objects being triggered, obtaining to-be-processed information according to the candidate information associated with the triggered three types of second objects and the object information of the first object;
[0036] invoking the second model to perform an information processing task corresponding to the to-be-processed information.
[0037] Optionally, the at least one second object comprises four types of second objects.
[0038] The method further comprises:
[0039] in response to the four types of second objects being triggered, performing a prompt operation to obtain input information input by a user;
[0040] obtaining to-be-processed information according to the input information and the object information of the first object;
[0041] invoking the second model to perform an information processing task corresponding to the to-be-processed information.
[0042] The second aspect of the present application provides an information processing apparatus, comprising:
[0043] an obtaining unit configured to obtain a target operation in an electronic device display output process; the target operation is used to select at least one first object displayed by the electronic device;
[0044] a generating unit configured to generate at least one second object based on a first model and the selected first object;
[0045] a processing unit configured to display the at least one second object in a display window of the at least one first object; the at least one second object can be triggered to invoke a second model to perform a corresponding information processing task, and the first model and the second model are one model or different models. BRIEF DESCRIPTION OF DRAWINGS
[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative effort on the basis of the provided drawings.
[0047] Figure 1 is a flowchart of an information processing method provided by an embodiment of the present application;
[0048] Figure 2 is a schematic diagram of a first display interface provided by an embodiment of the present application;
[0049] Figure 3 is a schematic diagram of a second display interface provided by an embodiment of the present application;
[0050] Figure 4 is a schematic diagram of a third display interface provided by an embodiment of the present application;
[0051] Figure 5 is a schematic diagram of a fourth display interface provided by an embodiment of the present application;
[0052] Figure 6 is a structural schematic diagram of an information processing device provided by an embodiment of the present application;
[0053] Figure 7 is a structural schematic diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0054] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative effort fall within the scope of the present application.
[0055] The present embodiment provides an information processing method, please refer to Figure 1 is a flowchart of the method, which can include the following steps.
[0056] S101, in the process of displaying and outputting the electronic device, a target operation is obtained; the target operation is used for selecting at least one first object displayed by the electronic device.
[0057] The electronic device can be a device with a display function and a user interaction interface. For example, the electronic device can be a smart phone, a tablet computer, a computer, etc.
[0058] The first object can be a text object or an image object displayed by the electronic device.
[0059] The text object refers to information content expressed in the form of text presented on the interface displayed by the electronic device. The text object can be a word, a sentence or a paragraph. For example, the text object can be the conversation content selected by the target operation in the text chat conversation picture displayed and output by the electronic device, or the text paragraph selected by the target operation in the text picture displayed and output by the electronic device.
[0060] The image object refers to information content presented in the form of graphics on the interface of the electronic device. The image object can correspond to an actual object existing in a real environment, or a virtual object in a virtual environment. For example, the image object can be a cup selected by the target operation in the camera picture displayed and output by the electronic device, or a virtual character selected by the target operation in the animation picture displayed and output by the electronic device.
[0061] It should be noted that the image object can contain text content. For example, when the image object selected by the target operation is a beverage bottle, the beverage bottle packaging can contain text content such as ingredient composition and production information.
[0062] The target operation can be a specific interaction action of the user on the electronic device for the at least one first object, used for selecting the at least one first object and triggering subsequent response. The target operation can be clicking, selecting or long pressing.
[0063] The target operation can be input through various input devices, which can be a touch screen, a mouse or a keyboard.
[0064] In some examples, the user can input an information processing start instruction by triggering an information processing shortcut key on the electronic device. For example, an information processing shortcut key can be displayed on the touch screen of the electronic device, and after the user clicks the information processing shortcut key, the electronic device can determine the first object in the currently displayed content based on the target operation.
[0065] S102, at least one second object is generated based on the first model and the selected first object.
[0066] The first model refers to a large language model (LLM) used to process the selected first object. The first model can identify the selected first object and determine the specific information content of the first object. For example, the first model can identify that the first object is a "water bottle", an "apple", a "user interface (UI)", text, or code.
[0067] The second object is an element generated by the first model based on the selected first object and capable of further interaction with the user.
[0068] It should be noted that the second object of the embodiment is not a preset second object, but a dynamically generated second object based on the currently selected first object. On the one hand, the generated second object can be different for different first objects, and on the other hand, the second object generated after this selection and the second object generated after the next selection can be different for the same first object.
[0069] The generation method of the second object can be: the first model identifies the first object to obtain an identification result, searches for task information matching the identification result according to the identification result, and generates a second object corresponding to the task information.
[0070] If the selected first object is a text object or an image object, the first model receives the first object as input, the identification result is the text contained in the text object or the image object, at least one task information matching the text contained in the text object or the image object is searched, at least one second object is generated based on the at least one task information, and the second object can include a second object associated with the task information or a second object not associated with the task information.
[0071] For example, the selected first object is the text in the interface shown in Figure 2 The identification result can be the text "copy file A to file C".
[0072] If the selected first object is an image object, the first model receives the first object as input, the identification result is object information of the first object, at least one task information matching the object information is searched, a second object corresponding to each task information is generated, and a custom second object is generated.
[0073] The object information can be the name of the item of the first object. For example, the selected first object is the water bottle shown in Figure 3 The obtained identification result can be "water bottle".
[0074] S103, display at least one second object in the display window of the at least one first object; the at least one second object can be triggered to call a second model to perform a corresponding information processing task, and the first model and the second model are one model or different models.
[0075] The display mode of the at least one second object can be that the second object is added to a picture displayed by the electronic device, and the second object can be displayed around the position of the corresponding first object. For example, the distance between the first object and the multiple second objects corresponding to the first object can be less than a certain threshold.
[0076] The at least one second object can be displayed in the same display window as the first object. For example, the electronic device displays a preview image currently captured by a camera in a preview window of a camera preview interface in real time, and displays at least one second object in the preview window after a first object in the preview image is selected.
[0077] The first object and the multiple second objects correspond to each other, that is, the task information used to generate the second objects is searched based on the identification result of the first object.
[0078] As some examples, refer to FIGS. 1A and 1B. Figure 2 , Figure 2 As shown in FIGS. 1A and 1B, the electronic device displays second objects, and the first model searches for matching task information based on the text of the identified text object or the text of the image object. The task information can include: “task information 1: translate the text into English”, “task information 2: generate code based on the input text”, and “task information 3: reply to the input text”. Based on the task information 1 to 3, three second objects of “translation”, “code”, and “inquiry” are generated in sequence, and a “custom” second object is generated. The first three second objects are associated with the task information 1 to 3 in sequence, and the “custom” second object is not associated with the task information.
[0079] As some examples, refer to FIGS. 2A and 2B. Figure 3 , Figure 3 As shown in FIGS. 2A and 2B, the electronic device displays second objects, and the first model searches for matching task information based on the object information of the identified image object, that is, the name of the selected first object “water bottle”. The task information can include: “task information 1: how much is the price of the product”, “task information 2: how long will it take to complete”, and “task information 3: what is the material of the product”. Based on the task information 1 to 3, three second objects of “price”, “timing”, and “material” are generated in sequence, and a “custom” second object is generated. The first three second objects are associated with the task information 1 to 3 in sequence, and the “custom” second object is not associated with the task information.
[0080] If the information processing task to be performed when a second object is triggered is a task of generating text, such as a task of translating the text of the first object, a task of inquiring about the price of the first object, etc., the second model can be a generative large language model. If the information processing task to be performed when a second object is triggered includes a task that requires execution of a specific operation instruction, such as a task of starting a timer of an electronic device to perform timing, the second model can be a large model having the capability of executing the corresponding instruction, and is not limited to the generative large language model described above.
[0081] The embodiments of the present application can train one large language model to perform object recognition processing tasks, and train another large language model to perform information processing tasks. In addition, the same large language model can also be trained to be capable of performing both object recognition processing tasks and information processing tasks.
[0082] The information processing task can include various types of tasks. For example, the information processing task can include generating corresponding text information from input information and outputting the generated text information, or the information processing task can include generating corresponding text information from input information and executing a specific operation instruction based on the generated text information, such as an operation instruction of starting a timer to perform timing.
[0083] It should be noted that the first model and the second model in the embodiments can be deployed locally on an electronic device for performing the above method, or can be deployed on a cloud device (such as a server device) in communication with the electronic device. In other words, the inference capability required for generating the second object and performing the information processing task can be provided locally by the electronic device performing the above method, or can be provided by the corresponding cloud device.
[0084] The embodiments have the beneficial effect that by displaying the second object based on the target operation in the display output process of the electronic device, the interaction process between the user and the electronic device can be simplified. Specifically, after displaying the second object, the user can directly trigger the second object to call the second model to perform specific information processing, without the need for the user to input specific information, such as inputting a specific question, thereby improving the convenience of the user interacting with the electronic device, reducing the complexity of the user operation, and also helping to improve the response speed and information processing efficiency of the electronic device, so that the user can complete the task more efficiently.
[0085] Optionally, in the display output process of the electronic device, the target operation is obtained, including:
[0086] In the process of displaying the image captured by the acquisition module in real time by the electronic device, the target operation is obtained, and the first object is an object in the image displayed by the electronic device in real time.
[0087] In some embodiments, the electronic device can display the image captured by the image capture module in real time in a camera preview interface. Figure 3
[0088] The image capture module refers to a component in the electronic device for capturing and processing external visual images. The image capture module can be a camera. For example, the image capture module in a smartphone is a camera, which captures images of the surrounding environment in real time, and the image data is output in real time by the touch screen of the smartphone.
[0089] During the process of the electronic device capturing and displaying images of the surrounding environment in real time through the image capture module, the electronic device can actively identify and respond to the target operation of the user, and determine the specific object selected by the user in the image displayed by the electronic device in real time. In this process, the interaction between the user and the electronic device becomes more intuitive, because the user can directly select on the real-time image without inputting complex instructions.
[0090] During the process of the electronic device displaying and outputting the image captured by the image capture module in real time, the user selects a first object in the displayed image through a target operation, and then processes the selected first object using a first model to generate at least one second object. These second objects are then displayed to the user, and the user can trigger these second objects to call a second model to complete an information processing task.
[0091] As some examples, the electronic device displays and outputs images of the surrounding environment captured by the camera in real time, the user selects a first object in the displayed image by performing a target operation on the touch screen of the electronic device, inputs the first object into a first model for identification, obtains an identification result, searches for task information matching the identification result according to the identification result, generates and displays second objects corresponding to the task information for the user to select, and after the user triggers the second object, calls a second model to execute a corresponding information processing task.
[0092] The beneficial effects of the present embodiment are that the image is captured and displayed in real time by the image capture module of the electronic device, and the user can intuitively select and operate the object in front of them at any time and anywhere on the electronic device, thereby improving the interactive convenience of the overall interaction process, not only effectively solving the problem of complex interaction between the user and the electronic device, but also expanding the scenario of information processing by the electronic device using a model.
[0093] In addition to the process of displaying and outputting the image collected by the collection module in real time, the method of the embodiment can also be applied to other display and output processes, not limited to the real-time output case described above, for example, it can also be applied to the process of displaying and outputting any image or video stored in advance by the electronic device, and the process of displaying and outputting the application interface of the application program by the electronic device. Optionally, generating at least one second object based on the first model and the selected first object includes:
[0094] Generating at least two second objects based on the first model and the selected at least two first objects, the at least two second objects including a first type of second object and a second type of second object;
[0095] The method further includes:
[0096] In response to the first type of second object being triggered, calling the second model to perform a non-associated information processing task related to only one first object;
[0097] In response to the second type of second object being triggered, calling the second model to perform an associated information processing task related to at least two first objects.
[0098] If only one first object is selected, only the first type of second object can be generated and displayed, and if multiple first objects are selected, the first type of second object can be generated and displayed for each first object.
[0099] If multiple first objects are selected, the second type of second object can be generated and displayed for each two first objects, that is, the number of second type of second objects can be equal to the number of combinations of two first objects, or one second type of second object can be generated for all first objects.
[0100] The first type of second object refers to an object directly related to the selected at least one first object. The first type of second object is used to provide operation options related to the corresponding first object. Each first type of second object corresponds to only one first object.
[0101] The second type of second object refers to an object related to the selected at least two first objects, that is, one second type of second object can correspond to at least two first objects. The second type of second object is used to provide an information processing task involving at least two first objects, that is, an associated information processing task. For example, by triggering the second type of second object, the second model can be called to process a task of comparing the at least two first objects corresponding to the second type of second object.
[0102] As some examples, please refer to Figure 4 , Figure 4The electronic device displays a first object, a first object, and a second object. The selected first object is a kettle and a pot. The first object generated for the kettle can include "price", "timing", "material", and "customization". The first object generated for the pot can include "material", "usage", "price", and "customization". The second object generated by the first model by associating the kettle and the pot is "comparison".
[0103] The non-associated information processing task refers to a task related to only one first object. The execution of the non-associated information processing task does not involve other objects.
[0104] The associated information processing task refers to a task related to at least two first objects. The execution of the associated information processing task requires comprehensive consideration of the information of multiple first objects. For example, if the selected first objects are apples and bananas, the associated information processing task can be "comparison of the nutritional components of apples and bananas" or "preparation method of mixed juice".
[0105] During the display and output of the electronic device, after the user selects at least two first objects through a target operation, the first model is used to independently process each selected first object to generate a corresponding first object. The first model is used to associate the processing of the at least two selected first objects to generate a corresponding second object. In response to the first object being triggered: the second model is called to execute a non-associated information processing task for only one first object. In response to the second object being triggered: the second model is called to execute an associated information processing task related to both objects.
[0106] The embodiment has the beneficial effect that by supporting the simultaneous selection of at least two first objects, not only can a first object related to each of the objects be generated, but also a second object involving the comprehensive information of at least two first objects can be generated. When the user triggers the first object, the processing result of the non-associated information processing task related to a single first object can be quickly obtained, and when the user triggers the second object, the processing result of the associated information processing task related to multiple objects can be obtained. This improves the convenience of user interaction with the electronic device, provides a more rich information interaction experience, and meets the user's information processing needs for multiple objects.
[0107] Optionally, the second model is called to execute an associated information processing task related to at least two first objects, including:
[0108] At least one sub-second object is displayed. The at least one sub-second object can be triggered to call the second model to execute a corresponding associated task. Different sub-second objects correspond to different associated information processing tasks.
[0109] In response to the sub-second object being triggered, the second model is invoked to perform an associated information processing task corresponding to the triggered sub-second object.
[0110] The sub-second object is generated before the second object of the second type is triggered or after the second object of the second type is triggered.
[0111] The sub-second object refers to an object further refined on the basis of the second object of the second type. The sub-second object can be triggered by the user to invoke the second model to perform a specific associated information processing task. Each sub-second object corresponds to a different associated information processing task, and can provide more refined information processing services for the user.
[0112] The display mode of the at least one sub-second object can be that the sub-second object is displayed around the position of the second object of the second type, that is, the distance between the sub-second object and the second object of the second type can be less than a certain threshold. As an example, the positional relationship between the sub-second object and the second object of the second type can be seen from Figure 5 .
[0113] Each sub-second object can be triggered by the user, which means that the user can interact with these sub-second objects. When the user triggers a certain sub-second object, the second model is invoked to perform an associated information processing task related to the sub-second object.
[0114] The user selects at least one first object through a target operation, generates corresponding second objects based on the first model, and displays these objects on the electronic device. The second object of the second type is further refined into a sub-second object as a comprehensive processing of multiple first objects, so as to provide more specific associated information processing services. When the user triggers the first object of the first type, a non-associated information processing task related to a single first object is performed; and when the user triggers the second object of the second type, the sub-second object is generated so that the user can further select a specific associated information processing task.
[0115] The timing of generating the sub-second object can be before the user triggers the second object of the second type, or after the user triggers the second object of the second type. If the sub-second object is generated before the user triggers the second object of the second type, the electronic device can simultaneously display each first object of the first type, the second object of the second type, and the sub-second object. If the sub-second object is generated after the user triggers the second object of the second type, the electronic device can first display the first object of the first type and the second object of the second type after determining the first object, and the display interface can be as shown in Figure 4 After the second object of the second type is triggered, the electronic device can display the second object of the second type and the corresponding sub-second object together, and can hide the originally displayed first object of the first type, that is, no longer display the first object of the first type, as shown in Figure 5 .
[0116] The beneficial effects of the embodiment are that by displaying at least one sub-second object, the user can more finely select the information they want to process. Each sub-second object can be triggered individually to call the second model to process a specific associated task, and the hierarchical design of the first-class second object, the second-class second object and the sub-second object further simplifies the originally cumbersome user operation, improves the user's interaction experience, and also enhances the accuracy and convenience of information processing, so that the user can more efficiently obtain the information processing result. In addition, the generation of the sub-second object can be performed before or after the user triggers the second-class second object, so that the electronic device can customize the accuracy of information processing according to device resources or user settings, thereby more efficiently serving the needs of the user and promoting the accuracy and diversity of information processing.
[0117] Optionally, the at least one sub-second object is generated by:
[0118] Two candidate information sets matched with the two first objects are obtained, and each first object matches a candidate information set;
[0119] Target candidate information with a similarity greater than a threshold between the two candidate information sets is extracted to generate the at least one sub-second object according to the target candidate information.
[0120] The candidate information set refers to a set constructed according to the task information matched with the selected first object.
[0121] The target candidate information refers to the task information with a similarity greater than a set threshold in the two candidate information sets after comparison and calculation of the similarity.
[0122] The extraction method of the target candidate information can be: extracting the features of each candidate information in the candidate information set, and converting the extracted features into numerical vectors using vectorization technology. For example: TF-IDF (Term Frequency-Inverse Document Frequency) or Word2Vec can be used for vectorization. Then select a similarity calculation method such as cosine similarity or Euclidean distance to calculate the similarity of two task information in each candidate information pair, wherein one task information in the candidate information pair comes from the candidate information set of one first object, and the other task information comes from the candidate information set of another first object. According to the pre-set similarity threshold, the target candidate information with a similarity higher than the threshold is selected.
[0123] When generating the sub-second object in the above manner, the sub-second object can include an object keyword to represent to the user what information processing task can be processed by triggering the sub-second object.
[0124] As some examples, please refer to Figure 5 , the selected first objects are a kettle and a pot, the matched task information searched for the first object "kettle" includes: "task information 1: what is the price of the commodity", "task information 2: how long will it take to finish", and "task information 3: what is the material of the article", then the task information 1 to 3 constitute the candidate information set of the first object "kettle". The matched task information searched for the first object "pot" includes: "task information 4: what is the price of the commodity", "task information 5: what is the material of the article", and "task information 6: how to use this thing", then the task information 4 to 6 constitute the candidate information set of the first object "pot". By comparison, it is found that the similarity of the task information 1 and the task information 4 is greater than the threshold, and the similarity of the task information 3 and the task information 5 is greater than the threshold, then based on the task information 1 and the task information 4, a sub-second object "price" belonging to the second object "comparison" of the second category is generated, and based on the task information 3 and the task information 5, a sub-second object "material" belonging to the second object "comparison" of the second category is generated.
[0125] The above method of generating sub-second objects can be applied to generating sub-second objects belonging to the second object of the second category corresponding to two first objects, and can also be applied to generating sub-second objects belonging to the second object of the second category corresponding to more than two first objects. For the latter case, only need to replace "two candidate information sets matched by two first objects" in the above embodiment with "two candidate information sets matched by multiple first objects", and no further description is given.
[0126] Optionally, generating at least one second object based on the first model and the selected first object includes:
[0127] obtaining at least one candidate information matched with the selected first object;
[0128] extracting at least one keyword corresponding to the at least one candidate information according to the first model;
[0129] generating at least one second object corresponding to the at least one keyword, each keyword is used to generate a corresponding second object.
[0130] The candidate information can be task information used to describe an information processing task and matched with the selected first object. For example, the information processing task can include a translation task, a question and answer task, etc., the candidate information can be task information "translate the text into English" related to the translation task, or task information "what is the price of the commodity" related to the question and answer task. As another example, the task information 1 to 3 matched in the example in the foregoing text are equivalent to the at least one candidate information here.
[0131] In this embodiment, one piece of candidate information can contain object information of a specific object, or can not contain object information of a specific object. As an example, the object information of a specific object can be “how to use a kettle”, and the object information not containing a specific object can be “how much is the price of an article”.
[0132] The matching of a first object and candidate information means that when performing an information processing task for the first object, the information processing task is a task that is probably consistent with the description of the candidate information.
[0133] As some examples, the first object can be an image of a kettle, and when performing an information processing task for the kettle, the information processing task is probably a question and answer task about the price or material of the kettle, but is generally not a translation task about English of the kettle, nor a question and answer task about a person's name; the first object can be a face image, and when performing an information processing task for the person, the information processing task is probably a question and answer task about the person's name or age, but is generally not a question and answer task about material or price.
[0134] Therefore, in the case that the first object is an image of a kettle, the matched candidate information can be “how much is the price of an article”, “how is the material of this thing”, and is basically not “how do you say this thing in English”; in the case that the first object is a face image, the matched candidate information can be “which star is this person”, “how old does this person look”, and is basically not “how much is the price of an article”, “how is the material of this thing”.
[0135] The method of extracting keywords can refer to related technologies, which will not be described here.
[0136] For each candidate information, a second object can be generated based on the keyword corresponding to the candidate information, and the generated second object is associated with the candidate information.
[0137] The generation method can be to call a program in the electronic device for generating a related operation control (such as a generation button), to generate a corresponding operation control with the keyword, and this operation control is the second object generated based on the keyword.
[0138] As an example, for the candidate information “how much is the price of an article”, the extracted keyword can be “price”, and the generated second object can be the “price” second object shown in the figure, and the second object generated in this way is associated with the candidate information used when generating, that is, the “price” second object is associated with the candidate information “how much is the price of an article”. Figure 2
[0139] As described above, the second object can be generated in association with the task information or can be generated without association with the task information.
[0140] If the second object is not associated with the task information, the second object generally needs user input of certain information after being triggered, and the second model is called to perform the corresponding task based on the user input information.
[0141] Therefore, the method of generating the second object without association with the task information can be to obtain a preset input prompt word, call a program in the electronic device for generating a corresponding operation control (for example, a generation button), and generate a corresponding operation control with the input prompt word displayed, which is the second object without association with the task information.
[0142] The input prompt word here can be any prompt word that can play a prompting role, and the prompting role refers to prompting the user to input information after triggering the second object.
[0143] As an example, the preset input prompt word "custom" can be obtained, and then the "custom" second object shown in the figure can be generated. Figure 2
[0144] The above method of generating the second object according to the candidate information can also be applied to the generation of the sub-second object according to the target candidate information in the foregoing embodiments, and will not be described in detail.
[0145] When two or more first objects are selected, the two types of second objects described above can be generated, and the generation method can be to obtain a preset association prompt word, call a program in the electronic device for generating a corresponding operation control (for example, a generation button), and generate a corresponding operation control with the association prompt word displayed, which is the second object. The association prompt word can be any prompt word that can represent joint processing of multiple first objects, for example, "compare", "contrast", and the like.
[0146] The method of obtaining at least one piece of candidate information matched with the first object can be:
[0147] The selected first object is recognized to obtain corresponding object information;
[0148] The object information is compared with multiple pieces of pre-stored information to determine at least one piece of candidate information matched with the selected first object in the multiple pieces of pre-stored information.
[0149] The method of recognizing the object information can refer to the description of the determination of the recognition result in step S102, and the above recognition result is equivalent to the object information of the embodiment.
[0150] The plurality of pieces of pre-stored information can be a plurality of pieces of pre-stored information contained in a pre-constructed information library. The pre-stored information can include information that any user has called the second model to process in the past, for example, including a question that user A has asked the second model in the past, and can also include information generated according to collected corpus, such as corpus collected from network forums, books, and the like. The information generated according to the corpus analysis of information that is usually called to process the second model is taken as pre-stored information.
[0151] The manner of comparing the object information with the plurality of pieces of pre-stored information to determine the candidate information can be:
[0152] Based on any algorithm capable of encoding information in the related art, the object information is encoded into a corresponding object information feature vector, the pre-stored information is encoded into a corresponding pre-stored information feature vector, the similarity between the object information feature vector and the pre-stored information feature vector is calculated, and the plurality of pieces of pre-stored information are sorted in descending order of similarity. The first N pieces of pre-stored information can be used as candidate information matched by the selected first object.
[0153] Processing each selected first object in this manner can obtain candidate information matched by each first object.
[0154] The above N is a predetermined positive integer, and the value thereof is not limited. As an example, N can be equal to 3.
[0155] Optionally, when determining the candidate information, the historical data corresponding to the candidate information can also be further considered to obtain candidate information that is more in line with the user's demand. The specific manner is as follows:
[0156] According to the similarity between the object information and the plurality of pieces of pre-stored information, a plurality of initial information is determined;
[0157] According to the historical data, at least one piece of candidate information matched by the selected first object is determined from the plurality of initial information, and the historical data represents the frequency of performing an information processing task corresponding to different initial information based on the second model.
[0158] The method of determining the plurality of initial information according to the similarity can refer to the method of determining the matched candidate information in the above embodiments, and will not be described in detail.
[0159] The number M of initial information can be greater than the number N of required candidate information, for example, M can be equal to 10.
[0160] The historical data can include a number of times that the information processing task corresponding to the different initial information is performed based on the second model in a recent period of time. For example, an initial information can be "what is the price of the commodity", and the electronic device has answered this question 10 times in the last week, for example, answering "what is the price of the pot", "what is the price of the computer", and so on, a total of 10 times, so the historical data corresponding to this initial information can be 10.
[0161] After obtaining the initial information, the order of the plurality of initial information can be adjusted according to the historical data. The adjustment manner can be to arrange the plurality of initial information in descending order of the historical data, or to appropriately move the initial information with historical data greater than a certain threshold forward and to appropriately move the initial information with historical data less than a certain threshold backward on the basis of the previous order based on the similarity, and the specific adjustment manner is not limited.
[0162] After the adjustment is completed, the first N initial information can be selected as the candidate information according to the adjusted order.
[0163] Optionally, the at least one second object includes three types of second objects, and each of the three types of second objects is associated with a candidate information.
[0164] The method further includes:
[0165] In response to the three types of second objects being triggered, obtaining the to-be-processed information according to the candidate information associated with the triggered three types of second objects and the object information of the first object.
[0166] Calling the second model to perform an information processing task corresponding to the to-be-processed information.
[0167] The method of obtaining the to-be-processed information can be that the first model or the second model is called to identify the to-be-processed candidate information, to determine the to-be-replaced information in the to-be-processed candidate information that needs to be replaced, to replace the to-be-replaced information with the object information of the first object, and to obtain the to-be-processed information after the replacement.
[0168] Alternatively, the method of obtaining the to-be-processed information can be that the first model or the second model is called to insert the object information of the first object into a specified position of the to-be-processed candidate information to obtain the to-be-processed information.
[0169] The to-be-identified candidate information refers to the candidate information associated with the triggered three types of second objects.
[0170] As an example, assuming that the first object is an image of a kettle, after being selected, the object information of the first object is identified as "kettle", and the electronic device displays Figure 3The three second objects shown are "price", "timing" and "material", among which the "price" second object is generated based on the candidate information "what is the price of the product" according to the foregoing method.
[0171] After the "price" second object is triggered, the object information "water bottle" of the corresponding first object is obtained, the to-be-replaced information "product" in the to-be-processed candidate information "what is the price of the product" is replaced with the object information, the to-be-processed information "what is the price of the water bottle" is obtained, and then the second model can be called to perform an information processing task corresponding to "what is the price of the water bottle" to obtain a corresponding processing result. The processing result obtained can be reply information about the price of the water bottle. For example, it can be "the price of this water bottle on the following e-commerce platforms is generally XX yuan, YY yuan, etc.".
[0172] When the second model is called to perform the information processing task, the selected first object itself can be input into the second model in the form of an image for reference by the second model.
[0173] In combination with the above example, when the second model is called to perform the information processing task, the to-be-processed information and the water bottle image cut out from the image displayed on the screen can be input into the second model.
[0174] In some embodiments, after the processing result is obtained based on the to-be-processed information, the processing result can also represent executable operation instructions. In this case, after the processing result is obtained, the processing result can be output, and the operation instructions represented based on the second model can be determined and executed.
[0175] As an example, Figure 3 The second object triggered in the above example is "timing", the candidate information associated with the second object is "how long is it left to boil", the to-be-processed information "how long is it left to boil the water bottle" is obtained by combining the candidate information and the object information of the first object, and the processing result obtained after the second model executes the information processing task corresponding to the to-be-processed information can be "it is left to boil the water bottle for 10 minutes". After the processing result, on the one hand, the processing result can be output, and on the other hand, the operation instructions corresponding to the timing of 10 minutes of the processing result can be determined based on the second model, and then the timing function of the electronic device is started to start 10 minutes of timing.
[0176] Alternatively, if the processing result represents executable operation instructions, the processing result can also not be output, and the corresponding operation instructions can be directly executed. For example, in the above example, the text "it is left to boil the water bottle for 10 minutes" can not be output, and the timing function of the electronic device can be directly started to start 10 minutes of timing.
[0177] In some embodiments, the to-be-processed information can also be acquired in a manner other than the above-described manner, and the candidate information associated with the triggered third-type second object is directly taken as the to-be-processed information. The to-be-processed information and the image of the first object are input into the second model, so as to invoke the second model to perform an information processing task corresponding to the to-be-processed information.
[0178] Optionally, the at least one second object includes a fourth-type second object.
[0179] The method further includes:
[0180] In response to the fourth-type second object being triggered, a prompt operation is performed to obtain input information input by a user.
[0181] The to-be-processed information is acquired according to the input information and the object information of the first object.
[0182] The second model is invoked to perform an information processing task corresponding to the to-be-processed information.
[0183] The fourth-type second object is a second object that is not associated with candidate information. As an example, the fourth-type second object can be Figure 3 a "customized" second object as shown in FIG. 1.
[0184] The input information can be any text information, for example, "what is the price", "how to use this", "translate this", and the like.
[0185] After the input information is acquired, the to-be-processed information can be generated according to the input information and the object information of the first object by using the text generation capability of the first model or the second model, or the to-be-processed information can be acquired in the manner of processing the recognized candidate information in the foregoing embodiments.
[0186] After the to-be-processed information is acquired, the second model can be invoked to perform an information processing task corresponding to the to-be-processed information. The related execution manner can be referred to the foregoing embodiments, and thus will not be described herein.
[0187] The first-type second object in the foregoing embodiments can be divided into a third-type second object and a fourth-type second object according to whether the first-type second object is associated with candidate information.
[0188] On the other hand, the sub-second object in the foregoing embodiments can also be divided into a third-type second object and a fourth-type second object according to whether the sub-second object is associated with candidate information. An example of the former can be Figure 5 the "price" sub-second object and the "material" sub-second object around the "compare" second object in FIG. 1, and an example of the latter can be Figure 5 the "customized" sub-second object around the "compare" second object in FIG. 1.
[0189] If the triggered second object is the third type of second object in the sub second object, the first model or the text generation capability of the second model can be used to generate the to-be-processed information according to the target candidate information associated with the sub second object, the object information of the first object associated with the sub second object, and the association prompt word used to generate the second object of the second type.
[0190] The target candidate information associated with the sub second object can refer to the target candidate information used when generating a sub second object according to the method of the foregoing embodiments. For example, when generating the sub second object of "material" according to the target candidate information "what is the material of the article" of the first object of the kettle and the target candidate information "what is the material" of the first object of the pot, the target candidate information associated with the sub second object of "material" is "what is the material of the article" and "what is the material".
[0191] In combination Figure 5 For example, when the sub second object of "material" is triggered, the target candidate information "what is the material of the article" and "what is the material", the object information of the two associated first objects "kettle" and "pot", and the association prompt word "compare" are obtained, and based on the text generation capability of the first model or the second model, the to-be-processed information "compare the materials of the kettle and the pot" is generated, and then the second model is called to execute the information processing task corresponding to the to-be-processed information, and the processing result obtained is output. The processing result can be "the material of the kettle is..., the material of the pot is..., and the relationship between the two is as follows...".
[0192] If the triggered second object is the fourth type of second object in the sub second object, the input information can be obtained according to the foregoing method, and the to-be-processed information is generated based on the input information, and the second model is called to execute the information processing task corresponding to the to-be-processed information.
[0193] The method of generating the to-be-processed information based on the input information can refer to the method of generating the to-be-processed information when the third type of second object in the sub second object is triggered according to the foregoing embodiments, and only the target candidate information associated with the sub second object needs to be replaced by the obtained input information, and details are not described herein.
[0194] The embodiments of the present application provide an information processing device, please refer to Figure 6 The device can include the following units.
[0195] The obtaining unit 601 is configured to obtain a target operation in an electronic device display output process; the target operation is used to select at least one first object displayed by the electronic device;
[0196] The generating unit 602 is configured to generate at least one second object based on a first model and the selected first object;
[0197] The processing unit 603 is configured to display at least one second object; the at least one second object can be triggered to call the second model to perform a corresponding information processing task, and the first model and the second model are one model or different models.
[0198] Optionally, the obtaining unit 601 obtains the target operation in the process of displaying and outputting the electronic device, and the target operation includes:
[0199] In the process of displaying and outputting the image collected by the collection module in real time by the electronic device, the target operation is obtained, and the first object is an object in the image displayed in real time by the electronic device.
[0200] Optionally, the generating unit 602 generates at least one second object based on the first model and the selected first object, and the generating includes:
[0201] The at least two second objects include a first type of second object and a second type of second object.
[0202] The processing unit 603 is further configured to:
[0203] In response to the first type of second object being triggered, the second model is called to perform a non-associated information processing task related to only one first object.
[0204] In response to the second type of second object being triggered, the second model is called to perform an associated information processing task related to at least two first objects.
[0205] Optionally, the processing unit 603 calls the second model to perform the associated information processing task related to at least two first objects, and the calling includes:
[0206] The processing unit 603 is configured to display at least one sub-second object, the at least one sub-second object can be triggered to call the second model to perform a corresponding associated task, and different sub-second objects correspond to different associated information processing tasks.
[0207] In response to the sub-second object being triggered, the second model is called to perform an associated information processing task corresponding to the triggered sub-second object.
[0208] The sub-second object is generated before the second type of second object is triggered or is generated after the second type of second object is triggered.
[0209] Optionally, the at least one sub-second object is generated by the generating unit 602, and the generating includes:
[0210] Two candidate information sets matched with the two first objects are obtained, and each first object matches one candidate information set.
[0211] extract target candidate information between the two candidate information sets with similarity greater than a threshold value, to generate at least one second object according to the target candidate information.
[0212] Optionally, the generating unit 602 generates at least one second object based on the first model and the selected first object, including:
[0213] obtaining at least one candidate information matched with the selected first object;
[0214] extracting at least one keyword corresponding to the at least one candidate information according to the first model;
[0215] generating at least one second object corresponding to the at least one keyword, each keyword being used to generate a corresponding second object.
[0216] Optionally, the generating unit 602 obtains at least one candidate information matched with the selected first object, including:
[0217] identifying the selected first object to obtain corresponding object information;
[0218] comparing the object information with a plurality of pre-stored information to determine at least one candidate information matched with the selected first object in the plurality of pre-stored information.
[0219] Optionally, the generating unit 602 compares the object information with a plurality of pre-stored information to determine at least one candidate information matched with the selected first object in the plurality of pre-stored information, including:
[0220] determining a plurality of initial information according to the similarity of the object information and the plurality of pre-stored information;
[0221] determining at least one candidate information matched with the selected first object in the plurality of initial information according to historical data, the historical data representing the frequency of performing information processing tasks corresponding to different initial information based on the second model.
[0222] Optionally, the at least one second object includes three types of second objects, each of which is associated with a candidate information;
[0223] The processing unit 603 is further configured to:
[0224] in response to the three types of second objects being triggered, obtaining to-be-processed information according to the candidate information associated with the triggered three types of second objects and the object information of the first object;
[0225] calling the second model to perform an information processing task corresponding to the to-be-processed information.
[0226] Optionally, the at least one second object includes four types of second objects.
[0227] The processing unit 603 is further configured to:
[0228] In response to the four types of second objects being triggered, a prompt operation is performed to obtain input information input by the user;
[0229] Obtain the to-be-processed information according to the input information and the object information of the first object;
[0230] Call the second model to perform an information processing task corresponding to the to-be-processed information.
[0231] The information processing device of the embodiment can refer to the related steps of the information processing method of the foregoing embodiments for working principles, which will not be described herein.
[0232] The embodiment further provides an electronic device, which can refer to Figure 7 The electronic device can include a display screen 701, at least one processor 702, and a memory 703 for storing computer instructions.
[0233] The processor 702 is configured to load the computer instructions described above to perform the following: obtaining a target operation during the display output process of the electronic device; and generating at least one second object based on a first model and a selected first object.
[0234] The display screen 701 is configured to display at least one object and at least one second object in the same window.
[0235] The target operation is configured to select at least one first object displayed by the electronic device. The first model and the second model can be loaded by the processor 702 to implement the processing process described above.
[0236] The working principles of the above electronic device can refer to the related steps of the information processing method of the foregoing embodiments, which will not be described herein.
[0237] It should be noted that each embodiment in the present specification adopts a progressive manner for description, and each embodiment focuses on the different places from other embodiments. The same and similar parts between each embodiment can be referred to.
[0238] For the convenience of description, the above system or device is described as various modules or units respectively described in terms of functions. Of course, the functions of each unit can be implemented in the same or multiple software and / or hardware in the implementation of the present application.
[0239] Those skilled in the art can clearly understand the application by the description of the above embodiments. The technical solutions of the application can be implemented by means of software and necessary universal hardware platforms. Based on such an understanding, the technical solutions of the application can be embodied in the form of a software product, which can be stored in a storage medium, such as a ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the application.
[0240] Finally, it should be noted that the terms such as first, second, third, and fourth, etc. are merely used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between these entities or operations. Moreover, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. Without more limitations, the element defined by the statement "including a" does not exclude the presence of other identical elements in the process, method, article or device including the element.
[0241] The above description is only the preferred embodiments of the application, and it should be pointed out that those skilled in the art can make some improvements and refinements without departing from the principles of the application, and these improvements and refinements should also be regarded as the protection scope of the application.
Claims
1. An information processing method comprising: obtaining a target operation in an electronic device display output process; the target operation is used to select at least one first object displayed by the electronic device; generating at least one second object based on a first model and the selected first object; displaying the at least one second object in a display window of the at least one first object; the at least one second object can be triggered to call a second model to perform a corresponding information processing task, and the first model and the second model are one model or different models.
2. The method of claim 1, wherein the obtaining a target operation in an electronic device display output process comprises: obtaining a target operation in a process of displaying an image collected by an image collection module in real time by the electronic device; and the first object is an object in the image displayed in real time by the electronic device.
3. The method of claim 1, wherein the generating at least one second object based on a first model and the selected first object comprises: generating at least two second objects based on a first model and the selected at least two first objects, the at least two second objects including a first type of second object and a second type of second object; the method further comprises: in response to the first type of second object being triggered, calling a second model to perform a non-associated information processing task related to only one of the first objects; and in response to the second type of second object being triggered, calling a second model to perform an associated information processing task related to at least two of the first objects.
4. The method of claim 3, wherein the calling a second model to perform an associated information processing task related to at least two of the first objects comprises: displaying at least one sub-second object, the at least one sub-second object can be triggered to call a second model to perform a corresponding associated task, and different sub-second objects correspond to different associated information processing tasks; in response to the sub-second object being triggered, calling a second model to perform an associated information processing task corresponding to the triggered sub-second object; and the sub-second object is generated before the second type of second object is triggered or after the second type of second object is triggered.
5. The method of claim 4, wherein the generating at least one sub-second object comprises: obtaining two candidate information sets matched with two first objects, each first object matching a candidate information set; extracting target candidate information with a similarity greater than a threshold value between the two candidate information sets to generate at least one sub-second object according to the target candidate information.
6. The method of claim 1, wherein the generating at least one second object based on a first model and the selected first object comprises: obtaining at least one candidate information matched with the selected first object; extracting at least one keyword corresponding to the at least one candidate information according to the first model; generating at least one second object corresponding to the at least one keyword based on the at least one keyword, each keyword being used to generate a corresponding second object.
7. The method of claim 6, wherein the obtaining at least one candidate information matching the selected first object comprises: identifying the selected first object to obtain corresponding object information; and comparing the object information with a plurality of pre-stored information to determine at least one candidate information matching the selected first object from the plurality of pre-stored information.
8. The method of claim 7, wherein the comparing the object information with a plurality of pre-stored information to determine at least one candidate information matching the selected first object from the plurality of pre-stored information comprises: determining a plurality of initial information according to a similarity between the object information and the plurality of pre-stored information; and determining at least one candidate information matching the selected first object from the plurality of initial information according to historical data, the historical data representing a frequency of performing an information processing task corresponding to different initial information based on the second model.
9. The method of claim 1, wherein the at least one second object comprises three types of second objects, each of the three types of second objects being associated with a candidate information, and the method further comprises: in response to the three types of second objects being triggered, obtaining to-be-processed information according to the candidate information associated with the triggered three types of second objects and the object information of the first object; and invoking the second model to perform an information processing task corresponding to the to-be-processed information.
10. The method of claim 1, wherein the at least one second object comprises four types of second objects, and the method further comprises: in response to the four types of second objects being triggered, performing a prompting operation to obtain input information input by a user; obtaining to-be-processed information according to the input information and the object information of the first object; and invoking the second model to perform an information processing task corresponding to the to-be-processed information.