Information extraction method and device

Through the end-side multimodal large model to recognize image text information and automatically fill it to the target interface, the complex image information extraction operation in the prior art is solved, and efficient information extraction and filling are achieved.

CN120407070APending Publication Date: 2025-08-01VIVO MOBILE COMM CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510543406.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

In the prior art, image information extraction operations are complex, especially when multiple segments of text need to be copied, it is necessary to make multiple jumps between different applications, which reduces the efficiency of information extraction.

Method used

By receiving user input, the text information in the image is identified using the end-side multimodal large model, the associated application and function item identification is displayed, and after the user selects the target identification, the text information is automatically extracted and filled into the corresponding interface, reducing user manual operations.

Benefits of technology

It realizes intelligent extraction and automatic filling of text information in the image, reduces user operation steps, improves information extraction efficiency, and avoids network dependence and privacy risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120407070A_ABST
    Figure CN120407070A_ABST
Patent Text Reader

Abstract

The invention discloses an information extraction method and device, and belongs to the technical field of image recognition. The information extraction method comprises the steps of receiving first input of a user to a target image; in response to the first input, displaying at least one identifier associated with first text information in the target image, the identifier comprising at least one of an icon of the application program and an icon of a function item in the application program; receiving a second input of moving the target image to a target identifier in the at least one identifier by the user; and in response to the second input, second text information in the target image is extracted to a target interface corresponding to the target identifier, and the first text information comprises the second text information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of image recognition, and particularly relates to an information extraction method and apparatus thereof. Background Art

[0002] Currently, image applications are becoming more and more popular on intelligent devices. Compared with text, the information contained in images is more abundant, and moreover, compared with recording information through text, it is easier to record information through images. For example, for business card information, as long as the business card is photographed as an image by a camera, the business card information can be stored in an electronic device in electronic file form.

[0003] On this basis, when the user needs to use the information in the image, the corresponding text information can be obtained by extracting the image content, and then other operations can be performed. For example, to create a new contact based on a business card image, first, as Figure 2 shown, extract the text information in the business card image, and then as Figure 3 shown, copy the required text paragraph by paragraph to the new contact page, thus omitting the process of manual typing.

[0004] However, the above information extraction method is complex in operation. Especially when multiple paragraphs of text need to be copied, it is necessary to jump between different applications multiple times, increasing the user's operation steps and reducing the information extraction efficiency. Summary of the Invention

[0005] The purpose of the embodiments of this application is to provide an information extraction method and apparatus thereof, which can realize intelligent extraction of text information in images, reduce the user's operation steps, and improve the information extraction efficiency.

[0006] In a first aspect, the embodiments of this application provide an information extraction method, which includes: receiving a first input from a user for a target image; in response to the first input, displaying at least one identifier associated with the first text information in the target image, where the identifier includes at least one of an icon of an application program and an icon of a function item in the application program; receiving a second input from the user to move the target image to a target identifier among the at least one identifier; in response to the second input, extracting the second text information in the target image to a target interface corresponding to the target identifier, where the first text information includes the second text information.

[0007] Second aspect, an information extraction device is provided in an embodiment of the present application. The device includes: a receiving unit configured to receive a first input from a user for a target image; a display unit configured to, in response to the first input, display at least one identifier associated with first text information in the target image, where the identifier includes at least one of an icon of an application program and an icon of a function item in the application program; the receiving unit is further configured to receive a second input from the user for moving the target image to a target identifier among the at least one identifier; a processing unit configured to, in response to the second input, extract second text information in the target image to a target interface corresponding to the target identifier, and the first text information includes the second text information.

[0008] Third aspect, an electronic device is provided in an embodiment of the present application. The electronic device includes a processor and a memory. The memory stores a program or instruction that can run on the processor. When the program or instruction is executed by the processor, the steps of the information extraction method in the first aspect are implemented.

[0009] Fourth aspect, a readable storage medium is provided in an embodiment of the present application. A program or instruction is stored on the readable storage medium. When the program or instruction is executed by the processor, the steps of the information extraction method in the first aspect are implemented.

[0010] Fifth aspect, a chip is provided in an embodiment of the present application. The chip includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is configured to run a program or instruction to implement the steps of the information extraction method in the first aspect.

[0011] Sixth aspect, a computer program product is provided in an embodiment of the present application. The program product is stored in a storage medium. The program product is executed by at least one processor to implement the steps of the information extraction method in the first aspect.

[0012] In the information extraction method provided by the embodiments of the present application, a first input of a user for a target image is received; in response to the first input, at least one identifier associated with the first text information in the target image is displayed, and the identifier includes at least one of an icon of an application program and an icon of a function item in the application program; a second input of the user moving the target image to a target identifier among the at least one identifier is received; in response to the second input, second text information in the target image is extracted to a target interface corresponding to the target identifier, and the first text information includes the second text information. Through the above information extraction method, based on the first input of the user for the target image, the identifiers of the application program and its function items associated with the first text information in the target image can be displayed, and then, based on the target identifier selected by the user, the second text information in the target image is extracted to the target interface corresponding to the target identifier. In this way, intelligent extraction of text information in the image is realized, and the text information in the image can be automatically and intelligently filled into the corresponding interface without manual operation by the user, reducing the operation steps of the user and improving the efficiency of information extraction. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 FIG. 1 is one of the schematic flowcharts of the information extraction method provided by the embodiments of the present application;

[0014] Figure 2 FIG. 2 is one of the schematic operation interfaces of the information extraction method in the related art;

[0015] Figure 3 FIG. 3 is another schematic operation interface of the information extraction method in the related art;

[0016] Figure 4 FIG. 4 is one of the schematic operation interfaces of the information extraction method provided by the embodiments of the present application;

[0017] Figure 5 FIG. 5 is another schematic operation interface of the information extraction method provided by the embodiments of the present application;

[0018] Figure 6 FIG. 6 is another schematic operation interface of the information extraction method provided by the embodiments of the present application;

[0019] Figure 7 FIG. 7 is the schematic principle diagram of the information extraction method provided by the embodiments of the present application;

[0020] Figure 8 FIG. 8 is another schematic flowchart of the information extraction method provided by the embodiments of the present application;

[0021] Figure 9 FIG. 9 is the structural block diagram of the information extraction device provided by the embodiments of the present application;

[0022] Figure 10Block diagram of the electronic device provided by the embodiment of the present application;

[0023] Figure 11 Schematic diagram of the hardware structure of the electronic device provided by the embodiment of the present application. Detailed implementation manners

[0024] Next, the technical solutions in the embodiments of the present application will be clearly described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art belong to the scope of protection of the present application.

[0025] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are usually of the same type, and the number of objects is not limited. For example, the first object can be one or multiple. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the associated objects before and after.

[0026] Next, with reference to the accompanying drawings, the information extraction method and its device provided by the embodiments of the present application will be described in detail through specific embodiments and their application scenarios.

[0027] As Figure 1 shown, the embodiment of the present application provides an information extraction method, and the method may include the following S102 to S108:

[0028] S102: Receive a first input from the user for the target image.

[0029] The information extraction method proposed by the embodiment of the present application is executed by an electronic device, which may specifically be a smart electronic device such as a smart phone, a tablet computer, a notebook computer, and a smart watch, and no specific limitation is made here.

[0030] Among them, the above first input is a touch input from the user for the target image, and the touch input may specifically be a click input, a long press input, a double click input, a slide input, an input along a preset trajectory, etc. Those skilled in the art can set the specific form of the above first input according to the actual situation, and no specific limitation is made here.

[0031] Further, the above target image is an image included in the display interface that is real-time opened by the electronic device during the user's use of the electronic device, such as images included in the desktop interface, application interface, chat interface, and other interfaces.

[0032] In the actual application process, the target image can specifically be an image stored in the album, an image shared in the chat, an image included in the application program such as a product introduction image in a shopping application program, a dish introduction image in a food delivery application program, etc., and no specific limitation is made here.

[0033] S104: In response to the first input, display at least one identifier associated with the first text information in the target image.

[0034] Wherein, the first text information is all the text information in the target image.

[0035] Further, one identifier corresponds to one function item, such as function items like creating a new contact, dialing, navigation, sending an email, and sending a message.

[0036] Further, the identifier includes at least one of the icon of the application program and the icon of the function item in the application program. In the case where the application program has only one function, the corresponding identifier is displayed as the icon of the application program. In the case where the application program has multiple functions, the corresponding identifier is displayed as the icon of the specific function item in the application program, so that the user can directly reach the interface that needs to be operated through the input of the identifier.

[0037] Further, when the text information in the target image can be used to execute the function item corresponding to a certain identifier, it means that the text information is associated with the identifier, and one text information can be associated with one or more identifiers. For example, a phone number can be used for the function of creating a new contact and can also be used for the dialing function. The phone number is simultaneously associated with the identifiers of the functions of creating a new contact and dialing.

[0038] Specifically, in the information extraction method provided in the embodiments of the present application, during the user's use of the electronic device, the electronic device can receive the first input of the user on the display interface including the target image. The electronic device responds to the first input, determines at least one function item that the first text information in the target image can be used to execute, and displays at least one identifier corresponding to the at least one function item on the display interface.

[0039] Exemplarily, as Figure 4 shown, a target image 204 is displayed in the display interface 202 of the electronic device, and the target image 204 is a business card image. The electronic device receives the long-press input of the user on the display interface 202, and the electronic device responds to the long-press input, as Figure 5As shown, it is determined that the function items that the first text information 206 in the target image 204 can be used to execute are creating a new contact, sending an email, navigation, and dialing, and the identifiers of the above function items are displayed on the display interface 202.

[0040] S106: Receive a second input from the user to move the target image to a target identifier among at least one identifier.

[0041] Among them, the above second input is a touch input of the user on the target image.

[0042] In the actual application process, the above first input can specifically be a long - press input of the user on the target image in the display interface. On this basis, the above second input can be a touch input when the user continues to drag the target image to the position of the target identifier without releasing the target image during the long - press process.

[0043] Among them, the target identifier corresponds to a target interface.

[0044] Furthermore, the target interface is the page that the electronic device needs to open when the electronic device executes the target function corresponding to the target identifier. For example, if the target function is dialing, the target identifier is the icon of the dialing application, and the target interface is the dialing page.

[0045] S108: In response to the second input, extract the second text information in the target image to the target interface corresponding to the target identifier.

[0046] Among them, the first text information includes the second text information. The second text information can be all the text information in the target image or part of the text information in the target image. The specific content of the second text information is determined by the information required to be filled in the target interface.

[0047] Specifically, in the information extraction method provided in the embodiments of the present application, after displaying at least one identifier associated with the first text information in the target image, the electronic device receives a second input from the user to move the target image to a target identifier among at least one identifier, and the electronic device extracts the second text information in the target image to the target interface corresponding to the target identifier in response to the second input. In this way, on the one hand, all the information required by the target interface can be intelligently extracted from the target image at one time, and the extracted information can be automatically filled into the target interface, which can avoid the user from manually pasting and copying information, and can avoid the jump operation between different applications, reducing the user's operation steps; on the other hand, the second text information is intelligently selected from the first text information of the target image for extraction, which can avoid extracting redundant information, thus avoiding the user from performing secondary editing on the extracted information, further reducing the user's operation steps.

[0048] Exemplarily, such as Figure 5As shown, the electronic device receives a drag input from the user to drag the target image 204 to the position of the target identifier 208. The target identifier 208 is the icon of the function item for creating a new contact. In response to the above drag input, as Figure 6 shown, the second text information 214 in the target image 204 is intelligently extracted into the information display control 212 of the target interface 210. The target interface 210 is the new contact page. Among them, compared with the Figure 5 first text information in, the second text information 214 lacks the redundant information of parentheses, and there is no need for the user to perform secondary editing on the target interface 210 to delete the parentheses, reducing the user's operation steps.

[0049] The information extraction method provided by the embodiments of the present application receives a first input from the user for a target image; in response to the first input, displays at least one identifier associated with the first text information in the target image, where the identifier includes at least one of an icon of an application and an image of a function item in the application; receives a second input from the user to move the target image to a target identifier among the at least one identifier; and in response to the second input, extracts the second text information in the target image into a target interface corresponding to the target identifier, where the first text information includes the second text information. Through the above information extraction method, based on the first input from the user for the target image, the identifiers of the application and its function items associated with the first text information in the target image can be displayed, and then based on the target identifier selected by the user, the second text information in the target image is extracted into the target interface corresponding to the target identifier. In this way, the intelligent extraction of text information in the image is realized, and the text information in the image can be automatically and intelligently filled into the corresponding interface without manual operation by the user, reducing the user's operation steps and improving the efficiency of information extraction.

[0050] In the embodiments of the present application, the above S104 may specifically include the following S104a to S104c:

[0051] S104a: In response to the first input, call the target model and generate a first instruction.

[0052] Among them, the target model is an edge-side model, and the edge-side model is a model that can run independently and complete tasks on the electronic device without the need to connect to the network.

[0053] In the actual application process, the above target model may specifically be an edge-side multi-modal large model, and the multi-modal large model is a large model that can process and understand various types of data such as images, videos, text, etc.

[0054] Further, the first instruction is used to instruct the target model to determine the user intention information according to the first text information in the target image.

[0055] In the actual application process, the first instruction can specifically be "what possible intents are included in the target image", and no specific limitation is made here.

[0056] Specifically, in the information extraction method provided in the embodiments of the present application, during the process of a user using an electronic device, the screen of the electronic device senses the user's actions and determines whether the user's finger is pressing on the screen. When the electronic device detects that the duration for which the user's finger presses on the screen exceeds a preset duration, such as 1 s, the user's action is determined as a long press. At this time, the electronic device identifies its current display interface and determines whether there is an image in the current display interface. In the case where the display interface includes a target image, the electronic device extracts the target image in the display interface, invokes the multi-modal large model on the terminal side, that is, the target model, and generates a first instruction, so that the target model performs intent recognition on the target image according to the first instruction.

[0057] It can be understood that when extracting image text in an application program, for an application program that does not support the text extraction function, it is necessary to call a third-party program to extract the text in the image of the application program. Calling a third-party program not only increases the operation complexity but also involves user privacy issues. Moreover, most of the text extraction technologies in the related art rely on cloud models and cannot provide services in a scenario without network connection.

[0058] However, in the information extraction method provided in the embodiments of the present application, the terminal-side model is called based on the user input to trigger the image understanding ability at the system level. Thus, it is not necessary to consider whether the current application program has the image understanding ability, nor is it necessary to call a third-party program for image understanding, getting rid of the dependence on the capabilities of a single application program. Further, in the information extraction method provided in the embodiments of the present application, using the multi-modal large model on the terminal side for image understanding and information extraction can avoid the usability problems caused by network factors, provide services in a scenario without network connection, and can also not upload user data to the server, thereby protecting user privacy. Moreover, it can also make the types of extracted data no longer restricted.

[0059] S104b: Input the target image and the first instruction into the target model, and obtain the user intent information output by the target model.

[0060] Specifically, in the information extraction method provided in the embodiments of the present application, after generating the first instruction, the target image and the first instruction are input into the target model, so that the target model performs intent recognition on the target image according to the indication of the first instruction, determines the user intent information that the target image may correspond to based on the first text information in the target image, and outputs the user intent information.

[0061] For example, the first text information in the target image includes a phone number, and the user intent information can be creating a new contact, making a call, or sending a message; the first text information in the target image includes address information, and the user intent information can be navigation; the first text information in the target image includes a product name, and the user intent information can be searching for a product.

[0062] S104c: Display at least one identifier according to the user intent information.

[0063] Specifically, in the information extraction method provided in the embodiments of the present application, after obtaining the user intent information output by the target model, for each piece of user intent information, check whether there is an application installed in the electronic device that can implement the user intent information. For example, if the user intent information is navigation, the application that implements this user intent information is a map application; if the user intent information is searching for a product, the application that implements this user intent information is a shopping application. When there is an application installed in the electronic device that can implement the user intent information, display the identifier corresponding to this user intent information on the display interface, that is, display the icon of the application corresponding to this user intent information or the icon of a function item in the application. Otherwise, ignore this user intent information.

[0064] It can be understood that an application can include one or more function items. In the information extraction method provided in the embodiments of the present application, displaying the icon of the application associated with the first text information in the target image or the icon of a function item in the application can make the user intent and function items at the finest granularity, which is convenient for the target model to accurately extract the second text information required to execute the target function from the target image according to the user's needs later.

[0065] In the above embodiments provided by the present application, in response to the first input, the target model is called and a first instruction is generated. The target model is a side model, and the first instruction is used to instruct the target model to determine the user intent information according to the first text information in the target image; the target image and the first instruction are input into the target model, and the user intent information output by the target model is obtained; at least one identifier is displayed according to the user intent information. In this way, using the side model for image understanding can quickly process data locally, reduce data transmission latency and privacy risks, and can more accurately grasp the user's operation intent based on the image text, making the provided identifier more in line with the user's needs.

[0066] In the embodiments of the present application, the above S108 may specifically include the following S108a to S108c:

[0067] S108a: In response to the second input, obtain the control parameters of at least one information display control in the target interface.

[0068] Among them, when the information display control executes the target function corresponding to the target interface, it is the control for the information to be input in the target interface. For example, when the target interface is the new contact page, at this time, the above information display control is the control for inputting information such as name, company, position, landline, email, and address in the new contact page.

[0069] Furthermore, the above control parameter is used to indicate the information type or information name of the information to be input in the information display control, that is, the above control parameter is used to indicate the information type or information name required to execute the target function.

[0070] For example, as Figure 7 shown, the target interface 210 is the new contact page, and this new contact page includes controls for inputting information such as name, company, position, mobile phone, landline, email, fax, and address. At this time, the above "name, company, position, mobile phone, landline, email, fax, and address" are the control parameters corresponding to the target interface 210.

[0071] S108b: According to the control parameter, extract the second text information in the target image to obtain the parameter value of each control parameter.

[0072] Among them, the above parameter value is the specific content to be filled in the information display control corresponding to the control parameter.

[0073] For example, continuing with the above example, if the control parameter is name, the parameter value is the contact name; if the control parameter is company, the parameter value is the company name; if the control parameter is position, the parameter value is the contact position; if the control parameter is mobile phone, the parameter value is the mobile phone number; if the control parameter is landline, the parameter value is the landline number; if the control parameter is email, the parameter value is the email address; if the control parameter is fax, the parameter value is the fax number; if the control parameter is address, the parameter value is the company address.

[0074] S108c: Display the parameter value in the corresponding information display control.

[0075] Among them, the information display control, the control parameter, and the parameter value correspond to each other one by one.

[0076] Specifically, in the information extraction method provided in the embodiments of the present application, the electronic device responds to the second input, obtains the control parameters of at least one information display control in the target interface, extracts the second text information in the target image according to the control parameters to obtain the parameter value of each control parameter, and then further displays the parameter value in the corresponding information display control.

[0077] In the actual application process, the electronic device can specifically call system-level APIs (Application Programming Interfaces), such as AccessibilityService or Intent interfaces, and directly fill the extracted parameter values into the corresponding information display controls.

[0078] In the above embodiments provided by this application, in response to a second input, control parameters of at least one information display control in the target interface are obtained; according to the control parameters, second text information in the target image is extracted to obtain parameter values of each control parameter; and the parameter values are displayed in the corresponding information display controls. In this way, an accurate mapping of the text information in the image to the controls in the function page is realized, the automatic filling of relevant parameters in the function page is realized, the manual operations of the user are reduced, the efficiency and accuracy of information entry are improved, and the usability of the function page is enhanced.

[0079] In the embodiments of this application, the above S108a may specifically include the following S110 and S112:

[0080] S110: In response to a second input, generate a second instruction.

[0081] Among them, the second instruction is used to instruct the target model to identify the control parameters of each information display control in the target interface.

[0082] In the actual application process, the above second instruction may specifically be "what necessary parameters are required for the target interface", and no specific limitation is made here.

[0083] S112: Input the target interface and the second instruction into the target model, and obtain the control parameters output by the target model.

[0084] Specifically, in the information extraction method provided in the embodiments of this application, after generating the second instruction, as Figure 7 shown, input the target interface 210 and the second instruction into the target model, so that the target model identifies the target interface 210 according to the instruction of the second instruction, determines the control parameters corresponding to the information display controls in the target interface 210, and outputs the control parameters.

[0085] In the above embodiments provided by this application, in response to a second input, a second instruction is generated, and the second instruction is used to instruct the target model to identify the control parameters of each information display control in the target interface; the target interface and the second instruction are input into the target model, and the control parameters output by the target model are obtained. In this way, the automatic identification and parsing of relevant parameters in the function page are realized, there is no need for manual pre-definition and configuration, the efficiency and accuracy of information filling are improved, and different function page structures can be adapted, improving the compatibility and adaptability to diverse pages.

[0086] In the embodiments of the present application, the above S108b may specifically include the following S114 and S116:

[0087] S114: Determine a third instruction according to the control parameter.

[0088] Among them, the third instruction is used to instruct the target model to extract second text information from the first text information in the target image as the parameter value of the control parameter. In the actual application process, the above third instruction may specifically be "extract the parameter value of the control parameter", which is not specifically limited herein.

[0089] S116: Input the target image and the third instruction into the target model, and obtain the parameter value output by the target model.

[0090] Specifically, in the information extraction method provided in the embodiments of the present application, after obtaining the control parameter corresponding to the target interface, a third instruction is generated based on the control parameter, and then the target image and the third instruction are input into the target model, so that the target model extracts the second text information in the target image according to the control parameter according to the instruction of the third instruction, to obtain the parameter value of each control parameter, and output the extracted parameter value.

[0091] In the actual application process, the output format of the target model may specifically be "control parameter = parameter value". For example, as Figure 6 shown, "Name = Name1", "Company = Company1", "Position = Position1", "Email = Email1", "Mobile = + area code mobile phone number", etc.

[0092] It can be understood that the text extraction technology in the related art can only extract the original text and cannot directly extract the final result according to the requirements. For example, when copying a landline number, the parentheses outside the area code are not required, but the text extraction solution in the related art will still extract the original text together with the parentheses, and the user needs to perform secondary editing.

[0093] In the information extraction method provided in the embodiments of the present application, by identifying and parsing the control parameters in the target interface based on the multimodal large model, and understanding the content and extracting information from the target image, the extraction result can be accurately output according to the requirements, avoiding secondary editing of the target interface, and avoiding the process of docking the control parameters in the target interface one by one, thus improving the operation efficiency.

[0094] In the above embodiments provided by the present application, a third instruction is determined according to the control parameter, and the third instruction is used to instruct the target model to extract second text information from the first text information in the target image as the parameter value of the control parameter; the target image and the third instruction are input into the target model, and the parameter value output by the target model is obtained. In this way, the automation and accuracy of extracting relevant parameter values in the function page are enhanced, information can be flexibly extracted according to different parameter requirements, and the versatility and reliability of information extraction are improved.

[0095] In summary, as Figure 8 shown, the information extraction method provided by the embodiments of the present application may specifically include the following S302 to S312:

[0096] S302: Detect that the user long-presses the target image, triggering the target model to understand the target image.

[0097] S304: Identify the user intention information corresponding to the target image, and display at least one identifier corresponding to the user intention information.

[0098] S306: Detect that the user drags the target image to the target identifier.

[0099] S308: Obtain the control parameters required when executing the target interface corresponding to the target identifier.

[0100] S310: According to the control parameters, use the target model to extract the parameter values corresponding to each control parameter from the target image.

[0101] S312: Fill the parameter values into the information display control of the target interface correspondingly.

[0102] In the information extraction method provided by the embodiments of the present application, a multimodal large model is used for information extraction. Since the multimodal large model has the ability to input or output multiple types of files, for example, the information extraction method provided by the embodiments of the present application can also be applied to the information extraction of different types of files such as documents, videos, and audios, which is not specifically limited herein.

[0103] Furthermore, in the information extraction method provided by the embodiments of the present application, a single target image is used as the input information of the multimodal large model. In the actual application process, information such as the user's current environment, device status, and device operation history can also be combined as the input information of the multimodal large model to further improve the accuracy of user intention judgment and parameter extraction.

[0104] For the information extraction method provided by the embodiments of the present application, the execution subject may be an information extraction device. In the embodiments of the present application, taking the information extraction device executing the above information extraction method as an example, the information extraction device provided by the embodiments of the present application is described.

[0105] AsFigure 9 As shown in Figure 9 , an information extraction device 500 is provided in an embodiment of the present application. The device may include a receiving unit 502, a display unit 504, and a processing unit 506 described below.

[0106] The receiving unit 502 is configured to receive a first input from a user for a target image;

[0107] The display unit 504 is configured to display at least one identifier associated with the first text information in the target image in response to the first input. The identifier includes at least one of an icon of an application program and an image of a function item in the application program;

[0108] The receiving unit 502 is further configured to receive a second input from the user for moving the target image to a target identifier among the at least one identifier;

[0109] The processing unit 506 is configured to extract the second text information in the target image to a target interface corresponding to the target identifier in response to the second input. The first text information includes the second text information.

[0110] The information extraction device 500 provided in the embodiment of the present application receives a first input from a user for a target image; in response to the first input, displays at least one identifier associated with the first text information in the target image, where the identifier includes at least one of an icon of an application program and an image of a function item in the application program; receives a second input from the user for moving the target image to a target identifier among the at least one identifier; and in response to the second input, extracts the second text information in the target image to a target interface corresponding to the target identifier. The first text information includes the second text information. Through the above information extraction device 500, based on the first input from the user for the target image, the identifiers of the application program and its function items associated with the first text information in the target image can be displayed, and then based on the target identifier selected by the user, the second text information in the target image is extracted to the target interface corresponding to the target identifier. In this way, the intelligent extraction of text information in the image is realized, and the text information in the image can be automatically and intelligently filled into the corresponding interface without manual operation by the user, reducing the operation steps of the user and improving the efficiency of information extraction.

[0111] In the embodiment of the present application, the processing unit 506 is further configured to: in response to the first input, call a target model and generate a first instruction. The target model is a side model, and the first instruction is used to instruct the target model to determine user intention information according to the first text information in the target image; input the target image and the first instruction into the target model to obtain the user intention information output by the target model; and the display unit 504 is specifically configured to: display at least one identifier according to the user intention information.

[0112] In the above embodiments provided by the present application, in response to a first input, a target model is called and a first instruction is generated. The target model is a side model, and the first instruction is used to instruct the target model to determine user intention information according to the first text information in the target image. The target image and the first instruction are input into the target model, and the user intention information output by the target model is obtained. According to the user intention information, at least one identifier is displayed. In this way, by using the side model for image understanding, data can be processed quickly locally, reducing data transmission latency and privacy risks, and being able to more accurately grasp the user's operation intention based on the image text, so that the provided identifier is more in line with the user's needs.

[0113] In the embodiments of the present application, the processing unit 506 is specifically configured to: in response to a second input, obtain the control parameters of at least one information display control in the target interface; according to the control parameters, extract the second text information in the target image to obtain the parameter values of each control parameter; the display unit 504 is further configured to: display the parameter values in the corresponding information display controls.

[0114] In the above embodiments provided by the present application, in response to a second input, obtain the control parameters of at least one information display control in the target interface; according to the control parameters, extract the second text information in the target image to obtain the parameter values of each control parameter; display the parameter values in the corresponding information display controls. In this way, an accurate mapping from the text information in the image to the controls in the function page is realized, the automatic filling of relevant parameters in the function page is realized, the manual operation of the user is reduced, the efficiency and accuracy of information entry are improved, and the usability of the function page is enhanced.

[0115] In the embodiments of the present application, the processing unit 506 is specifically configured to: in response to a second input, generate a second instruction, where the second instruction is used to instruct the target model to identify the control parameters of each information display control in the target interface; input the target interface and the second instruction into the target model, and obtain the control parameters output by the target model.

[0116] In the above embodiments provided by the present application, in response to a second input, generate a second instruction, where the second instruction is used to instruct the target model to identify the control parameters of each information display control in the target interface; input the target interface and the second instruction into the target model, and obtain the control parameters output by the target model. In this way, the automatic recognition and parsing of relevant parameters in the function page are realized, without the need for manual pre-definition and configuration, the efficiency and accuracy of information filling are improved, and different function page structures can be adapted, improving the compatibility and adaptability to diverse pages.

[0117] In an embodiment of the present application, the processing unit 506 is specifically configured to: determine a third instruction according to a control parameter, where the third instruction is used to instruct a target model to extract second text information from first text information in a target image as a parameter value of the control parameter; input the target image and the third instruction into the target model, and obtain the parameter value output by the target model.

[0118] In the above embodiment provided by the present application, a third instruction is determined according to a control parameter, where the third instruction is used to instruct a target model to extract second text information from first text information in a target image as a parameter value of the control parameter; the target image and the third instruction are input into the target model, and the parameter value output by the target model is obtained. In this way, the automation and accuracy of extracting relevant parameter values in a function page are enhanced, information can be flexibly extracted according to different parameter requirements, and the generality and reliability of information extraction are improved.

[0119] The information extraction device 500 in an embodiment of the present application may be an electronic device or a component in an electronic device, such as an integrated circuit or a chip. The electronic device may be a terminal or other devices other than the terminal. Exemplarily, the electronic device may be a mobile phone, a tablet computer, a laptop computer, a handheld computer, a vehicle-mounted electronic device, a Mobile Internet Device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc., and may also be a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc. The embodiments of the present application do not make specific limitations.

[0120] The information extraction device 500 in an embodiment of the present application may be a device with an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems. The embodiments of the present application do not make specific limitations.

[0121] The information extraction device 500 provided in the embodiment of the present application can implement Figure 1 and Figure 8 each process implemented by the method embodiment. To avoid repetition, it will not be elaborated here.

[0122] Optionally, asFigure 10 As shown in the figure, an embodiment of the present application further provides an electronic device 600, including a processor 602 and a memory 604. A program or instruction that can run on the processor 602 is stored on the memory 604. When the program or instruction is executed by the processor 602, each step of the information extraction method embodiment described above is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be elaborated here.

[0123] It should be noted that the electronic devices in the embodiments of the present application include the above-mentioned mobile electronic devices and non-mobile electronic devices.

[0124] Figure 11 It is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present application.

[0125] The electronic device 700 includes but is not limited to: a radio frequency unit 701, a network module 702, an audio output unit 703, an input unit 704, a sensor 705, a display unit 706, a user input unit 707, an interface unit 708, a memory 709, a processor 710, and other components.

[0126] Those skilled in the art can understand that the electronic device 700 may further include a power source (such as a battery) for supplying power to each component. The power source can be logically connected to the processor 710 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. Figure 11 The structure of the electronic device shown in the figure does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0127] Among them, the user input unit 707 is used to receive a first input from the user for a target image.

[0128] The display unit 706 is used to display at least one identifier associated with the first text information in the target image in response to the first input. The identifier includes at least one of an icon of an application program and an image of a function item in the application program.

[0129] The user input unit 707 is further used to receive a second input from the user for moving the target image to a target identifier among at least one identifier.

[0130] The processor 710 is used to extract the second text information in the target image to a target interface corresponding to the target identifier in response to the second input. The first text information includes the second text information.

[0131] In an embodiment of the present application, a first input of a user for a target image is received; in response to the first input, at least one identifier associated with first text information in the target image is displayed, and the identifier includes at least one of an icon of an application program and an image of a function item in the application program; a second input of the user moving the target image to a target identifier among the at least one identifier is received; in response to the second input, second text information in the target image is extracted into a target interface corresponding to the target identifier, and the first text information includes the second text information. In the embodiment of the present application, based on the first input of the user for the target image, the identifiers of the application program and its function items associated with the first text information in the target image can be displayed, and then based on the target identifier selected by the user, the second text information in the target image is extracted into the target interface corresponding to the target identifier. In this way, the intelligent extraction of text information in the image is realized, and the text information in the image can be automatically and intelligently filled into the corresponding interface without manual operation by the user, reducing the operation steps of the user and improving the efficiency of information extraction.

[0132] Optionally, the processor 710 is further configured to: in response to the first input, call a target model and generate a first instruction, where the target model is a side model, and the first instruction is used to instruct the target model to determine user intention information according to the first text information in the target image; input the target image and the first instruction into the target model to obtain the user intention information output by the target model; specifically, the display unit 706 is configured to: display at least one identifier according to the user intention information.

[0133] In the above embodiment provided by the present application, in response to the first input, a target model is called and a first instruction is generated, where the target model is a side model, and the first instruction is used to instruct the target model to determine user intention information according to the first text information in the target image; the target image and the first instruction are input into the target model to obtain the user intention information output by the target model; at least one identifier is displayed according to the user intention information. In this way, by using the side model for image understanding, data can be quickly processed locally, reducing data transmission delay and privacy risks, and being able to more accurately grasp the operation intention of the user based on the image text, so that the provided identifiers are more in line with the user's needs.

[0134] Optionally, the processor 710 is specifically configured to: in response to the second input, obtain control parameters of at least one information display control in the target interface; according to the control parameters, extract the second text information in the target image to obtain parameter values of each control parameter; the display unit 706 is further configured to: display the parameter values in the corresponding information display controls.

[0135] In the above embodiments provided by the present application, in response to a second input, control parameters of at least one information display control in a target interface are obtained; according to the control parameters, second text information in a target image is extracted to obtain parameter values of each control parameter; and the parameter values are displayed in the corresponding information display controls. In this way, an accurate mapping of the text information in the image to the controls in the function page is achieved, automatic filling of relevant parameters in the function page is realized, manual operations of the user are reduced, the efficiency and accuracy of information entry are improved, and the usability of the function page is enhanced.

[0136] Optionally, the processor 710 is specifically configured to: in response to a second input, generate a second instruction for instructing a target model to identify control parameters of each information display control in the target interface; input the target interface and the second instruction into the target model, and obtain the control parameters output by the target model.

[0137] In the above embodiments provided by the present application, in response to a second input, a second instruction is generated for instructing a target model to identify control parameters of each information display control in the target interface; the target interface and the second instruction are input into the target model, and the control parameters output by the target model are obtained. In this way, automatic identification and parsing of relevant parameters in the function page are achieved, without manual pre-definition and configuration, the efficiency and accuracy of information filling are improved, and different function page structures can be adapted, improving the compatibility and adaptability to diverse pages.

[0138] Optionally, the processor 710 is specifically configured to: determine a third instruction according to the control parameters, where the third instruction is used to instruct the target model to extract second text information from the first text information in the target image as the parameter value of the control parameter; input the target image and the third instruction into the target model, and obtain the parameter value output by the target model.

[0139] In the above embodiments provided by the present application, a third instruction is determined according to the control parameters, where the third instruction is used to instruct the target model to extract second text information from the first text information in the target image as the parameter value of the control parameter; the target image and the third instruction are input into the target model, and the parameter value output by the target model is obtained. In this way, the automation and accuracy of extracting relevant parameter values in the function page are enhanced, information can be flexibly extracted according to different parameter requirements, and the generality and reliability of information extraction are improved.

[0140] It should be understood that in the embodiments of the present application, the input unit 704 may include a Graphics Processing Unit (GPU) 7041 and a microphone 7042. The GPU 7041 processes the image data of static pictures or videos obtained by an image capture device (such as a camera) in the video capture mode or the image capture mode. The display unit 706 may include a display panel 7061, and the display panel 7061 may be configured in the form of a liquid crystal display, an organic light emitting diode, etc. The user input unit 707 includes at least one of a touch panel 7071 and other input devices 7072. The touch panel 7071 is also referred to as a touch screen. The touch panel 7071 may include two parts: a touch detection device and a touch controller. The other input devices 7072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, power on / off keys, etc.), a trackball, a mouse, a joystick, which will not be elaborated here.

[0141] The memory 709 can be used to store software programs and various data. The memory 709 mainly includes a first storage area for storing programs or instructions and a second storage area for storing data. Among them, the first storage area can store an operating system, applications or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory 709 may include a volatile memory or a non-volatile memory, or the memory 709 may include both a volatile memory and a non-volatile memory. Among them, the non-volatile memory may be a Read-Only Memory (ROM), a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically Erasable PROM (EEPROM), or a flash memory. The volatile memory may be a Random Access Memory (RAM), a Static RAM (SRAM), a Dynamic RAM (DRAM), a Synchronous DRAM (SDRAM), a Double Data Rate SDRAM (DDR SDRAM), an Enhanced SDRAM (ESDRAM), a Synch link DRAM (SLDRAM), and a Direct Rambus RAM (DRRAM). The memory 709 in the embodiments of the present application includes, but is not limited to, these and any other suitable types of memories.

[0142] The processor 710 may include one or more processing units; optionally, the processor 710 integrates an application processor and a modem processor. Among them, the application processor mainly processes operations related to the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication signals, such as a baseband processor. It can be understood that the above-mentioned modem processor may not be integrated into the processor 710 either.

[0143] The embodiments of the present application further provide a readable storage medium. Programs or instructions are stored on the readable storage medium. When the programs or instructions are executed by a processor, each process of the above-mentioned information extraction method embodiment is implemented, and the same technical effects can be achieved. To avoid repetition, details are not described here again.

[0144] Among them, the processor is the processor in the electronic device in the above-mentioned embodiment. The readable storage medium includes computer-readable storage media, such as computer read-only memory ROM, random access memory RAM, magnetic disks, or optical discs, etc.

[0145] The embodiments of the present application further provide a chip. The chip includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement each process of the above-mentioned information extraction method embodiment, and the same technical effects can be achieved. To avoid repetition, details are not described here again.

[0146] It should be understood that the chip mentioned in the embodiments of the present application may also be referred to as a system-on-chip, system chip, chip system, or system-on-chip, etc.

[0147] The embodiments of the present application provide a computer program product. The program product is stored in a storage medium. The program product is executed by at least one processor to implement each process of the above-mentioned information extraction method embodiment, and the same technical effects can be achieved. To avoid repetition, details are not described here again.

[0148] It should be noted that in this article, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without more limitations, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including such element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0149] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described example methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of the present application.

[0150] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them fall within the protection scope of the present application.

Claims

1. An information extraction method, characterized in that including: Receiving a first input of a user for a target image; In response to the first input, displaying at least one identifier associated with first text information in the target image, the identifier including at least one of an icon of an application program and an icon of a function item in the application program; Receiving a second input of the user moving the target image to a target identifier among the at least one identifier; In response to the second input, extracting second text information in the target image into a target interface corresponding to the target identifier, where the first text information includes the second text information.

2. The information extraction method according to claim 1, wherein The responding to the first input and displaying at least one identifier associated with the first text information in the target image includes: In response to the first input, calling a target model and generating a first instruction, where the target model is a side model, and the first instruction is used to instruct the target model to determine user intention information according to the first text information in the target image; Inputting the target image and the first instruction into the target model to obtain user intention information output by the target model; Displaying at least one identifier according to the user intention information.

3. The information extraction method according to claim 2, wherein The responding to the second input and extracting the second text information in the target image into a target interface corresponding to the target identifier includes: In response to the second input, obtaining control parameters of at least one information display control in the target interface; According to the control parameters, extracting the second text information in the target image to obtain parameter values of each control parameter; Displaying the parameter values in corresponding information display controls.

4. The information extraction method according to claim 3, wherein The responding to the second input and obtaining control parameters of at least one information display control in the target interface includes: In response to the second input, generating a second instruction, where the second instruction is used to instruct the target model to identify control parameters of each information display control in the target interface; Inputting the target interface and the second instruction into the target model to obtain control parameters output by the target model.

5. The information extraction method according to claim 3, wherein The extracting the second text information in the target image according to the control parameters to obtain parameter values of each control parameter includes: Determining a third instruction according to the control parameters, where the third instruction is used to instruct the target model to extract second text information from the first text information in the target image as the parameter value of the control parameter; Inputting the target image and the third instruction into the target model to obtain parameter values output by the target model.

6. An information extraction device, characterized in that, including: A receiving unit, configured to receive a first input of a user for a target image; A display unit, configured to, in response to the first input, display at least one identifier associated with first text information in the target image, the identifier including at least one of an icon of an application program and an icon of a function item in the application program; The receiving unit is further configured to receive a second input of the user moving the target image to a target identifier among the at least one identifier; A processing unit, configured to extract second text information in the target image to a target interface corresponding to the target identifier in response to the second input, where the first text information includes the second text information.

7. The information extraction device according to claim 6, wherein The processing unit is further configured to: Call a target model and generate a first instruction in response to the first input, where the target model is an edge-side model, and the first instruction is used to instruct the target model to determine user intention information according to the first text information in the target image; Input the target image and the first instruction into the target model, and obtain the user intention information output by the target model; The display unit is specifically configured to: Display at least one identifier according to the user intention information.

8. The information extraction device according to claim 7, wherein The processing unit is specifically configured to: Obtain control parameters of at least one information display control in the target interface in response to the second input; Extract the second text information in the target image according to the control parameters to obtain parameter values of each control parameter; The display unit is further configured to: Display the parameter values in the corresponding information display controls.

9. The information extraction device according to claim 8, wherein The processing unit is specifically configured to: Generate a second instruction in response to the second input, where the second instruction is used to instruct the target model to identify the control parameters of each information display control in the target interface; Input the target interface and the second instruction into the target model, and obtain the control parameters output by the target model.

10. The information extraction device according to claim 8, wherein The processing unit is specifically configured to: Determine a third instruction according to the control parameters, where the third instruction is used to instruct the target model to extract the second text information from the first text information in the target image as the parameter value of the control parameter; Input the target image and the third instruction into the target model, and obtain the parameter values output by the target model.