Information entry method and apparatus based on large model, storage medium, and computer device

By using a large-scale model-based information entry method, an efficient and accurate information entry process was achieved, solving the problems of low efficiency and poor accuracy in existing technologies, improving user experience and reducing merchant churn on the platform.

WO2026086421A1PCT designated stage Publication Date: 2026-04-30RAJAX NETWORK &TECHNOLOGY (SHANGHAI) CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
RAJAX NETWORK &TECHNOLOGY (SHANGHAI) CO LTD
Filing Date
2025-08-29
Publication Date
2026-04-30

AI Technical Summary

Technical Problem

The existing information entry methods are inefficient, inaccurate, and result in a poor user experience. In particular, when merchants join the online platform, errors, omissions, and incorrect material locations are prone to occur, leading to cumbersome processes and user resistance.

Method used

The system employs a large-scale model-based information entry method, allowing users to input multiple images at once through an image entry page. The large-scale model is then used for image type recognition and text information analysis to automatically populate the application information page. The system also performs integrity checks and forged image identification, and provides modification and confirmation interfaces.

Benefits of technology

It improved the efficiency and accuracy of information entry, shortened the time users spent filling out forms, reduced the risk of errors, omissions, and incorrect material delivery, improved the user experience, reduced merchant churn on the platform, and promoted the standardization and normalization of information entry.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025117856_30042026_PF_FP_ABST
    Figure CN2025117856_30042026_PF_FP_ABST
Patent Text Reader

Abstract

An information entry method and apparatus based on a large model, a storage medium, and a computer device. The method comprises: displaying a picture entry page for a target function, wherein the target function is a function applied for by providing specific types of information, and the specific types of information include picture information of at least one required picture type and text information of at least one required text type (101); acquiring at least one picture entered on the basis of the picture entry page, and performing picture type recognition and text information analysis on the picture by means of a large model (102); populating an application information page for the target function on the basis of the picture, a target picture type recognized by means of the large model and target text information analyzed by means of the large model, wherein the application information page is an editable page containing the specific types of information (103); and displaying the application information page, and when the application information page is confirmed, submitting entered information for the target function on the basis of the confirmed application information page (104).
Need to check novelty before this filing date? Find Prior Art

Description

Information entry methods, devices, storage media, and computer equipment based on large models

[0001] This application claims priority to Chinese Patent Application No. CN202411498777.5, filed on October 24, 2024, entitled "Information Input Method, Apparatus, Storage Medium and Computer Equipment Based on Large Model", the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the field of information processing technology, and in particular to an information input method, apparatus, storage medium and computer equipment based on a large model. Background Technology

[0003] In related technologies, exemplary information entry methods often rely on users manually inputting or uploading files, followed by manual review to ensure the accuracy and completeness of the information. However, with the rapid development of information technology and the increasing demand from users for efficient and convenient services, these exemplary information entry methods have gradually revealed many shortcomings. The exemplary information entry methods have significant deficiencies in terms of information entry efficiency, accuracy, and user experience, urgently requiring an efficient, accurate, and convenient information entry method to solve these problems. Summary of the Invention

[0004] In view of this, the embodiments of this application provide an information entry method, apparatus, storage medium and computer equipment based on a large model, which helps to improve the efficiency and accuracy of information entry, shorten the user's filling time, reduce the risk of incorrect writing, omissions and wrong material location transmission, and at the same time improve the user experience and reduce user resistance and platform merchant loss caused by cumbersome processes.

[0005] According to one aspect of this application, a method for information entry based on a large model is provided, comprising:

[0006] Displays an image input page for the target function, wherein the target function is a function applied for by providing specific type of information, and the specific type of information includes image information of at least one required image type and text information of at least one required text type;

[0007] Obtain at least one image entered based on the image entry page, and perform image type recognition and text information analysis on the image using a large model;

[0008] Based on the image, the target image type identified by the large model, and the analyzed target text information, the application information page for the target function is populated, wherein the application information page is a modifiable page containing the specific type of information;

[0009] The application information page is displayed, and if the application information page is confirmed, the input information for the target function is submitted based on the confirmed application information page.

[0010] In some exemplary embodiments, before performing image type recognition on the image using a large model, the method further includes:

[0011] Perform text recognition on each image to obtain the text information for each image; and,

[0012] For any image, identify whether the image text information matches any keyword corresponding to any required image type, and determine the target image type of the image based on the required image type corresponding to the matched keyword;

[0013] Accordingly, image type identification is performed on the image using a large model, including:

[0014] Based on the required image type and the identified target image type, determine the second image type to be identified; and,

[0015] For each image whose image type is not determined, the feature description information corresponding to the second image type is obtained. The image type identification prompt template is filled according to the image, the second image type, and the feature description information to obtain the image type identification prompt information. The target image type of the image is identified at least in the second image type using a large model based on the image type identification prompt information.

[0016] In some exemplary embodiments, before performing text information analysis on the image using a large model, the method further includes:

[0017] Based on the required image type, perform image integrity verification on the target image type;

[0018] If the image integrity verification passes, a first image confirmation page is displayed based on each image and its target image type. In response to a confirmation operation on the first image confirmation page, the image type of each confirmed image is used as its target image type.

[0019] If the image integrity check fails, a second image confirmation page is displayed based on each image, the target image type of each image, and the missing image type. The second image confirmation page includes an image upload control corresponding to the missing image type. If an image of the missing image type is uploaded based on the image upload control, in response to the confirmation operation on the second image confirmation page, the image type of each confirmed image is used as the target image type of each image.

[0020] In some exemplary embodiments, text information analysis of the image is performed using a large model, including:

[0021] For each image, text extraction is performed on the corresponding image text information based on the preset text type; and,

[0022] For each image, based on the image, the text extraction results of the image, and the preset text type corresponding to the image, text information analysis prompts are constructed, and a large model is used to analyze text information that matches the preset text type based on the text information analysis prompts.

[0023] Specifically, based on the image, the text extraction results of the image, and the preset text type corresponding to the image, a text information analysis prompt is constructed, including:

[0024] For each image from which text information has been extracted, a text information confirmation prompt template is populated based on the image, the text extraction result of the image, and the preset text type corresponding to the image, to obtain the text information analysis prompt information. The text information confirmation prompt template includes content configured to guide the large model to verify the text extraction result of the image based on the image and the preset text type corresponding to the image; and...

[0025] For each image from which no text information has been extracted, a text information generation prompt template is filled in according to the image and the preset text type corresponding to the image to obtain the text information analysis prompt information. The text information generation prompt template includes content configured to guide the large model to generate text information based on the image and the preset text type corresponding to the image.

[0026] In some exemplary embodiments, before populating the application information page for the target function based on the image, the target image type identified by the large model, and the analyzed target text information, the method further includes:

[0027] Based on the required image type, perform image integrity verification on the target image type; and based on the required text type, perform text integrity verification on the target text information.

[0028] If the image integrity check and the text integrity check pass, then the step of filling the application information page for the target function based on the image, the target image type of the image, and the target text information is executed; and,

[0029] If at least one of the image integrity checks and the text integrity checks fails, a missing item that failed the integrity check is identified. If the missing item is a predicted item, based on the target text information and the image text information corresponding to the image, the associated item information corresponding to the missing item is obtained. Based on the associated item information and the missing item, a missing item prediction prompt is constructed. The missing item is predicted using a large model based on the missing item prediction prompt. The application information page for the target function is filled based on the predicted missing item information, the image, the target image type of the image, and the target text information. If the missing item is not a predicted item, the application information page for the target function is filled directly based on the image, the target image type of the image, and the target text information.

[0030] In some exemplary embodiments, the target function includes opening an online store; the required image types include storefront images, store environment images, business license images, and operating permit images; the required text types include store name, store address, product category, and business hours; and the prediction items include product category and business hours.

[0031] When the missing item is a predicted item, based on the target text information and the image text information corresponding to the image, the associated item information corresponding to the missing item is obtained, including:

[0032] If the missing item includes a product category, the associated item information corresponding to the product category is determined based on at least one of the following: the store name, the image and text information corresponding to the storefront image, the image and text information corresponding to the business license image, the image and text information corresponding to the operating permit image, the menu image information, and the relevant store product category information corresponding to the store name; wherein, the relevant store includes chain stores, and the menu image information includes image and text information identified in images with a target image type of menu and / or images with a target image type of dish; and,

[0033] If the missing item includes business hours, determine the range of surrounding stores corresponding to the store address, obtain information on each surrounding store within the range of surrounding stores, the surrounding store information includes the business categories and business hours of the surrounding stores, and determine the associated item information corresponding to the business hours based on the surrounding store information and the business categories.

[0034] In some exemplary embodiments, the method further includes:

[0035] The image is subjected to watermark recognition and / or similar image recognition based on a preset image library to determine whether the image is a forged image, wherein the forged images include watermarked images and images that meet similarity criteria; and,

[0036] If the image is a forged image, a reminder page will be displayed to prompt the user to correct the forged image.

[0037] The process of identifying similar images based on a preset image library includes:

[0038] For any image, based on the target image type of the image, at least one comparison image is determined from each preset image in the preset image library, and the similarity between the image and each comparison image is calculated according to the comparison conditions corresponding to the target image type of the image. The comparison conditions for text images include comparing text similarity, and the comparison conditions for image images include comparing pixel grayscale value similarity.

[0039] In some exemplary embodiments, after displaying the application information page, the method further includes:

[0040] Obtain information modification data for the application information page, and modify specific types of information on the application information page based on the information modification data; and,

[0041] The modified information data is recorded, and large model training samples are constructed based on the modified information data to optimize the training of the large model.

[0042] According to another aspect of this application, an information input device based on a large model is provided, comprising:

[0043] The display section is configured to display an image input page for a target function, wherein the target function is a function requested by providing specific type information, and the specific type information includes image information of at least one required image type and text information of at least one required text type;

[0044] The analysis section is configured to acquire at least one image entered based on the image entry page, and perform image type recognition and text information analysis on the image using a large model;

[0045] The filling part is configured to fill the application information page of the target function based on the image, the target image type identified by the large model, and the analyzed target text information, wherein the application information page is a modifiable page containing the specific type of information;

[0046] The display section is further configured to display the application information page, and, if the application information page is confirmed, to submit the input information for the target function based on the confirmed application information page.

[0047] In some exemplary embodiments, the analysis section is further configured to:

[0048] Perform text recognition on each image to obtain the text information for each image; and,

[0049] For any image, identify whether the image text information matches any keyword corresponding to any required image type, and determine the target image type of the image based on the required image type corresponding to the matched keyword;

[0050] Based on the required image type and the identified target image type, determine the second image type to be identified; and,

[0051] For each image whose image type is not determined, the feature description information corresponding to the second image type is obtained. The image type identification prompt template is filled according to the image, the second image type, and the feature description information to obtain the image type identification prompt information. The target image type of the image is identified at least in the second image type using a large model based on the image type identification prompt information.

[0052] In some exemplary embodiments, the analysis section is further configured to: perform image integrity verification on the target image type based on the required image type;

[0053] The display section is further configured to: if the image integrity verification passes, display a first image confirmation page based on each image and its target image type, and in response to a confirmation operation on the first image confirmation page, use the image type of each confirmed image as its target image type; and,

[0054] If the image integrity check fails, a second image confirmation page is displayed based on each image, the target image type of each image, and the missing image type. The second image confirmation page includes an image upload control corresponding to the missing image type. If an image of the missing image type is uploaded based on the image upload control, in response to the confirmation operation on the second image confirmation page, the image type of each confirmed image is used as the target image type of each image.

[0055] In some exemplary embodiments, the analysis section is further configured to:

[0056] For each image, text extraction is performed on the image text information corresponding to the image according to the preset text type corresponding to the image;

[0057] For each image, based on the image, the text extraction results of the image, and the preset text type corresponding to the image, text information analysis prompts are constructed, and a large model is used to analyze text information that matches the preset text type based on the text information analysis prompts.

[0058] For each image from which text information has been extracted, a text information confirmation prompt template is populated based on the image, the text extraction result of the image, and the preset text type corresponding to the image, to obtain the text information analysis prompt information. The text information confirmation prompt template includes content configured to guide the large model to verify the text extraction result of the image based on the image and the preset text type corresponding to the image; and...

[0059] For each image from which no text information has been extracted, a text information generation prompt template is filled in according to the image and the preset text type corresponding to the image to obtain the text information analysis prompt information. The text information generation prompt template includes content configured to guide the large model to generate text information based on the image and the preset text type corresponding to the image.

[0060] In some exemplary embodiments, the analysis section is further configured to:

[0061] Based on the required image type, perform image integrity verification on the target image type; and based on the required text type, perform text integrity verification on the target text information.

[0062] The filling portion is further configured to: if the image integrity check and the text integrity check pass, then execute the step of filling the application information page of the target function based on the image, the target image type of the image, and the target text information; and,

[0063] If at least one of the image integrity checks and the text integrity checks fails, a missing item that failed the integrity check is identified. If the missing item is a predicted item, based on the target text information and the image text information corresponding to the image, the associated item information corresponding to the missing item is obtained. Based on the associated item information and the missing item, a missing item prediction prompt is constructed. The missing item is predicted using a large model based on the missing item prediction prompt. The application information page for the target function is filled based on the predicted missing item information, the image, the target image type of the image, and the target text information. If the missing item is not a predicted item, the application information page for the target function is filled directly based on the image, the target image type of the image, and the target text information.

[0064] In some exemplary embodiments, the target function includes opening an online store; the required image types include storefront images, store environment images, business license images, and operating permit images; the required text types include store name, store address, product category, and business hours; and the prediction items include product category and business hours.

[0065] If the missing item is a predicted item, the analysis section is further configured as follows:

[0066] If the missing item includes a product category, the associated item information corresponding to the product category is determined based on at least one of the following: the store name, the image and text information corresponding to the storefront image, the image and text information corresponding to the business license image, the image and text information corresponding to the operating permit image, the menu image information, and the relevant store product category information corresponding to the store name; wherein, the relevant store includes chain stores, and the menu image information includes image and text information identified in images with a target image type of menu and / or images with a target image type of dish; and,

[0067] If the missing item includes business hours, determine the range of surrounding stores corresponding to the store address, obtain information on each surrounding store within the range of surrounding stores, the surrounding store information includes the business categories and business hours of the surrounding stores, and determine the associated item information corresponding to the business hours based on the surrounding store information and the business categories.

[0068] In some exemplary embodiments, the analysis section is further configured to:

[0069] The image is subjected to watermark recognition and / or similar image recognition based on a preset image library to determine whether the image is a forged image, wherein the forged image includes watermarked images and images that meet the similarity criteria;

[0070] If the image is a forged image, a reminder page will be displayed to prompt correction of the forged image; and,

[0071] For any image, based on the target image type of the image, at least one comparison image is determined from each preset image in the preset image library, and the similarity between the image and each comparison image is calculated according to the comparison conditions corresponding to the target image type of the image. The comparison conditions for text images include comparing text similarity, and the comparison conditions for image images include comparing pixel grayscale value similarity.

[0072] In some exemplary embodiments, the analysis section is further configured to:

[0073] Obtain information modification data for the application information page, and modify specific types of information on the application information page based on the information modification data; and,

[0074] The modified information data is recorded, and large model training samples are constructed based on the modified information data to optimize the training of the large model.

[0075] According to another aspect of this application, a storage medium is provided that stores a computer program thereon, which, when executed by a processor, implements the above-described information entry method based on a large model.

[0076] According to another aspect of this application, a computer device is provided, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein the processor executes the program to implement the above-described information input method based on a large model.

[0077] By employing the above technical solutions, this application provides a method, apparatus, storage medium, and computer device for information entry based on a large model. When entering information for a target function, multiple images can be entered at once on the image entry page. The large model is used to identify the image type and analyze the text information of the entered images. Based on the entered images, the identified target image type, and the analyzed target text information, an application information page is displayed, allowing users to modify and confirm the information on the page. The information entry for the target function is completed based on the confirmed information on the application information page. This application, by introducing an information entry method based on a large model, helps improve the efficiency and accuracy of information entry, shortens user filling time, reduces the risk of errors, omissions, and incorrect material transmission, and improves user experience, reducing user resistance and platform merchant churn caused by cumbersome processes. Furthermore, this method promotes the standardization and normalization of information entry, laying a solid foundation for efficient information processing and utilization, thereby optimizing the overall information entry process and meeting users' needs for efficient and convenient services.

[0078] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description

[0079] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application or in the background art will be described below.

[0080] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments of the application and are configured to explain the application, but do not constitute an undue limitation of the application. In the drawings:

[0081] Figure 1 shows a flowchart of an information entry method based on a large model provided in an embodiment of this application;

[0082] Figure 2 shows a flowchart of another information entry method based on a large model provided in an embodiment of this application;

[0083] Figure 3 shows a schematic diagram of the structure of an information input device based on a large model provided in an embodiment of this application. Detailed Implementation

[0084] The present application will be described in detail below with reference to the accompanying drawings and embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in the embodiments of the present application can be combined with each other.

[0085] In the description of this application, it should be understood that the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. Those skilled in the art can understand the specific meaning of these terms in this application based on the specific circumstances. Furthermore, in the description of this application, unless otherwise stated, "multiple" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.

[0086] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of this application. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0087] In related technologies, exemplary information entry methods often rely on users manually inputting or uploading files, and then undergoing manual review to ensure the accuracy and completeness of the information. However, with the rapid development of information technology and the increasing demand from users for efficient and convenient services, these exemplary information entry methods have gradually revealed many shortcomings.

[0088] Currently, there are many scenarios in the industry where merchants need to fill out qualification information. For example, merchants joining online platforms need to fill out forms containing numerous supporting documents, and they also need to simultaneously fill in the corresponding text information according to the content of the documents. This easily leads to errors, omissions, and incorrect transmission of relevant documents. For instance, an image box that should be uploaded as a food license might be uploaded as a business license, or the store name might be entered incorrectly, with missing or extra characters. The forms also have too many fields, making it time-consuming for merchants to complete them. Therefore, the success rate of filling out forms on the first attempt is low, requiring repeated modifications and submissions. This can further lead to users having a poor experience with filling out large forms, being unwilling to fill them out, or being unwilling to resubmit after a rejection, resulting in merchant churn on the platform.

[0089] In summary, the existing information entry methods have significant shortcomings in terms of efficiency, accuracy, and user experience, and there is an urgent need for an efficient, accurate, and convenient information entry method to solve these problems.

[0090] This embodiment provides an information entry method based on a large model. Figure 1 shows a flowchart of an information entry method based on a large model provided in this embodiment. As shown in Figure 1, the method includes:

[0091] Step 101: Display the image input page for the target function, wherein the target function is a function applied for by providing specific type information, and the specific type information includes image information of at least one required image type and text information of at least one required text type.

[0092] This application provides an information entry method based on a large model, aiming to simplify and optimize the process for users to apply for specific target functions by providing specific types of information, such as applying for the online store function by filling in various posture information. First, under the application process for the target function, a dedicated image entry page designed for the target function is displayed. This page serves as the entry point for users to submit the required application information. The target function refers to the function that users wish to apply for by providing certain specific types of information. These specific types of information include at least one type of image information and at least one type of text information. Taking the online store function as an example, the required image types may include storefront images, store environment images, business license images, operating permit images, front image of ID card (portrait side), and back image of ID card (reverse side of the portrait side). The required text types may include store name, store address, business license number, product category, business hours, etc. In some implementations, the image entry page may include an image upload box, allowing users to drag and drop images to upload multiple images at once or upload multiple images one by one. Additionally, the image upload box may display prompts to indicate the required image type to the user.

[0093] Step 102: Obtain at least one image entered based on the image entry page, and perform image type recognition and text information analysis on the image using a large model.

[0094] In this embodiment, after a user submits at least one image through the image entry page, a large-scale model can be used to perform image type recognition and text information analysis on the submitted image to identify the image type and extract the required information. The large-scale model's image type recognition includes at least the recognition of the aforementioned required image types, and the text information analysis includes at least the analysis of the aforementioned required text types. By using a large-scale model to achieve image type recognition and text information analysis, the application information page is automatically populated, reducing the possibility of human error, improving the accuracy of information, and helping to reduce the time users spend manually inputting data and the workload of manual review, thereby improving the efficiency of information entry.

[0095] Step 103: Based on the image, the target image type identified by the large model, and the analyzed target text information, populate the application information page of the target function, wherein the application information page is a modifiable page containing the specific type of information.

[0096] Step 104: Display the application information page, and if the application information page is confirmed, submit the input information for the target function based on the confirmed application information page.

[0097] In this embodiment, based on the user-submitted images and the image types and text information identified by the large model, the application information page for the target function can be automatically populated and displayed. The application information page is an editable page containing specific types of information, allowing users to manually modify or supplement it. If the user believes that the image types and text information automatically identified by the large model deviate from or are missing from their expectations, they can modify or supplement the information on the application information page. If the user has no objection to the information on the page (or the modified information), after confirmation, the system will submit the entry information for the target function based on the confirmed application information page, completing the entire application process. By simplifying the information entry process, reducing user operation steps, and providing clear reminder pages and guidance, users can complete information entry more easily and quickly. Furthermore, the automated processing reduces user waiting time, improves user satisfaction and trust, and reduces the difficulty and resistance of filling out forms, increasing user willingness and satisfaction, thereby helping to reduce the churn of platform merchants.

[0098] By applying the technical solution of this embodiment, when entering information for a target function, multiple images can be entered at once on the image entry page. A large model is used to identify the image type and analyze the text information of the entered images. Based on the entered images, the identified target image type, and the analyzed target text information, the application information page is displayed, allowing users to modify and confirm the information on the page. The information entry for the target function is then completed based on the confirmed information on the application information page. This embodiment, by introducing a large model-based information entry method, helps improve the efficiency and accuracy of information entry, shortens user filling time, reduces the risk of errors, omissions, and incorrect material transmission, and improves user experience, reducing user resistance and merchant churn caused by cumbersome processes. Furthermore, this method promotes the standardization and normalization of information entry, laying a solid foundation for efficient information processing and utilization, thereby optimizing the overall information entry process and meeting users' needs for efficient and convenient services.

[0099] In some embodiments of this application, the method further includes: performing watermark recognition and / or similar image recognition based on a preset image library to determine whether the image is a forged image, wherein the forged image includes watermarked images and images that meet similarity conditions; if the image is a forged image, a reminder page is displayed based on the forged image to remind the user to correct the forged image.

[0100] In the above embodiments, in addition to improving the efficiency and accuracy of information entry, this application embodiment further enhances the security and authenticity of information by identifying forged images. Forged image identification includes watermark recognition and similar image recognition based on a preset image library. This mechanism effectively identifies and intercepts forged images, including those with watermarks or those highly similar to other images, thereby avoiding information entry errors or fraudulent activities caused by the use of false materials. When the system detects a forged image, it can display a reminder page to guide the user to correct or re-upload the image. This not only improves the accuracy of the information but also enhances the user's trust in the system, further ensuring the security and authenticity of the information and providing users with a more reliable and efficient information entry experience.

[0101] In some embodiments of this application, the similarity recognition of the image based on a preset image library includes: for any image, determining at least one comparison image in each preset image of the preset image library based on the target image type of the image, and calculating the similarity between the image and each comparison image according to the comparison conditions corresponding to the target image type of the image. The comparison conditions for text images include comparing text similarity, and the comparison conditions for image images include comparing pixel grayscale value similarity.

[0102] In the above embodiments, different similarity image recognition rules can be adopted for different types of images. In some implementations, for any image to be identified, at least one relevant comparison image can be selected from a preset image library based on its target image type (such as business license, ID card, shop environment image, etc.). The preset image library can be a database containing multiple types of images, for example, containing multiple business license images, multiple ID card images (used with user permission), etc. Then, according to the target image type of the image, the corresponding comparison conditions are selected to calculate the similarity between the image and the comparison image. For text-type images, such as business licenses and permits, the comparison conditions mainly focus on the similarity of text content. Text recognition technology and natural language processing technology can be used to extract the text content and determine whether the text content is consistent in order to determine whether the uploaded image is a misappropriation of other people's information. For image-type images, such as photos and screenshots, the comparison conditions focus more on the similarity of pixel grayscale values. Image processing technology is used to convert the image into a grayscale image and calculate the similarity score between grayscale values. This mechanism accurately identifies images highly similar to those in a pre-set image library, effectively blocking forged or tampered images and ensuring the authenticity and accuracy of information entry. When a highly similar image is detected, a notification page is triggered, guiding the user to correct or re-upload the image, preventing the use of forged images for information entry. The similar image recognition function in this embodiment provides additional security for the information entry process through precise comparison conditions and efficient calculation methods, further enhancing system reliability and user experience.

[0103] Furthermore, as a refinement and extension of the specific implementation of the above embodiments, in order to fully illustrate the specific implementation process of this embodiment, another information entry method based on a large model is provided. Figure 2 shows a flowchart of another information entry method based on a large model provided in this application embodiment. As shown in Figure 2, the method includes:

[0104] Step 201: Display the image input page for the target function, wherein the target function is a function applied for by providing specific type information, and the specific type information includes image information of at least one required image type and text information of at least one required text type.

[0105] In this embodiment, an image entry page for a specific target function is first displayed. This target function is requested by providing specific types of information, including image information of at least one required image type and text information of at least one required text type. This step provides the user with a clear and intuitive entry interface, guiding the user to upload relevant images and fill in the necessary text information as required.

[0106] Step 202: Obtain at least one image entered based on the image input page; perform text recognition on each image to obtain the image text information of each image; for any image, identify whether the image text information of the image matches any keyword corresponding to any required image type, and determine the target image type of the image based on the required image type corresponding to the matched keyword; for images whose image text information does not match the keyword, perform image type recognition on the image using a large model to determine the target image type of the image.

[0107] In this embodiment, at least one image entered by the user on the image input page is acquired. Then, text recognition processing is performed on these images to extract the text information, i.e., image text information. This step utilizes advanced OCR (Optical Character Recognition) technology to automatically recognize and extract text from images, providing a foundation for subsequent image classification and information input. Next, for any given image, the following operations are performed: First, it is determined whether the image text information contains keywords corresponding to any desired image type. If a keyword is matched, the target image type is determined based on the desired image type corresponding to that keyword. This step uses keyword matching technology to achieve rapid and accurate determination of image type. For example, images containing keywords such as "resident ID card" and "validity period" can be identified as images of the back of an ID card; images containing keywords such as "food business license" can be identified as images of a business license. For images that are not classified in this way or whose text information does not match the keywords, image type identification is performed using a large model. By using large-scale models to analyze and understand the content of images, the target image type can be determined. This step further improves the accuracy and adaptability of image classification, especially for image types with limited text information or those difficult to identify through keyword matching. This embodiment not only enables rapid and accurate input of image information but also improves the efficiency and accuracy of information entry through intelligent image classification technology. Simultaneously, this process provides a reliable data foundation for subsequent information processing and review, contributing to improved performance and user experience of the entire information entry system.

[0108] In some embodiments of this application, step 202, identifying whether the image text information of the image matches any keyword corresponding to a required image type, includes: identifying whether the image text information of the image matches a keyword corresponding to a first image type, wherein the first image type includes the standard image type among the required image types; correspondingly, performing image type recognition on the image using a large model to determine the target image type of the image includes: determining a second image type to be recognized based on the required image type and the identified target image type; for each image whose image type is not determined, constructing image type recognition prompt information based on the image and the second image type, and using the large model to identify the target image type of the image at least in the second image type based on the image type recognition prompt information. The large model is pre-trained with images of different required image types.

[0109] In the above embodiments, when identifying whether the text information of an image matches a keyword, the primary focus is on the first image type, namely the standard image type. Standard image types typically refer to standard images with fixed formats, content, and requirements, such as business licenses and ID cards. Using a predefined keyword library, the text information of these images is precisely matched to quickly determine their type. This keyword matching method is particularly effective for standard image types, enabling rapid and accurate identification. However, for non-standard image types or images where the text information is insufficient for keyword matching, a more flexible and intelligent identification method is employed. In some implementations, based on the required image type and the identified target image type, the remaining unidentified second image types are determined. These may include non-standard, more complex, or diverse images, such as product photos and scene screenshots. Next, for images whose types have not yet been determined, corresponding image type identification prompts can be constructed to help the large model more accurately understand the image content and identify the target image type within the second image type. By utilizing the deep learning and image understanding capabilities of the large model, it is helpful to accurately locate the type matching the image content, thereby achieving intelligent identification of image types. The information entry process in this embodiment combines keyword matching and large-scale image type recognition, which not only improves the accuracy and efficiency of image type recognition but also enhances the system's flexibility and adaptability.

[0110] In some embodiments of this application, constructing image type recognition prompt information based on the image and the second image type includes: obtaining feature description information corresponding to the second image type; and filling the image type recognition prompt template according to the image, the second image type, and the feature description information to obtain the image type recognition prompt information.

[0111] In the above embodiments, when constructing image type recognition prompts, the feature description information corresponding to the second image type is first obtained. This feature description information may include visual features such as key elements, colors, textures, and shapes that should be present in the image, as well as text features such as text, symbols, or patterns that may appear in the image. This feature description information is derived in advance based on a deep understanding and analysis of the second image type, aiming to help the large model more accurately identify the image type. Next, the image type recognition prompt template is populated according to the image, the second image type, and the feature description information. The image type recognition prompt template is a predefined framework configured to organize and present information related to image type recognition. By populating this template, the system can generate a comprehensive prompt information that includes the image itself, its possible types, and the feature description information corresponding to that type. This comprehensive prompt information not only provides direct visual information about the image but also provides additional clues and context for image type recognition through feature description information. This helps the large model to understand the image content more comprehensively and identify the target type of the image more accurately within the second image type. This application embodiment obtains feature description information of the second image type and fills the image type recognition prompt template with this information, which can generate more accurate and useful image type recognition prompt information, further improving the accuracy and efficiency of the large model in image type recognition, thereby optimizing the performance and user experience of the entire information entry process.

[0112] Furthermore, after determining the image type of each image, in some embodiments of this application, the method further includes: performing an image integrity check on the target image type based on the required image type; if the image integrity check passes, then displaying a first image confirmation page based on each image and its target image type, and in response to a confirmation operation on the first image confirmation page, using the image type of each confirmed image as the target image type of each image; if the image integrity check fails, then displaying a second image confirmation page based on each image, its target image type, and the missing image type, wherein the second image confirmation page includes an image upload control corresponding to the missing image type; and if an image of the missing image type uploaded based on the image upload control is obtained, then in response to a confirmation operation on the second image confirmation page, using the image type of each confirmed image as the target image type of each image.

[0113] In the above embodiments, after determining the image types of each image, to ensure comprehensiveness of image types, an image integrity check can be performed on the target image types based on the required image types. This step aims to check whether the user has uploaded all the necessary image types as required, and whether these images have been correctly categorized. If the image integrity check passes, it means that after automatic type recognition, the system considers the user to have uploaded all the necessary images. At this time, the system will display a first image confirmation page based on each image and its target image type. This page will list all uploaded images and their corresponding image types for the user to confirm. If the user confirms that everything is correct, these confirmed image types can be used as the final target image types for each image, and the subsequent information processing flow can continue. However, if the image integrity check fails, it means that the user may have missed uploading some necessary image types. At this time, a second image confirmation page can be displayed based on each image, its target image type, and the missing image types. This page not only lists the uploaded images and their corresponding image types, but also specifically indicates which image types are missing, and provides image upload controls for these missing image types. Users can use these upload controls to upload the missing images and make necessary categorization adjustments. After receiving the missing image type uploaded via the image upload control, the user can confirm it on the second image confirmation page. Once confirmed, these confirmed image types will be used as the final target image types for each image, and the subsequent information processing flow will continue. Through this image integrity verification and confirmation process, the information entry process of this embodiment ensures that all necessary images have been correctly uploaded and categorized, thereby improving the accuracy and reliability of information entry. Simultaneously, this process also provides a good user experience, allowing users to clearly understand which images have been uploaded, which images are missing, and how to make necessary adjustments and additions.

[0114] Step 203: For each image, construct text information analysis prompts based on the preset text type corresponding to the image and the target image type of the image, and analyze the text information that matches the preset text type using a large model based on the text information analysis prompts.

[0115] In this embodiment, to perform in-depth analysis and extraction of text information contained in images, a text information analysis prompt can be constructed based on the preset text type corresponding to the image and its target image type. Then, a large model is used to parse this prompt to extract text information matching the preset text type. In some implementations, one or more preset text types are determined based on each image and its target image type. These preset text types are defined according to the characteristics of the target image type and the information input requirements. For example, for an ID card image, the preset text type may include name, gender, date of birth, etc.; for a business license image, the preset text type may include the scope of business license, etc. Next, the text information analysis prompt is constructed. This information may include some basic features of the image (such as color, texture, shape, etc.), recognized image text information (such as text extracted through OCR technology), and relevant descriptions or keywords of the preset text type. This combination of information aims to provide a comprehensive context for the large model, enabling more accurate analysis and extraction of text information matching the preset text type. Then, the text information analysis prompt is input into a large model for analysis. The large model utilizes its powerful natural language processing capabilities and deep learning algorithms to deeply analyze the prompt and attempt to extract text information matching a preset text type. The large model is pre-trained using different types of images and their corresponding preset text types. Finally, based on the large model's analysis results, the text information matching the preset text type is extracted and further processed as part of the image information. This information can be configured for subsequent information entry, review, and analysis stages, providing strong support for the entire information entry process. The information entry process of this embodiment enables accurate extraction and analysis of text information contained in images, thereby improving the accuracy and efficiency of information entry. Simultaneously, this step also provides a more comprehensive and accurate data foundation for subsequent information processing and review.

[0116] In some embodiments of this application, before performing text information analysis on the images using a large model, the method further includes: for each image, extracting text information from the image text information corresponding to the image according to a preset text type corresponding to the image; correspondingly, constructing text information analysis prompt information based on the image and the preset text type corresponding to the target image type of the image, including: constructing text information analysis prompt information based on the image, the text extraction result of the image, and the preset text type corresponding to the image.

[0117] In this embodiment, to further improve the accuracy and efficiency of text information analysis, text extraction can be performed before using a large model to analyze the text information of the image. This step aims to extract text information from the image in a targeted manner according to the preset text type corresponding to the image, thereby providing more accurate and useful data for subsequent text information analysis. In some implementations, for each image, the system extracts text information from the image according to its corresponding preset text type. This extraction process may involve steps such as recognizing text in the image (e.g., through OCR technology), filtering and screening text content (e.g., removing irrelevant text information and retaining text information related to the preset text type). Through this step, text information closely related to the preset text type can be extracted from the image, providing strong support for subsequent analysis. For example, the preset text type corresponding to the front image of an ID card includes the ID card number and name. After completing the text extraction, a text information analysis prompt is constructed based on the extraction results, the image itself, and the preset text type corresponding to the target image type. This prompt not only includes basic image information and extracted text information, but also incorporates relevant descriptions or keywords for the preset text type, aiming to provide a comprehensive contextual environment for the large model. Through this prompt, the large model can more accurately understand the relationship between image content and the preset text type, thereby more effectively extracting and analyzing text information that matches the preset text type.

[0118] In some embodiments of this application, text information analysis prompts are constructed based on the image, the text extraction result of the image, and the preset text type corresponding to the image. This includes: for each image from which text information has been extracted, filling a text information confirmation prompt template based on the image, the text extraction result of the image, and the preset text type corresponding to the image to obtain the text information analysis prompts. The text information confirmation prompt template includes content configured to guide the large model to verify the text extraction result of the image based on the image and the preset text type corresponding to the image. For each image from which no text information has been extracted, filling a text information generation prompt template based on the image and the preset text type corresponding to the image to obtain the text information analysis prompts. The text information generation prompt template includes content configured to guide the large model to generate text information based on the image and the preset text type corresponding to the image.

[0119] In this embodiment, to process text information in images more accurately, two different strategies can be used to construct text information analysis prompts, one for images with extracted text information and the other for images without extracted text information. For images with extracted text information, a text information confirmation prompt template is populated based on the image itself, the text extraction results, and the corresponding preset text type. This template is designed to guide the large model to verify the extracted text information, ensuring it matches the image content and the preset text type. In some implementations, the template may include guiding statements or questions, such as "Please confirm whether the following text information matches the [preset text type] in the image," along with the extracted text information for the large model to compare and verify. In this way, the system can further improve the accuracy and reliability of the text information. For images without extracted text information, a text information generation prompt template is populated based on the image itself and the corresponding preset text type. This template is designed to guide the large model to generate text information based on the image content and the preset text type. In some implementations, the template may include suggestive statements or prompts, such as "Please generate corresponding text information based on [image features] and [preset text type] in the image," to stimulate the creativity and comprehension of the large model, thereby generating text information that matches the image content and the preset text type. Through these two different strategies, the system can flexibly handle different types of images and provide targeted text information analysis prompts based on the actual situation. This not only improves the accuracy and efficiency of text information processing but also enhances the system's adaptability and robustness.

[0120] Step 204: Perform image integrity verification on the target image type based on the required image type, and perform text integrity verification on the target text information based on the required text type.

[0121] In the information entry process of this application embodiment, to ensure that all necessary information has been entered correctly and accurately, thereby supporting the accurate filling of the application information page for the target function, image integrity verification can also be performed on the target image type according to the required image type to check whether all necessary images have been uploaded and correctly categorized. Simultaneously, the system will also perform text integrity verification on the target text information according to the required text type to ensure that all necessary text information has been correctly extracted or generated.

[0122] Step 205: If the image integrity check and the text integrity check pass, then based on the image, the target image type of the image, and the target text information, fill in the application information page of the target function.

[0123] In this embodiment, if both image integrity verification and text integrity verification pass, it means that all necessary information has been entered completely and accurately. At this point, based on the image, the target image type, and the target text information, the application information page for the target function is automatically populated. This step achieves automated information entry, greatly improving the efficiency and accuracy of information entry.

[0124] Step 206: If at least one of the image integrity check and the text integrity check fails, the missing item that failed the integrity check is determined; if the missing item is a predicted item, the missing item is predicted based on the image and the target text information using a large model, and the application information page of the target function is filled based on the predicted missing item information, the image, the target image type of the image, and the target text information; and if the missing item is not a predicted item, the application information page of the target function is filled directly based on the image, the target image type of the image, and the target text information.

[0125] In this embodiment, if at least one of the image integrity check and text integrity check fails, the missing items that failed the integrity check can be identified first. These missing items may include images that were not uploaded, images that were misclassified, missing text information, etc. Next, the system will adopt different processing strategies based on the nature of the missing items: If the missing item is a predicted item, meaning it can be predicted using other entered information (such as images and target text information), then the large model will continue to predict the missing item, and the application information page for the target function will be filled based on the predicted missing item information, the image, the target image type of the image, and the target text information. This step fully utilizes the predictive capabilities of the large model, which can, to some extent, compensate for the inconvenience caused by missing information. If the missing item is not a predicted item, meaning it cannot be predicted using other entered information, then the application information page for the target function will be filled directly based on the entered information (image, target image type of the image, and target text information), and the user will be prompted on the page to pay attention to the missing items. This step achieves continuity and completeness of information entry, while reminding the user to promptly supplement the missing information. Through this processing flow, the information entry process of this application embodiment can ensure that all necessary information has been correctly entered and provide reasonable processing strategies in the event of missing information, thereby supporting the accurate filling of the application information page for the target function. This process not only improves the efficiency and accuracy of information entry but also enhances the adaptability and user-friendliness of the system.

[0126] In some embodiments of this application, the predicted item includes a demand text type; in step 206, when the missing item is a predicted item, the missing item is predicted by a large model based on the image and the target text information, including: when the missing item is a predicted item, obtaining the associated item information corresponding to the missing item based on the target text information and the image text information corresponding to the image, constructing missing item prediction prompt information based on the associated item information and the missing item, and predicting the missing item by a large model based on the missing item prediction prompt information.

[0127] In this embodiment, when a missing item is detected as belonging to a predicted item (a certain type of requirement text), the system first obtains related item information based on the target text information and the corresponding image text information. This related item information may include other text information associated with the target text information, specific regions or features in the image, etc., which have some logical or semantic connection with the missing item. After obtaining the related item information, a missing item prediction prompt is constructed based on this information and the missing item itself. This prompt aims to provide the large model with a comprehensive contextual environment to more accurately understand and predict the content of the missing item. The prompt may contain a detailed description of the related item information, the requirements of the missing item, or the expected prediction result. Finally, this missing item prediction prompt is input into the large model for prediction. The large model utilizes its powerful natural language processing capabilities and deep learning algorithms to deeply analyze the prompt and attempt to infer the content of the missing item. The prediction result will serve as a supplement to the missing item information and, along with other entered information, will be configured to fill the application information page for the target function. Through this processing flow, the information entry process of this application embodiment can more accurately predict and process missing text information. Especially when the required text type is missing, this strategy not only improves the completeness and accuracy of information entry, but also enhances the intelligence of the system and the user experience.

[0128] In some embodiments of this application, the target function includes online store opening; the required image types include storefront image, store environment image, business license image, and operating permit image; the required text types include store name, store address, product category, and business hours; the predicted items include product category and business hours; when the missing item is a predicted item, based on the target text information and the image text information corresponding to the image, obtain the associated item information corresponding to the missing item, including: when the missing item includes product category, based on the store name, the image text information corresponding to the storefront image, the image text information corresponding to the business license image, and the operating permit image... The system determines the associated information for the business category by considering at least one of the following: image text information corresponding to the certificate image, food image information, and relevant business category information corresponding to the store name. The relevant stores include chain stores, and the food image information includes image text information identified from images with a target image type of menu and / or images with a target image type of food. If the missing item includes business hours, the system determines the surrounding store range corresponding to the store address, obtains information on each surrounding store within that range, including the business category and business hours of the surrounding stores, and determines the associated information for the business hours based on the surrounding store information and the business category.

[0129] In this embodiment, for the target function of online store opening, the required image types include storefront images, store environment images, business license images, and operating permit images. The required text types include store name, store address, product category, and business hours. The predicted items include product category and business hours. Of course, the predicted items can also include other types of items, which are not limited here. For the predicted item of product category, when its absence is detected, the associated item information can be determined based on multiple information sources. These information sources include, but are not limited to: store name (which may imply the business direction or characteristics), image text information corresponding to the storefront image (such as the text on the doorplate), image text information corresponding to the business license image (such as the description of the business scope), image text information corresponding to the operating permit image (such as the permitted business items), menu image information (including image text information identified in images with the target image type of menu and / or dishes identified in images with the target image type of dishes, which can directly reflect the types of dishes offered by the store), and related store product category information corresponding to the store name (such as the product categories of chain stores or stores of the same brand, which may be obtained through big data analysis and other technical means). By integrating this information, the large model can more accurately infer the missing business categories. It's worth noting that to enhance the large model's predictive ability regarding business categories, it can be pre-trained. Regarding sample construction, due to the uneven distribution of merchant categories, an upsampling and downsampling strategy is used to balance the sample size of different categories, avoiding bias during model training. For the prediction of business hours, when a missing item is detected, the surrounding store range corresponding to the store address can be determined first. This range may be based on geographical location information (such as latitude and longitude, street names, etc.). For example, the surrounding store range can be defined by a predetermined distance from the store address's geographical location, or by the business district where the store address is located. Then, information on each surrounding store within this range is obtained, including their business categories and business hours. After obtaining this information, relevant surrounding store information is filtered based on the missing item's business category, and the large model attempts to infer the missing business hours from the business hours of these relevant stores. This strategy utilizes the business hours of surrounding stores as a reference, improving the accuracy and practicality of the prediction. Through this processing flow, the information entry process of this application embodiment can more accurately predict and process the missing business categories and business hours information in the online store opening application. For cases where the user-provided pictures do not contain certain items, item prediction can be achieved. This strategy not only improves the completeness and accuracy of information entry, but also enhances the intelligence of the system and the user experience.

[0130] Step 207: Display the application information page, and if the application information page is confirmed, submit the input information for the target function based on the confirmed application information page.

[0131] In this embodiment, a processed and fully populated application information page is displayed to the user. This page contains all necessary information, such as store name, store address, product category, business hours, and images of the storefront, store environment, business license, and operating permit. Users can view and confirm the accuracy of all information on this page. If any errors or omissions are found, the user has the right to modify or supplement them. After confirming that all information on the application information page is accurate, the user can submit the application. The system receives the user's submission request and, based on the confirmed application information page, submits the data for the target function to the corresponding backend processing system or database. This step not only ensures the accuracy and completeness of all necessary information but also provides users with an intuitive and convenient interface to view and confirm their application information. Furthermore, the automated submission process significantly improves the efficiency and accuracy of information entry, providing a better user experience.

[0132] In some embodiments of this application, after displaying the application information page in step 207, the method further includes: obtaining information modification data of the application information page, and modifying specific types of information on the application information page based on the information modification data; recording the information modification data, and constructing large model training samples based on the information modification data for large model optimization training.

[0133] In this embodiment, after a user modifies the information on the application information page, the page content can be updated accordingly, and the modifications can be used to optimize the training of a large model. In some implementations, after displaying the application information page, the user is allowed to view and confirm all information. If the user finds any errors or needs to supplement information, they can modify it through the interface provided by the system. These modifications may involve specific types of information, such as store name, product category, business hours, etc. If the user submits modified data, the corresponding information on the application information page can be updated based on this modified data to ensure that the submitted information is the most accurate and complete. Furthermore, in addition to updating the application information page, the user's submitted information modification data can be further recorded to optimize the training of the large model. For example, if the user modifies the image type, image classification samples can be constructed based on the modified image and the modified image type to train the large model's image classification ability. As another example, if the user modifies the text information of a certain demand text type (not a prediction item), text analysis samples can be constructed based on the modified demand text type, the modified text information, and the image corresponding to that demand text type to train the large model's text analysis ability. For example, if a user modifies the text information of a certain requirement text type (belonging to the prediction item), then text prediction samples can be constructed based on the modified requirement text type, the modified text information, and the related item information corresponding to that requirement text type to train the text prediction capability of the large model. This optimization training method based on actual user feedback not only improves the prediction capability of the large model but also makes it more intelligent and user-friendly. Through continuous iteration and optimization, a more accurate and personalized service experience can be provided to users. This application embodiment further enhances the system's flexibility and intelligence by introducing information modification and page update functions, as well as a strategy for recording modification data and optimizing the training of the large model. These functions not only enhance the user experience but also provide strong support for the continuous optimization and upgrading of the system.

[0134] Furthermore, as an implementation of the methods in Figures 1 and 2, this application embodiment provides an information entry device based on a large model. Figure 3 shows a schematic diagram of the structure of an information entry device based on a large model provided in this application embodiment. As shown in Figure 3, the information entry device 300 based on a large model includes:

[0135] Display section 310 is configured to display an image input page for a target function, wherein the target function is a function requested by providing specific type information, and the specific type information includes image information of at least one required image type and text information of at least one required text type;

[0136] Analysis section 320 is configured to acquire at least one image entered based on the image entry page, and perform image type recognition and text information analysis on the image using a large model;

[0137] The filling part 330 is configured to fill the application information page of the target function based on the image, the target image type identified by the large model, and the analyzed target text information, wherein the application information page is a modifiable page containing the specific type of information;

[0138] The display portion 310 is further configured to display the application information page, and, if the application information page is confirmed, to submit the input information for the target function based on the confirmed application information page.

[0139] In some embodiments, the analysis section 320 is further configured to:

[0140] Perform text recognition on each image to obtain the text information for each image; and,

[0141] For any image, identify whether the image text information matches any keyword corresponding to any required image type, and determine the target image type of the image based on the required image type corresponding to the matched keyword;

[0142] Based on the required image type and the identified target image type, determine the second image type to be identified; and,

[0143] For each image whose image type is not determined, the feature description information corresponding to the second image type is obtained. The image type identification prompt template is filled according to the image, the second image type, and the feature description information to obtain the image type identification prompt information. The target image type of the image is identified at least in the second image type using a large model based on the image type identification prompt information.

[0144] In some embodiments, the analysis section 320 is further configured to:

[0145] Based on the required image type, perform image integrity verification on the target image type;

[0146] The display portion 310 is further configured to: if the image integrity verification passes, display a first image confirmation page based on each image and its target image type, and in response to a confirmation operation on the first image confirmation page, use the image type of each confirmed image as its target image type; and,

[0147] If the image integrity check fails, a second image confirmation page is displayed based on each image, the target image type of each image, and the missing image type. The second image confirmation page includes an image upload control corresponding to the missing image type. If an image of the missing image type is uploaded based on the image upload control, in response to the confirmation operation on the second image confirmation page, the image type of each confirmed image is used as the target image type of each image.

[0148] In some embodiments, the analysis section 320 is further configured to:

[0149] For each image, text extraction is performed on the image text information corresponding to the image according to the preset text type corresponding to the image;

[0150] For each image, based on the image, the text extraction results of the image, and the preset text type corresponding to the image, text information analysis prompts are constructed, and a large model is used to analyze text information that matches the preset text type based on the text information analysis prompts.

[0151] For each image from which text information has been extracted, a text information confirmation prompt template is populated based on the image, the text extraction result of the image, and the preset text type corresponding to the image, to obtain the text information analysis prompt information. The text information confirmation prompt template includes content configured to guide the large model to verify the text extraction result of the image based on the image and the preset text type corresponding to the image; and...

[0152] For each image from which no text information has been extracted, a text information generation prompt template is filled in according to the image and the preset text type corresponding to the image to obtain the text information analysis prompt information. The text information generation prompt template includes content configured to guide the large model to generate text information based on the image and the preset text type corresponding to the image.

[0153] In some embodiments, the analysis section 320 is further configured to:

[0154] Based on the required image type, perform image integrity verification on the target image type; and based on the required text type, perform text integrity verification on the target text information.

[0155] The filling portion 330 is further configured to: if the image integrity check and the text integrity check pass, then execute the step of filling the application information page of the target function based on the image, the target image type of the image, and the target text information; and,

[0156] If at least one of the image integrity checks and the text integrity checks fails, a missing item that failed the integrity check is identified. If the missing item is a predicted item, based on the target text information and the image text information corresponding to the image, the associated item information corresponding to the missing item is obtained. Based on the associated item information and the missing item, a missing item prediction prompt is constructed. The missing item is predicted using a large model based on the missing item prediction prompt. The application information page for the target function is filled based on the predicted missing item information, the image, the target image type of the image, and the target text information. If the missing item is not a predicted item, the application information page for the target function is filled directly based on the image, the target image type of the image, and the target text information.

[0157] In some implementations, the target function includes opening an online store; the required image types include storefront images, store environment images, business license images, and operating permit images; the required text types include store name, store address, product category, and business hours; and the predicted items include product category and business hours.

[0158] If the missing item is a predicted item, the analysis section 320 is further configured as follows:

[0159] If the missing item includes a product category, the associated item information corresponding to the product category is determined based on at least one of the following: the store name, the image and text information corresponding to the storefront image, the image and text information corresponding to the business license image, the image and text information corresponding to the operating permit image, the menu image information, and the relevant store product category information corresponding to the store name; wherein, the relevant store includes chain stores, and the menu image information includes image and text information identified in images with a target image type of menu and / or images with a target image type of dish; and,

[0160] If the missing item includes business hours, determine the range of surrounding stores corresponding to the store address, obtain information on each surrounding store within the range of surrounding stores, the surrounding store information includes the business categories and business hours of the surrounding stores, and determine the associated item information corresponding to the business hours based on the surrounding store information and the business categories.

[0161] In some embodiments, the analysis section 320 is further configured to:

[0162] The image is subjected to watermark recognition and / or similar image recognition based on a preset image library to determine whether the image is a forged image, wherein the forged image includes watermarked images and images that meet the similarity criteria;

[0163] If the image is a forged image, a reminder page will be displayed to prompt correction of the forged image; and,

[0164] For any image, based on the target image type of the image, at least one comparison image is determined from each preset image in the preset image library, and the similarity between the image and each comparison image is calculated according to the comparison conditions corresponding to the target image type of the image. The comparison conditions for text images include comparing text similarity, and the comparison conditions for image images include comparing pixel grayscale value similarity.

[0165] In some embodiments, the analysis section 320 is further configured to:

[0166] Obtain information modification data for the application information page, and modify specific types of information on the application information page based on the information modification data; and,

[0167] The modified information data is recorded, and large model training samples are constructed based on the modified information data to optimize the training of the large model.

[0168] It should be noted that other corresponding descriptions of the various functional parts involved in the information input device 300 based on a large model provided in the embodiments of this application can be found in the corresponding descriptions in the methods of Figures 1 and 2, and will not be repeated here.

[0169] This application also provides a computer device, which can be a personal computer, server, network device, etc. The computer device includes a bus, processor, memory, and communication interface, and may also include input / output interface and display device. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database of the computer device is configured to store location information. The network interface of the computer device is configured to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements the steps in the various method embodiments.

[0170] Those skilled in the art will understand that the structure of the computer device described above is only a partial structure related to the present application and does not constitute a limitation on the computer device that should be configured on the present application. The computer device may include more or fewer components, or combine certain components, or have different component arrangements.

[0171] In one embodiment, a computer-readable storage medium is provided, which may be non-volatile or volatile, having stored thereon a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0172] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0173] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data configured for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0174] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, graphics processors, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0175] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0176] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. An information entry method based on a large model, wherein, The method includes: Displays an image input page for the target function, wherein the target function is a function applied for by providing specific types of information, and the specific types of information include image information of at least one required image type and text information of at least one required text type; Obtain at least one image entered based on the image entry page, and perform image type recognition and text information analysis on the image using a large model; Based on the image, the target image type identified by the large model, and the analyzed target text information, the application information page for the target function is populated, wherein the application information page is a modifiable page containing the specific type of information; and, The application information page is displayed, and if the application information page is confirmed, the input information for the target function is submitted based on the confirmed application information page.

2. The method according to claim 1, wherein, Before performing image type recognition on the image using a large model, the method further includes: Perform text recognition on each image to obtain the text information for each image; and, For any image, identify whether the image text information of the image matches the keyword corresponding to the first image type, and determine the target image type of the image based on the required image type corresponding to the matched keyword, wherein the first image type includes the standard image type in the required image type; Accordingly, image type identification is performed on the image using a large model, including: Based on the required image type and the identified target image type, determine the second image type to be identified; and, For each image whose image type is not determined, the feature description information corresponding to the second image type is obtained. The image type identification prompt template is filled according to the image, the second image type, and the feature description information to obtain the image type identification prompt information. The target image type of the image is identified at least in the second image type using a large model based on the image type identification prompt information.

3. The method according to claim 1, wherein, Before performing text information analysis on the image using a large model, the method further includes: Based on the required image type, perform image integrity verification on the target image type; If the image integrity verification passes, a first image confirmation page is displayed based on each image and its target image type. In response to a confirmation operation on the first image confirmation page, the image type of each confirmed image is used as its target image type. If the image integrity check fails, a second image confirmation page is displayed based on each image, the target image type of each image, and the missing image type. The second image confirmation page includes an image upload control corresponding to the missing image type. If an image of the missing image type is uploaded based on the image upload control, in response to the confirmation operation on the second image confirmation page, the image type of each confirmed image is used as the target image type of each image.

4. The method according to claim 1, wherein, Textual information analysis of the images is performed using a large model, including: For each image, text extraction is performed on the corresponding image text information based on the preset text type; and, For each image, based on the image, the text extraction results of the image, and the preset text type corresponding to the image, text information analysis prompts are constructed, and a large model is used to analyze text information that matches the preset text type based on the text information analysis prompts. Specifically, based on the image, the text extraction results of the image, and the preset text type corresponding to the image, a text information analysis prompt is constructed, including: For each image from which text information has been extracted, a text information confirmation prompt template is populated based on the image, the text extraction result of the image, and the preset text type corresponding to the image, to obtain the text information analysis prompt information. The text information confirmation prompt template includes content configured to guide the large model to verify the text extraction result of the image based on the image and the preset text type corresponding to the image; and... For each image from which no text information has been extracted, a text information generation prompt template is filled in according to the image and the preset text type corresponding to the image to obtain the text information analysis prompt information. The text information generation prompt template includes content configured to guide the large model to generate text information based on the image and the preset text type corresponding to the image.

5. The method according to claim 1, wherein, Before filling the application information page for the target function with the image, the target image type identified by the large model, and the analyzed target text information, the method further includes: Based on the required image type, perform image integrity verification on the target image type; and based on the required text type, perform text integrity verification on the target text information. If the image integrity check and the text integrity check pass, then the step of filling the application information page for the target function based on the image, the target image type of the image, and the target text information is executed; and, If at least one of the image integrity checks and the text integrity checks fails, a missing item that failed the integrity check is identified. If the missing item is a predicted item, based on the target text information and the image text information corresponding to the image, the associated item information corresponding to the missing item is obtained. Based on the associated item information and the missing item, a missing item prediction prompt is constructed. The missing item is predicted using a large model based on the missing item prediction prompt. The application information page for the target function is filled based on the predicted missing item information, the image, the target image type of the image, and the target text information. If the missing item is not a predicted item, the application information page for the target function is filled directly based on the image, the target image type of the image, and the target text information.

6. The method according to claim 5, wherein, The target function includes online store opening; the required image types include storefront image, store environment image, business license image, and operating permit image; the required text types include store name, store address, product category, and business hours; and the predicted items include product category and business hours. When the missing item is a predicted item, based on the target text information and the image text information corresponding to the image, the associated item information corresponding to the missing item is obtained, including: If the missing item includes a product category, the associated item information corresponding to the product category is determined based on at least one of the following: the store name, the image and text information corresponding to the storefront image, the image and text information corresponding to the business license image, the image and text information corresponding to the operating permit image, the menu image information, and the relevant store product category information corresponding to the store name; wherein, the relevant store includes chain stores, and the menu image information includes image and text information identified in images with a target image type of menu and / or images with a target image type of dish; and, If the missing item includes business hours, determine the range of surrounding stores corresponding to the store address, obtain information on each surrounding store within the range of surrounding stores, the surrounding store information includes the business categories and business hours of the surrounding stores, and determine the associated item information corresponding to the business hours based on the surrounding store information and the business categories.

7. The method according to any one of claims 1 to 6, wherein, The method further includes: The image is subjected to watermark recognition and / or similar image recognition based on a preset image library to determine whether the image is a forged image, wherein the forged images include watermarked images and images that meet similarity criteria; and, If the image is a forged image, a reminder page will be displayed to prompt the user to correct the forged image. The process of identifying similar images based on a preset image library includes: For any image, based on the target image type of the image, at least one comparison image is determined from each preset image in the preset image library, and the similarity between the image and each comparison image is calculated according to the comparison conditions corresponding to the target image type of the image. The comparison conditions for text images include comparing text similarity, and the comparison conditions for image images include comparing pixel grayscale value similarity.

8. The method according to any one of claims 1 to 6, wherein, After displaying the application information page, the method further includes: Obtain information modification data for the application information page, and modify specific types of information on the application information page based on the information modification data; and, The modified information data is recorded, and large model training samples are constructed based on the modified information data to optimize the training of the large model.

9. An information input device based on a large model, wherein, The device includes: The display section is configured to display an image input page for a target function, wherein the target function is a function requested by providing specific type information, and the specific type information includes image information of at least one required image type and text information of at least one required text type; The analysis section is configured to acquire at least one image entered based on the image entry page, and perform image type recognition and text information analysis on the image using a large model; The filling section is configured to fill the application information page of the target function based on the image, the target image type identified by the large model, and the analyzed target text information, wherein the application information page is a modifiable page containing the specific type of information; and, The display section is further configured to display the application information page, and, if the application information page is confirmed, to submit the input information for the target function based on the confirmed application information page.

10. The apparatus according to claim 9, wherein, Before performing image type identification on the image using a large model, the analysis section is further configured as follows: Perform text recognition on each image to obtain the text information for each image; and, For any image, identify whether the image text information of the image matches the keyword corresponding to the first image type, and determine the target image type of the image based on the required image type corresponding to the matched keyword, wherein the first image type includes the standard image type in the required image type; Based on the required image type and the identified target image type, determine the second image type to be identified; and, For each image whose image type is not determined, the feature description information corresponding to the second image type is obtained. The image type identification prompt template is filled according to the image, the second image type, and the feature description information to obtain the image type identification prompt information. The target image type of the image is identified at least in the second image type using a large model based on the image type identification prompt information.

11. The apparatus according to claim 9, wherein, The analysis section is further configured to: perform image integrity verification on the target image type based on the required image type; The display section is further configured to: if the image integrity verification is passed, then display a first image confirmation page based on each image and the target image type of each image, and in response to the confirmation operation of the first image confirmation page, use the image type of each confirmed image as the target image type of each image; as well as, If the image integrity check fails, a second image confirmation page will be displayed based on each image, the target image type of each image, and the missing image type. The second image confirmation page includes an image upload control corresponding to the missing image type. If an image of the missing image type is obtained based on the image upload control, in response to the confirmation operation on the second image confirmation page, the image type of each confirmed image is taken as the target image type of each image.

12. The apparatus according to claim 9, wherein, The analysis section is also configured as follows: For each image, text extraction is performed on the corresponding image text information based on the preset text type; and, For each image, based on the image, the text extraction results of the image, and the preset text type corresponding to the image, text information analysis prompts are constructed, and a large model is used to analyze text information that matches the preset text type based on the text information analysis prompts. Specifically, based on the image, the text extraction results of the image, and the preset text type corresponding to the image, a text information analysis prompt is constructed, including: For each image from which text information has been extracted, a text information confirmation prompt template is populated based on the image, the text extraction result of the image, and the preset text type corresponding to the image, to obtain the text information analysis prompt information. The text information confirmation prompt template includes content configured to guide the large model to verify the text extraction result of the image based on the image and the preset text type corresponding to the image; and... For each image from which no text information has been extracted, a text information generation prompt template is filled in according to the image and the preset text type corresponding to the image to obtain the text information analysis prompt information. The text information generation prompt template includes content configured to guide the large model to generate text information based on the image and the preset text type corresponding to the image.

13. The apparatus according to claim 9, wherein, The analysis section is also configured as follows: Based on the required image type, perform image integrity verification on the target image type; and based on the required text type, perform text integrity verification on the target text information. The filling part is further configured to: if the image integrity check and the text integrity check pass, then execute the step of filling the application information page of the target function based on the image, the target image type of the image, and the target text information; as well as, If at least one of the image integrity checks and the text integrity checks fails, a missing item that failed the integrity check is identified. If the missing item is a predicted item, based on the target text information and the image text information corresponding to the image, the associated item information corresponding to the missing item is obtained. Based on the associated item information and the missing item, a missing item prediction prompt is constructed. The missing item is predicted using a large model based on the missing item prediction prompt. The application information page for the target function is filled based on the predicted missing item information, the image, the target image type of the image, and the target text information. If the missing item is not a predicted item, the application information page for the target function is filled directly based on the image, the target image type of the image, and the target text information.

14. The apparatus according to claim 13, wherein, The target function includes online store opening; the required image types include storefront image, store environment image, business license image, and operating permit image; the required text types include store name, store address, product category, and business hours; and the predicted items include product category and business hours. If the missing item is a predicted item, the analysis section is further configured as follows: If the missing item includes a product category, the associated item information corresponding to the product category is determined based on at least one of the following: the store name, the image and text information corresponding to the storefront image, the image and text information corresponding to the business license image, the image and text information corresponding to the operating permit image, the menu image information, and the relevant store product category information corresponding to the store name; wherein, the relevant store includes chain stores, and the menu image information includes image and text information identified in images with a target image type of menu and / or images with a target image type of dish; and, If the missing item includes business hours, determine the range of surrounding stores corresponding to the store address, obtain information on each surrounding store within the range of surrounding stores, the surrounding store information includes the business categories and business hours of the surrounding stores, and determine the associated item information corresponding to the business hours based on the surrounding store information and the business categories.

15. The apparatus according to any one of claims 9 to 14, wherein, The analysis section is also configured as follows: The image is subjected to watermark recognition and / or similar image recognition based on a preset image library to determine whether the image is a forged image, wherein the forged image includes watermarked images and images that meet the similarity criteria; If the image is a forged image, a reminder page will be displayed to prompt correction of the forged image; and, For any image, based on the target image type of the image, at least one comparison image is determined from each preset image in the preset image library, and the similarity between the image and each comparison image is calculated according to the comparison conditions corresponding to the target image type of the image. The comparison conditions for text images include comparing text similarity, and the comparison conditions for image images include comparing pixel grayscale value similarity.

16. The apparatus according to any one of claims 9 to 14, wherein, The analysis section is also configured as follows: Obtain information modification data for the application information page, and modify specific types of information on the application information page based on the information modification data; and, The modified information data is recorded, and large model training samples are constructed based on the modified information data to optimize the training of the large model.

17. A storage medium having a computer program stored thereon, wherein, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 8.

18. A computer device comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein, When the processor executes the computer program, it implements the method of any one of claims 1 to 8.

Citation Information

Patent Citations

  • Data entry method, device and equipment based on artificial intelligence and storage medium

    CN116453125A

  • Big language model-based bill identification method and apparatus, and storage medium

    CN117253248A

  • Information processing method and device, storage medium and electronic equipment

    CN117422519A

  • Merchant information auditing method and device, electronic equipment and medium

    CN118195536A

  • Information input method based on large model, storage medium and computer equipment

    CN119477338A