Business data retrieval method and device, business data auditing method and device, equipment, medium and product

By enhancing and encoding business data and generating feature data using the CLIP model, the problems of low efficiency and accuracy in business data retrieval and review are solved, achieving high accuracy and generalizability in the absence of data.

CN120973885APending Publication Date: 2025-11-18CHINA MOBILE (XIONGAN) ICT CO LTD +3
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511010045.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-22
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing technologies suffer from low efficiency and accuracy in retrieving and reviewing business data from massive amounts of image content data, relying on manual or automated methods. In particular, when business data is lacking, it is impossible to train a model that meets business needs.

Method used

By performing retrieval enhancement operations on the target business data, it is enhanced into training data of the corresponding data type in the CLIP model training set. The CLIP model is then used for encoding to generate feature data, which is then combined with a classifier to verify the matching of text and image data.

Benefits of technology

Without requiring additional business data, it improves the accuracy and generalization of business data retrieval and auditing, and solves the problem of low precision caused by insufficient business data acquisition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120973885A_ABST
    Figure CN120973885A_ABST
Patent Text Reader

Abstract

The invention discloses a business data retrieval method and device, an auditing method and device, equipment, a medium and a product, and the method comprises the steps: carrying out retrieval enhancement operation on target business data, so as to enhance the target business data into training data of a corresponding data type in a training set of a CLIP model, the training set is used for training the CLIP model, and the training set is used for training the CLIP model; the target business data comprises text retrieval data and / or picture retrieval data; encoding the target business data after retrieval enhancement according to a CLIP model to obtain target feature data; and retrieving a matched retrieval picture based on the target feature data.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, and in particular to a business data retrieval method, an auditing method and device, equipment, medium and product. BACKGROUND

[0002] In today's digital society, information carried by pictures is increasing, and picture retrieval and auditing technology plays an indispensable role in many important fields such as security monitoring and social governance. In the face of massive picture content data, the traditional way of relying on manual retrieval or auditing exposes the problem of low efficiency.

[0003] With the development of deep learning technology, some automatic methods have also appeared, but such methods have generalization problems. When encountering targets outside the training data distribution, relevant business data needs to be collected for retraining, and in some business data sensitive scenarios, such data is not easy to obtain. Without business data, a model that meets business needs cannot be trained, and the retrieval and / or auditing of business data has the problem of low precision. SUMMARY

[0004] The purpose of the embodiments of the present application is to provide a business data retrieval method, an auditing method and device, equipment, medium and product to solve the problem of low precision of business data retrieval or auditing.

[0005] In order to solve the above technical problems, the present specification is implemented as follows: In a first aspect, a business data retrieval method is provided, comprising: performing a retrieval enhancement operation on target business data to enhance the target business data into training data of a corresponding data type in a training set of a CLIP model, the training set being used to train the CLIP model, the target business data including text-based retrieval data and / or picture-based retrieval data; encoding the target business data after retrieval enhancement according to a CLIP model to obtain target feature data; retrieving and matching a retrieval picture based on the target feature data.

[0006] In a second aspect, a business data auditing method is provided, comprising: performing a retrieval enhancement operation on target business data to enhance the target business data into training data of a corresponding data type in a training set of a CLIP model, the training set being used to train the CLIP model, the target business data including text-based retrieval data and picture-based retrieval data; The enhanced text retrieval data and enhanced image retrieval data are encoded according to the CLIP model to obtain the corresponding text feature data and image feature data. Image-text pairs, composed of text feature data and image feature data, are input into a preset classifier to check whether the text-class retrieval data and the image-class retrieval data match. The classifier is trained using image-text pairs as samples and whether the image-text pairs describe the same object as labels.

[0007] Thirdly, a business data retrieval device is provided, comprising: An enhancement module is used to perform retrieval enhancement operations on target business data, so as to enhance the target business data into training data of the corresponding data type in the training set of the CLIP model. The training set is used to train the CLIP model. The target business data includes text-based retrieval data and / or image-based retrieval data. The encoding module is used to encode the enhanced target business data according to the CLIP model to obtain target feature data; The retrieval module is used to retrieve matching images based on the target feature data.

[0008] Fourthly, a business data auditing device is provided, comprising: An enhancement module is used to perform retrieval enhancement operations on target business data, so as to enhance the target business data into training data of the corresponding data type in the training set of the CLIP model. The training set is used to train the CLIP model. The target business data includes text-based retrieval data and image-based retrieval data. The encoding module is used to encode the enhanced text retrieval data and the enhanced image retrieval data according to the CLIP model to obtain the corresponding text feature data and image feature data. The review module is used to input image-text pairs composed of text feature data and image feature data into a preset classifier to review whether the text-class retrieval data and the image-class retrieval data match. The classifier is trained using image-text pairs as samples and whether the image-text pairs describe the same object as labels.

[0009] Fifthly, an electronic device is provided, including a processor and a memory, the memory storing a program or instructions executable on the processor, the program or instructions, when executed by the processor, implementing the steps of the method as described in the first or second aspect.

[0010] In a sixth aspect, a readable storage medium is provided, on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method as described in the first or second aspect.

[0011] In a seventh aspect, a computer program product is provided, comprising a non-transitory computer-readable storage medium storing a computer program operable to cause a computer to perform the steps of the method as described in the first or second aspect.

[0012] In this embodiment, target business data is enhanced by performing a retrieval enhancement operation to transform it into training data of the corresponding data type in the CLIP model's training set. The training set is used to train the CLIP model, and the target business data includes text-based retrieval data and image-based retrieval data. The enhanced text-based and image-based retrieval data are encoded according to the CLIP model to obtain corresponding text feature data and image feature data. Image-text pairs composed of the text feature data and image feature data are input into a preset classifier to verify whether the text-based and image-based retrieval data match. The classifier is trained using image-text pairs as samples and whether the image-text pairs describe the same object as labels. This eliminates the need for extensive business data for model training. By enhancing the business data to be reviewed without requiring additional business data, the accuracy of subsequent review tasks is improved, and strong generalization is achieved. This alleviates the problem of insufficient or unavailable actual business data leading to unsatisfactory review results. Attached Figure Description

[0013] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 This is a flowchart illustrating the business data retrieval method according to an embodiment of this application.

[0014] Figure 2 This is a schematic diagram of the text retrieval enhancement process of the business data retrieval method according to an embodiment of this application.

[0015] Figure 3 This is a schematic diagram of the image retrieval enhancement process of the business data retrieval method according to an embodiment of this application.

[0016] Figure 4 This is a schematic diagram of the application process of the business data retrieval method according to an embodiment of this application.

[0017] Figure 5This is a flowchart illustrating the business data review method according to an embodiment of this application.

[0018] Figure 6 This is a schematic diagram illustrating the application process of the business data verification method in an embodiment of this application.

[0019] Figure 7 This is a structural block diagram of an electronic device according to an embodiment of this application. Detailed Implementation

[0020] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application. The drawing numbers in this application are only used to distinguish the various steps in the solution and are not used to limit the execution order of the various steps. The specific execution order is subject to the description in the specification.

[0021] To address the problems existing in the prior art, embodiments of this application provide a business data retrieval method, such as... Figure 1 As shown, the process includes steps 102 to 106.

[0022] Step 102: Perform retrieval enhancement operations on the target business data to enhance the target business data into training data of the corresponding data type in the training set of the CLIP model. The training set is used to train the CLIP model, and the target business data includes text-based retrieval data and / or image-based retrieval data.

[0023] Business data refers to data specific to a particular business scenario, such as proprietary or frequently occurring data in fields like transportation, medicine, and finance. The purpose of this application's embodiments is to retrieve corresponding matching images based on target business data input by the user. Target business data includes text-based retrieval data and / or image-based retrieval data. Text-based retrieval data refers to text data, and image-based retrieval data refers to image data. Users can input corresponding text data and / or image data to retrieve image data matching the input.

[0024] Before performing a search based on the target business data, a search enhancement operation is first performed to transform the target business data into training data of the corresponding data type in the CLIP model's training set. The CLIP model is a typical multimodal visual large-scale model that uses either text-based or image-based image search to achieve retrieval.

[0025] Text-to-image search places certain demands on the text input because there are different expressions for the same concept. For example, in the financial field, the English expression for "automatic teller machine" can be either "ATM" or "Cash machine." Since it's impossible to predict in advance which expression the CLIP model used for retrieval will adopt, it will affect the accuracy of subsequent text search results.

[0026] The text representations sampled by the CLIP model are consistent with the training data in the training set used to train the CLIP model. The training set includes both text and image training data. If the way the user inputs search terms is expressed differs from the way they are expressed in the CLIP model's training set for the same concept, the accuracy of the search will be greatly reduced.

[0027] In this embodiment of the application, even if the expression method corresponding to the CLIP model training set cannot be known in advance, by performing retrieval enhancement operations on the target business data, the enhanced target business data can be made consistent with the expression method in the CLIP model training set used for subsequent retrieval, thereby significantly improving the accuracy or precision of the retrieval.

[0028] Based on the solution provided in the above embodiments, optionally, in step 102, the retrieval enhancement operation on the target business data includes: when the target business data includes text-based retrieval data, deriving synonyms for the text-based retrieval data; matching the derived synonyms with text-based training data in the training set of the CLIP model; and filtering text-based training data related to the text-based retrieval data from the matched text-based training data to obtain the retrieval-enhanced target business data.

[0029] The following is combined with Figure 2 The text retrieval enhancement process of the business data retrieval method in the embodiments of this application is described.

[0030] like Figure 2 As shown, when the business data to be retrieved includes text data, users will use words commonly used in the business scenario to describe the target they want to retrieve (step 202). Then, synonym deduction is performed on the text data input by the user (step 204). In one embodiment, synonym deduction can be performed using the chatGPT model. For example, if the text data input by the user is "bus", then the chatGPT model can be asked a question such as "what are some common ways of referring to {bus}?#What are the common ways of referring to buses?", which can deduce synonyms for "bus" such as "City Bus".

[0031] Based on these synonyms, string matching is performed with the text class training data (pretraining caption) of the CLIP model (step 206) to obtain the text class training data matched by each synonym. A synonym may match one or more text class training data.

[0032] The matched text training data may contain descriptions unrelated to the user's input text data. For example, synonyms derived from "bus" might match "bus stop#bus station," which could cause ambiguity. Therefore, it is necessary to filter the matched text training data to find those related to the user's input text data, i.e., to select expressions that meet the requirements (step 208) and filter out irrelevant text training data.

[0033] One way to filter relevant text training data is to first match based on text similarity, and then compare it with a set similarity threshold to filter the data. However, the level of the similarity threshold will affect the filtering results.

[0034] In this embodiment, a specialized prompt word structure can be designed to filter irrelevant text-based training data. Taking the user-input text "bus" as an example, the Large Language Model Meta AI (LLAMA) and the input prompt word "does {bus stop} in the {caption}refer to {bus}#whether the 'bus stop' in the text-based training data indicates 'bus'" prompt the LLAMA model to obtain text-based training data related to the text data "bus".

[0035] Therefore, step 208 completes the process of enhancing the retrieval of text-based retrieval data, resulting in enhanced text-based retrieval data that matches and is related to the text-based training data in the CLIP model.

[0036] Based on the solution provided in the above embodiments, optionally, in step 102, the retrieval enhancement operation on the target business data includes: when the target business data includes image-type retrieval data, dividing the image-type retrieval data into multiple sub-image-type retrieval data; determining the image description text corresponding to each of the image-type retrieval data and the multiple sub-image-type retrieval data; grouping the image-text pairs composed of each image-type retrieval data and the corresponding image description text based on a preset classifier, wherein the image-text pairs in the same group describe the same object, and the classifier is trained using image-text pairs as samples and whether the image-text pairs describe the same object as labels; merging the image-type retrieval data in each image-text pair in the same group based on pixel position to obtain the retrieval-enhanced target business data.

[0037] The following is combined Figure 3 The image retrieval enhancement process of the business data retrieval method in the embodiments of this application is described.

[0038] like Figure 3 As shown, the business data to be retrieved input by the user includes image data (step 302). The image data is then divided into multiple sub-images (step 304). In one embodiment, the Segmentation and Masking Model (SAM) can be used to segment the image data into multiple sub-images.

[0039] The sub-images obtained by SAM may also be meaningless noise blocks that do not include objects, such as the background of the image, or include overlapping regions that belong to different objects. To solve the above problems, the image text descriptions of the whole image input by the user and the corresponding segmented sub-images are further extracted. In one embodiment, the above images can be input into the LLAVA multimodal model to extract and output the text descriptions of each image (step 306), thus obtaining the original expression of the semantic features of each image. If the user inputs an image-based retrieval data and it is segmented into M sub-images, after the image text description extraction, M+1 images and their corresponding image text descriptions can be obtained, thus forming M+1 image-text pairs.

[0040] By inputting these image-text pairs into a pre-trained classifier, it can be determined whether each image-text pair describes the same object. The classifier is trained using multiple image-text pairs as samples and the label indicating whether each image-text pair describes the same object. Based on the classifier's result indicating whether they describe the same object, the image-text pairs are grouped. Image-text pairs in the same group describe the same object, and the text within the image-text pairs can be used to determine whether two image-text pairs describe the same object.

[0041] Then, the pictures in each picture-text pair belonging to the same group are merged (step 308), and the merged pictures describe the same object. When merging the pictures in the picture-text pairs of the same group, the merging can be performed based on whether the pixel positions of the pictures are adjacent or whether there are adjacent regions, etc. Thus, for M + 1 picture-text pairs, since there is at least one merge, K + 1 pictures can be obtained, where K < M. The K + 1 pictures are the set of sub-pictures that meet the requirements and may also include the entire picture input by the user. The sub-pictures determined to be the picture background according to the picture description are meaningless noises and can be filtered out from the set of sub-pictures.

[0042] Through the above retrieval enhancement operation on picture-type data, the background interference in the picture-type data input by the user and the inconsistency between the picture-type data and the picture-type training data used by the CLIP model can be solved. The picture-type training data of CLIP usually contains one object, while the picture-type data in actual business may contain multiple objects and have large background interference. In the embodiment of the present application, by segmenting the target business data of the picture type as described above, extracting the picture description text in each of the multiple sub-pictures after segmentation, and performing object consistency grouping and same-object merging on the obtained picture-text pairs, the background interference in the picture-type retrieval data input by the user can be reduced, and the object main body in the picture can be highlighted.

[0043] In one embodiment, further, before grouping the picture-text pairs composed of each picture-type retrieval data and the corresponding picture description text based on a preset classifier, it further includes: obtaining multiple text-type business data of the business scenario to which the target business data belongs; performing a retrieval enhancement operation on each text-type business data to enhance each text-type business data into the text-type training data in the training set of the CLIP model; encoding the text-type business data after retrieval enhancement according to the CLIP model to obtain corresponding target feature data; retrieving and matching the retrieval pictures based on the target feature data; using the picture-text pairs composed of the text-type training data corresponding to each text-type business data and the retrieval pictures as samples and using whether the picture-text pairs describe the same object as a label to train the classifier.

[0044] In this embodiment, the classifier is trained based on multiple text-type business data of the business scenario to which the target business data input by the user belongs. Among them, the multiple text-type business data are enhanced through retrieval to be consistent with the text-type training data of the CLIP model. By encoding and retrieving and matching the text-type business data after retrieval enhancement, retrieval pictures consistent with the picture-type training data of the CLIP model can be further obtained.

[0045] In this way, by using image-text pairs consisting of text training data corresponding to each text-type business data and retrieved images as samples, and by labeling whether the image-text pairs describe the same object, a classifier can be trained. This not only solves the problem of inconsistency between the user-input retrieval text and / or retrieval images and the training data of the CLIP model used for retrieval in real business scenarios, but also establishes the association between image enhancement and text enhancement by classifying the image-type retrieval data included in the target business data through a classifier suitable for the corresponding business scenario, thereby improving the retrieval accuracy of image-type retrieval data.

[0046] Step 104: Encode the enhanced target business data according to the CLIP model to obtain target feature data.

[0047] The CLIP model performs feature encoding on the enhanced target business data, which can convert text-based or image-based retrieval data into corresponding vector-based feature data for subsequent retrieval.

[0048] Combination Figure 2 In step 210, the enhanced text-based retrieval data can be encoded using the CLIP model to obtain the corresponding text features.

[0049] Step 106: Retrieve matching images based on the target feature data.

[0050] Based on the solution provided in the above embodiments, optionally, in step 106 above, retrieving a matching search image based on the target feature data includes: calculating the similarity between the target feature data and feature data in a preset database to determine the matching feature data; determining the image corresponding to the matching target feature data as the matching search image, wherein the feature data in the preset database is obtained by encoding target data using the CLIP model, and the target data includes data after performing search enhancement operations on image-related business data of each image category in the business scenario to which the target business data belongs, and data after performing search enhancement operations on image description text corresponding to each image category business data.

[0051] For enhanced text-based search data, combined with Figure 2 The text features encoded by the CLIP model can be matched with the CLIP model's existing feature database using similarity calculations. The feature data in the existing feature database is obtained by encoding the text and image training data in the training set by the CLIP model. Based on the text features that match the enhanced text retrieval data, the corresponding text training data and the corresponding image training data (pretraining image data) can be determined, thus obtaining the matched retrieval image (step 212).

[0052] In one embodiment, to improve the retrieval accuracy of target business data in a specific business scenario, a feature database specifically for the business scenario to which the target business data belongs can be pre-established. That is, each image-type business data and its corresponding image description text within the business scenario to which the target business data belongs is acquired, and retrieval enhancement operations are performed on each. The CLIP model encodes the corresponding data after retrieval enhancement to obtain the corresponding text features and image features. These features correspond to the training data of the CLIP model for the corresponding data type.

[0053] As described above, the classifier in the above embodiment can be trained based on image-text pairs composed of text training data corresponding to each text class business data and retrieved images, along with corresponding labels. Here, the text in the image-text pair is the text feature from step 210, and the image in the image-text pair is the image feature obtained by encoding the retrieved image corresponding to the text feature in step 212 using the CLIP model. Thus, the classifier can be trained (step 214).

[0054] The following is combined with Figure 4 The application flow of the business data retrieval method according to the embodiments of this application is described. For example... Figure 4 This includes the database construction stage (steps 402 to 406) and the retrieval stage (steps 408 to 410).

[0055] Step 402: First, perform enhanced retrieval operations on the images and corresponding image description texts in the existing business dataset of the target business scenario. Step 404: Then, based on the CLIP model, each data after enhanced retrieval is vectorized to obtain the feature data represented by the corresponding vectors; Step 406: Use the vectors corresponding to the feature data as elements to create the preset database, and complete the database creation operation.

[0056] Step 408: The user only needs to provide a keyword or image. After corresponding retrieval enhancement, enhanced text training data or image training data is obtained. After being encoded by the CLIP model, similarity calculation is performed with the vectors in the database created in step 406. Step 410: Based on the similarity calculation results, retrieve the matching search results, such as the target image or text.

[0057] In this embodiment, by performing a retrieval enhancement operation on the target business data, the target business data is enhanced into training data of the corresponding data type in the training set of the CLIP model. The training set is used to train the CLIP model, and the target business data includes text-based retrieval data and / or image-based retrieval data. The enhanced target business data is encoded according to the CLIP model to obtain target feature data. Based on the target feature data, matching retrieval images are retrieved. This eliminates the need for a large amount of business data for model training. It allows for the enhancement of the business data to be retrieved without requiring additional business data, which not only improves the accuracy of subsequent retrieval tasks but also has strong generalization capabilities. This alleviates the problem that the retrieval results cannot meet business needs due to the inability or scarcity of actual business data.

[0058] To address the problems existing in the prior art, embodiments of this application also provide a business data auditing method, such as... Figure 5 As shown, the process includes steps 502 to 506.

[0059] Step 502: Perform retrieval enhancement operations on the target business data to enhance the target business data into training data of the corresponding data type in the training set of the CLIP model. The training set is used to train the CLIP model, and the target business data includes text-based retrieval data and image-based retrieval data.

[0060] Optionally, the target business data is subjected to retrieval enhancement operations, including: if the target business data includes text-based retrieval data, deriving synonyms for the text-based retrieval data; matching the derived synonyms with text-based training data in the training set of the CLIP model; and filtering text-based training data related to the text-based retrieval data from the matched text-based training data to obtain the retrieval-enhanced target business data.

[0061] Optionally, the retrieval enhancement operation on the target business data includes: when the target business data includes image-type retrieval data, segmenting the image-type retrieval data into multiple sub-image-type retrieval data; determining the image description text corresponding to each of the image-type retrieval data and the multiple sub-image-type retrieval data; grouping the image-text pairs composed of each image-type retrieval data and the corresponding image description text based on a preset classifier, wherein the image-text pairs in the same group describe the same object, and the classifier is trained using image-text pairs as samples and whether the image-text pairs describe the same object as labels; merging the image-type retrieval data in each image-text pair in the same group based on pixel position to obtain the retrieval-enhanced target business data.

[0062] Before grouping image-text pairs consisting of image retrieval data and corresponding image description text based on a preset classifier, the process further includes: acquiring multiple text-based business data from the business scenario to which the target business data belongs; performing retrieval enhancement operations on each text-based business data to enhance it into text-based training data in the training set of the CLIP model; encoding the enhanced text-based business data according to the CLIP model to obtain corresponding target feature data; retrieving matching images based on the target feature data; and training the classifier using image-text pairs consisting of text-based training data and retrieved images corresponding to each text-based business data as samples and whether the image-text pairs describe the same object as labels.

[0063] It is understood that step 502 in this embodiment is the same as or similar to step 102 in the above-described business data retrieval method embodiment. To avoid repetition, relevant descriptions are omitted appropriately.

[0064] Step 504: Encode the enhanced text retrieval data and enhanced image retrieval data according to the CLIP model to obtain the corresponding text feature data and image feature data.

[0065] It is understood that step 504 in this embodiment is the same as or similar to step 104 in the above embodiment of the business data retrieval method. To avoid repetition, relevant descriptions are omitted appropriately.

[0066] Step 506: Input the image-text pairs composed of text feature data and image feature data into a preset classifier to check whether the text data and the image data match. The classifier is trained using image-text pairs as samples and whether the image-text pairs describe the same object as labels.

[0067] Specifically, before inputting the image-text pairs composed of text feature data and image feature data into a preset classifier, the process further includes: acquiring multiple text-based business data of the business scenario to which the target business data belongs; performing retrieval enhancement operations on each text-based business data to enhance each text-based business data into text-based training data in the training set of the CLIP model; encoding the retrieval-enhanced text-based business data according to the CLIP model to obtain corresponding target feature data; retrieving matching retrieval images based on the target feature data; and training the classifier using the image-text pairs composed of the text-based training data corresponding to each text-based business data and the retrieval images as samples, and using whether the image-text pairs describe the same object as labels.

[0068] It is understood that the classifier training step in step 506 of this embodiment is the same as or similar to the classifier training step in the above-described business data retrieval method embodiment. To avoid repetition, relevant descriptions are omitted appropriately.

[0069] The following is combined with Figure 6 The application process of the business data verification method in the embodiments of this application is described. For example... Figure 6 This includes the following steps: Step 602: Perform retrieval enhancement processing on the target business data of the image-type retrieval data input by the user to obtain the image-type training data in the training set of the CLIP model. Step 604: Perform retrieval enhancement processing on the target business data of the text-based retrieval data input by the user to obtain the text-based training data in the training set of the CLIP model. Step 606: Based on the CLIP model, surface the enhanced text-type business data and the enhanced image-type business data respectively to obtain the corresponding feature data vectors; Step 608: Input the vectors of the enhanced text-based business data and the vectors of the enhanced image-based business data in pairs into the classifier trained on the text-based business data based on the corresponding business scenario, to determine whether the image-based retrieval data and the text-based retrieval data input by the user describe the same object. If they describe the same object, the two are considered to match; otherwise, they are considered not to match.

[0070] In this embodiment, target business data is enhanced by performing a retrieval enhancement operation to transform it into training data of the corresponding data type in the CLIP model's training set. The training set is used to train the CLIP model, and the target business data includes text-based retrieval data and image-based retrieval data. The enhanced text-based and image-based retrieval data are encoded according to the CLIP model to obtain corresponding text feature data and image feature data. Image-text pairs composed of the text feature data and image feature data are input into a preset classifier to verify whether the text-based and image-based retrieval data match. The classifier is trained using image-text pairs as samples and whether the image-text pairs describe the same object as labels. This eliminates the need for extensive business data for model training. By enhancing the business data to be reviewed without requiring additional business data, the accuracy of subsequent review tasks is improved, and strong generalization is achieved. This alleviates the problem of insufficient or unavailable actual business data leading to unsatisfactory review results.

[0071] Optionally, embodiments of this application also provide a business data retrieval device, including: An enhancement module is used to perform retrieval enhancement operations on target business data, so as to enhance the target business data into training data of the corresponding data type in the training set of the CLIP model. The training set is used to train the CLIP model. The target business data includes text-based retrieval data and / or image-based retrieval data. The encoding module is used to encode the enhanced target business data according to the CLIP model to obtain target feature data; The retrieval module is used to retrieve matching images based on the target feature data.

[0072] The business data retrieval device in this embodiment implements the following when executed: Figures 1 to 4 To avoid repetition, the steps of the business data retrieval method described herein have been appropriately omitted.

[0073] Optionally, embodiments of this application also provide a business data verification device, including: An enhancement module is used to perform retrieval enhancement operations on target business data, so as to enhance the target business data into training data of the corresponding data type in the training set of the CLIP model. The training set is used to train the CLIP model. The target business data includes text-based retrieval data and image-based retrieval data. The encoding module is used to encode the enhanced text retrieval data and the enhanced image retrieval data according to the CLIP model to obtain the corresponding text feature data and image feature data. The review module is used to input image-text pairs composed of text feature data and image feature data into a preset classifier to review whether the text-class retrieval data and the image-class retrieval data match. The classifier is trained using image-text pairs as samples and whether the image-text pairs describe the same object as labels.

[0074] The business data verification device in this embodiment implements the following when executed: Figures 5 to 6 To avoid repetition, the steps of the business data auditing method described herein have been appropriately omitted.

[0075] Optionally, such as Figure 7 As shown, this application embodiment also provides an electronic device 2000, including a processor 2400 and a memory 2200. The memory 2200 stores a program or instructions that can run on the processor 2400. When the program or instructions are executed by the processor 2400, they implement the various steps of the above-described business data retrieval method or business data auditing method embodiments and can achieve the same technical effect. To avoid repetition, they will not be described again here.

[0076] This application also provides a readable storage medium storing a program or instructions. When executed by a processor, the program or instructions implement the various processes of any of the above-described business data retrieval method or business data verification method embodiments, and achieve the same technical effect. To avoid repetition, further details are omitted here. The readable storage medium includes computer-readable storage media, such as read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0077] This application also provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program. The computer program is operable to enable a computer to execute various processes of any of the above-described business data retrieval method or business data auditing method embodiments, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0078] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0079] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0080] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A business data retrieval method, characterized in that, include: The target business data is enhanced by a retrieval enhancement operation to transform it into training data of the corresponding data type in the training set of the CLIP model. The training set is used to train the CLIP model, and the target business data includes text-based retrieval data and / or image-based retrieval data. The enhanced target business data is encoded according to the CLIP model to obtain target feature data; Retrieve matching images based on the target feature data.

2. The method according to claim 1, characterized in that, Perform retrieval enhancement operations on the target business data, including: When the target business data includes text-based retrieval data, deduce the synonyms of the text-based retrieval data; The derived synonyms are matched with the text training data in the training set of the CLIP model; The target business data is obtained by filtering the text class training data related to the text class retrieval data from the matched text class training data.

3. The method according to claim 1, characterized in that, Perform retrieval enhancement operations on the target business data, including: If the target business data includes image-based search data, the image-based search data is divided into multiple sub-image-based search data. Determine the image description text corresponding to each of the image category retrieval data and the multiple sub-image category retrieval data; Based on a preset classifier, the image-text pairs consisting of the image retrieval data of each image class and the corresponding image description text are grouped. The image-text pairs in the same group describe the same object. The classifier is trained using image-text pairs as samples and whether the image-text pairs describe the same object as labels. Image-type retrieval data from each image-text pair in the same group are merged based on pixel position to obtain the target business data with enhanced retrieval.

4. The method according to claim 3, characterized in that, Before grouping image-text pairs consisting of image retrieval data for each image category and corresponding image description text based on a preset classifier, the process also includes: Obtain multiple text-based business data sets belonging to the business scenario of the target business data; Retrieval enhancement operations are performed on each type of text business data to enhance each type of text business data into text training data in the training set of the CLIP model; The enhanced text-based business data is encoded according to the CLIP model to obtain the corresponding target feature data; Based on the target feature data, a matching image is retrieved. The classifier is trained using image-text pairs consisting of text training data corresponding to each text category of business data and retrieved images as samples, and using whether the image-text pairs describe the same object as labels.

5. The method according to any one of claims 1 to 4, characterized in that, Retrieving matching images based on the target feature data includes: The target feature data is compared with feature data in a preset database to determine the matching feature data; The image corresponding to the matched target feature data is determined as the matching retrieval image. The feature data in the preset database is obtained by encoding the target data using the CLIP model. The target data includes the data after retrieval and enhancement operations on the image-type business data of the business scenario to which the target business data belongs, and the data after retrieval and enhancement operations on the image description text corresponding to each image-type business data.

6. A method for auditing business data, characterized in that, include: The target business data is enhanced by a retrieval enhancement operation to transform it into training data of the corresponding data type in the training set of the CLIP model. The training set is used to train the CLIP model, and the target business data includes text-based retrieval data and image-based retrieval data. The enhanced text retrieval data and enhanced image retrieval data are encoded according to the CLIP model to obtain the corresponding text feature data and image feature data. Image-text pairs, composed of text feature data and image feature data, are input into a preset classifier to check whether the text-class retrieval data and the image-class retrieval data match. The classifier is trained using image-text pairs as samples and whether the image-text pairs describe the same object as labels.

7. The method according to claim 6, characterized in that, Before inputting the image-text pairs (composed of text feature data and image feature data) into the predefined classifier, the following steps are also included: Obtain multiple text-based business data sets belonging to the business scenario of the target business data; Retrieval enhancement operations are performed on each type of text business data to enhance each type of text business data into text training data in the training set of the CLIP model; The enhanced text-based business data is encoded according to the CLIP model to obtain the corresponding target feature data; Based on the target feature data, a matching image is retrieved. The classifier is trained using image-text pairs consisting of text training data corresponding to each text category of business data and retrieved images as samples, and using whether the image-text pairs describe the same object as labels.

8. A business data retrieval device, characterized in that, include: An enhancement module is used to perform retrieval enhancement operations on target business data, so as to enhance the target business data into training data of the corresponding data type in the training set of the CLIP model. The training set is used to train the CLIP model. The target business data includes text-based retrieval data and / or image-based retrieval data. The encoding module is used to encode the enhanced target business data according to the CLIP model to obtain target feature data; The retrieval module is used to retrieve matching images based on the target feature data.

9. A business data verification device, characterized in that, include: An enhancement module is used to perform retrieval enhancement operations on target business data, so as to enhance the target business data into training data of the corresponding data type in the training set of the CLIP model. The training set is used to train the CLIP model. The target business data includes text-based retrieval data and image-based retrieval data. The encoding module is used to encode the enhanced text retrieval data and the enhanced image retrieval data according to the CLIP model to obtain the corresponding text feature data and image feature data. The review module is used to input image-text pairs composed of text feature data and image feature data into a preset classifier to review whether the text-class retrieval data and the image-class retrieval data match. The classifier is trained using image-text pairs as samples and whether the image-text pairs describe the same object as labels.

10. An electronic device, characterized in that, It includes a processor and a memory, the memory storing a program or instructions that can run on the processor, the program or instructions being executed by the processor to implement the steps of the method as described in any one of claims 1-5 or claim 6 or 7.

11. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the method as described in any one of claims 1-5 or claim 6 or 7.

12. A computer program product, characterized in that, The computer program product includes a non-transitory computer-readable storage medium storing a computer program operable to cause a computer to perform the steps of the method as described in any one of claims 1-5 or claim 6 or 7.