Abnormality detection method and device, storage medium and electronic equipment

By acquiring credential images and their associated images, and utilizing large language models and rule-generated prompts, cross-sample anomaly detection is performed. This solves the problem of capturing complex relationships in images in traditional methods, achieving higher detection accuracy and rule generation efficiency.

CN120997843APending Publication Date: 2025-11-21XIAMEN UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510970076.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-15
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Traditional single-sample detection methods struggle to capture the complex relationships between document images, making it difficult to effectively detect anomalies such as document tampering and logical conflicts.

Method used

By acquiring the image to be detected and its associated images, and using the pre-trained large language model and the prompts generated by the stored rules, the model is guided to perform cross-sample anomaly detection, explore the relationships between images, and generate detection rules.

Benefits of technology

It improves the accuracy of anomaly detection, avoids biases caused by human subjectivity, reduces rule maintenance costs, and enhances the standardization and efficiency of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120997843A_ABST
    Figure CN120997843A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an anomaly detection method, and the method comprises the steps: inputting a to-be-detected image and all related images related to the to-be-detected image into a pre-trained large language model, through first prompt information generated based on stored rules, a large language model is guided to carry out cross-sample anomaly detection on a to-be-detected image based on each associated image and rule, so that the relationship between different images is mined through the reasoning ability of the large language model; and cross-sample anomaly detection is performed on the to-be-detected image based on the existing rule and the mined relationship, so that the anomaly detection accuracy can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present specification relates to the technical field of computer technology, and in particular, to an anomaly detection method and device, a storage medium, and an electronic device. BACKGROUND

[0002] With the digital transformation of the financial, tax, insurance and other industries, the electronic management of certificate data has become the mainstream. However, the number of abnormal situations such as certificate tampering and logical conflicts is increasing, and traditional single-sample detection methods cannot meet the complex business requirements.

[0003] Cross-sample certificate image anomaly detection is to analyze the logical association between multiple certificate images to identify potential abnormal patterns. Compared with traditional anomaly detection methods based on single-sample analysis, cross-sample anomaly detection can capture the complex relationships between multiple certificate images and perform anomaly detection on the certificate images accordingly.

[0004] Therefore, how to implement such a complex cross-sample anomaly detection is a problem to be solved. SUMMARY

[0005] Embodiments of the present specification provide an anomaly detection method, device, storage medium and electronic device to partially solve the problems existing in the prior art.

[0006] Embodiments of the present specification adopt the following technical solutions:

[0007] The anomaly detection method provided by the present specification comprises:

[0008] Obtaining a to-be-detected image;

[0009] According to the to-be-detected image, each associated image associated with the to-be-detected image is retrieved;

[0010] The to-be-detected image and each associated image are input into a pre-trained large language model, and a first prompt information generated based on a stored rule is input into the large language model; the rule is used for cross-sample anomaly detection of the to-be-detected image;

[0011] Obtaining a detection result of the large language model performing cross-sample anomaly detection on the to-be-detected image based on the associated images and the rule under the guidance of the first prompt information.

[0012] The anomaly detection device provided by the present specification comprises:

[0013] The obtaining module is configured to obtain a to-be-detected image;

[0014] retrieving, according to the image to be detected, each associated image associated with the image to be detected;

[0015] inputting, by an input module, the image to be detected and each associated image into a pre-trained large language model, and inputting first prompt information generated based on a stored rule into the large language model, the rule being used for cross-sample anomaly detection on the image to be detected;

[0016] detecting, by a detection module, a detection result of cross-sample anomaly detection on the image to be detected by the large language model based on each associated image and the rule under the guidance of the first prompt information.

[0017] The specification provides a computer-readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the anomaly detection method.

[0018] The specification provides an electronic device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the anomaly detection method when executing the program.

[0019] The above at least one technical solution adopted by the embodiment of the specification can achieve the following beneficial effects:

[0020] The embodiment of the specification discloses an anomaly detection method, which inputs an image to be detected and each associated image associated with the image to be detected into a pre-trained large language model, guides the large language model to perform cross-sample anomaly detection on the image to be detected based on each associated image and a rule through first prompt information generated based on a stored rule, thereby mining the relationship between different images through the reasoning capability of the large language model, and performing cross-sample anomaly detection on the image to be detected based on the existing rule and the mined relationship, which can improve the accuracy of anomaly detection. BRIEF DESCRIPTION OF DRAWINGS

[0021] The accompanying drawings described herein are used to provide further understanding of the specification, and form a part of the specification. The illustrative embodiments of the specification and their descriptions serve to explain the specification, and do not constitute an improper limitation on the specification. In the drawings:

[0022] Figure 1 An anomaly detection method flowchart provided by the embodiment of the specification;

[0023] Figure 2 An anomaly detection device schematic diagram provided by the embodiment of the specification;

[0024] Figure 3 A structure schematic diagram of an electronic device provided by the embodiment of the specification. DETAILED DESCRIPTION

[0025] For the purposes of the present description, the technical solutions and advantages thereof will be more apparent from the following detailed description of specific embodiments thereof, presented solely by way of non-limiting example, in conjunction with the accompanying drawings. Obviously, the described embodiments are only some of the embodiments of the present description and not all of them. Based on the embodiments described in the present description, all other embodiments obtained by a person of ordinary skill in the art without having to carry out an inventive act, fall within the scope of protection of the present description.

[0026] The technical solutions provided by the embodiments of the present description will be described in detail below in conjunction with the accompanying drawings.

[0027] Figure 1 A flowchart of an anomaly detection method provided by an embodiment of the present description includes the following steps:

[0028] S100: Obtain an image to be detected.

[0029] In the embodiments of the present description, the method for performing cross-sample anomaly detection on the image to be detected can be any electronic device, and the embodiments of the present description only take an electronic device as an example to illustrate a server. Figure 1

[0030] When performing cross-sample anomaly detection, the server first needs to obtain an image to be detected. The image to be detected described in the embodiments of the present description can be any image that needs to be detected for anomalies, and in particular can be a credential image.

[0031] A credential image refers to an image in electronic format that is a static or dynamic visual presentation form converted from a paper or entity credential (such as an invoice, a receipt, a contract, a check, an ID card, etc.) through digital photography, scanning or other image acquisition technology. Since users often need various credential images to pass the verification required by a business when performing a business, it is particularly important to verify whether the credential image is abnormal (such as whether it has been tampered with, whether it is counterfeit, etc.).

[0032] ​The traditional method of detecting anomalies in credential images through single-sample visual feature analysis generally needs to extract visual features (such as color, texture, shape, etc.) from a single credential image and detect anomalies through these features. This method is direct and intuitive, and is suitable for simple anomaly detection tasks, especially when the abnormal features are obvious. However, it only relies on the features of a single sample and is difficult to capture the complex relationships between samples. For example, when the credential image is an ID card image, if the ID card numbers in two ID card images are consistent, but the birth dates are inconsistent, or the faces on the ID cards are not the faces of the same person, then there is a high probability that at least one of the two ID card images is an abnormal ID card image, which is the so-called relationship between samples.

[0033] The server first obtains a credential image that needs to be detected for anomalies as a to-be-detected image. Of course, the to-be-detected image can also be other types of images that need to be detected for anomalies, and the embodiments of the present specification are only described by taking credential images as an example.

[0034] S102: According to the to-be-detected image, each associated image associated with the to-be-detected image is retrieved.

[0035] Since the to-be-detected image needs to be detected for anomalies in the embodiments of the present specification, after the server obtains the to-be-detected image, it also needs to retrieve each associated image associated with the to-be-detected image. Specifically, the server can retrieve each associated image associated with the to-be-detected image in a preset database. The preset database pre-stores a large number of credential images that have been authorized by users to be used by the server.

[0036] When retrieving the associated image from the preset database, the server can first perform visual encoding on the to-be-detected image to obtain the visual encoding features of the to-be-detected image, and then retrieve each associated image associated with the to-be-detected image according to the visual encoding features of the to-be-detected image. Wherein, when performing visual encoding on the to-be-detected image, any encoding method that can perform visual encoding on the image can be used, such as using a visual encoder ViT-B / 16 to perform visual encoding on the to-be-detected image, or using a dense text vector encoder RoBERTa-wwm-Base to perform visual encoding on the to-be-detected image, or using a sparse text vector encoder HashingVectorizer to perform visual encoding on the to-be-detected image, all of which can achieve the technical effects of the embodiments of the present specification.

[0037] It should be noted that the images stored in the above preset database also need to be visually encoded in the embodiments of the present specification (the encoding method is the same as that of visually encoding the to-be-detected image), and the visual encoding features of each image are stored as the index of each image. Therefore, after obtaining the visual encoding features of the to-be-detected image, the server can determine the associated images associated with the to-be-detected image according to the similarity between the visual encoding features of the to-be-detected image and the visual encoding features of each image stored in the preset database. For example, in order of similarity from high to low, n images stored in the preset database are selected as the associated images associated with the to-be-detected image, and n is a preset number.

[0038] S104: input the to-be-detected image and each associated image into a pre-trained large language model, and input the first prompt information generated based on the stored rule into the large language model.

[0039] The rule is used for cross-sample anomaly detection on the to-be-detected image.

[0040] The rule used for cross-sample anomaly detection in the present specification can be stored in a cross-sample anomaly detection rule library, so that the server can first construct the first prompt information according to the rule stored in the rule library, and then input the first prompt information into the large language model. The first prompt information is used to guide the large language model to perform cross-sample anomaly detection on the to-be-detected image based on each associated image and the rule in the above rule library.

[0041] In order to avoid introducing human subjective factors, the rule in the above rule library can also be generated by a pre-trained large language model in the embodiments of the present specification. That is, in the embodiments of the present specification, the server can perform cross-sample anomaly detection on the to-be-detected image by the large language model and the existing rule on the one hand, and use the reasoning ability of the large language model itself to generate a rule for cross-sample anomaly detection based on the to-be-detected image and the associated images retrieved in step S102 on the other hand. Therefore, before inputting the above first prompt information into the large language model, the server also needs to input second prompt information into the large language model, so as to guide the large language model to mine the relationship between the to-be-detected image and each associated image based on its own reasoning ability, and generate a rule for cross-sample anomaly detection on the to-be-detected image according to the relationship, and then add the generated rule to the cross-sample anomaly detection rule library for storage.

[0042] Specifically, since the to-be-detected image and each of the associated images associated with the to-be-detected image are all credential images, tampering and forgery of the credential images are mostly around key fields in the credential images, for example, if the credential image is an ID card image, the tampering and forgery are generally the name, date of birth, address, and ID card number in the ID card image, and if the credential image is a tax payment certificate image, the tampering and forgery are generally the taxpayer name, taxpayer identification number, tax payment date, and certificate number in the tax payment certificate image. Therefore, in the embodiments of the present specification, before inputting the second prompt information into the large language model, the server can first input third prompt information into the large language model, the third prompt information being used to guide the large language model to identify the text content in the to-be-detected image and each of the associated images and extract the key fields from the text content. The large language model can first identify the text content in the to-be-detected image and each of the associated images under the guidance of the third prompt information, and then extract the key fields from the text content based on its reasoning capability.

[0043] After obtaining the key fields extracted from the text content by the large language model, the second prompt information used to guide the large language model to generate the rules for cross-sample anomaly detection of the to-be-detected image based on the key fields can be constructed based on the key fields extracted by the large language model, and then the second prompt information is input into the large language model, so that the large language model generates the rules for cross-sample anomaly detection of the to-be-detected image based on the key fields, and finally the generated rules are added to the cross-sample anomaly detection rule library for storage.

[0044] During initialization, if the rule library does not store any rule for cross-sample anomaly detection, one or more rules for cross-sample anomaly detection set by humans can be added to the rule library for storage.

[0045] S106: Obtain the detection result of the cross-sample anomaly detection of the to-be-detected image by the large language model based on the associated images and the rules under the guidance of the first prompt information.

[0046] After inputting the to-be-detected image, each of the associated images, and the first prompt information into the large language model through the above step S104, the server can obtain the relationship between the to-be-detected image and each of the associated images mined by the large language model based on its reasoning capability, and perform cross-sample anomaly detection on the to-be-detected image based on this relationship and the rules in the rule library.

[0047] By the above method, the to-be-detected image and each associated image associated with the to-be-detected image are input into the pre-trained large language model, the first prompt information generated based on the stored rules is used to guide the large language model to perform cross-sample anomaly detection on the to-be-detected image based on each associated image and the rules, so as to mine the relationship between different images through the reasoning ability of the large language model, and perform cross-sample anomaly detection on the to-be-detected image based on the existing rules and the mined relationship, which can improve the accuracy of anomaly detection. At the same time, the large language model can also be used to generate rules for cross-sample anomaly detection, so that rules for cross-sample anomaly detection based on the relationship between images can be generated, and rules can be generated without relying on artificial generation, avoiding the deviation caused by the introduction of human subjective factors, improving the accuracy of subsequent detection, and efficiently generating rules, which not only facilitates the standardization of rules, but also reduces the cost of maintaining rules.

[0048] Further, in the above step S102, each associated image associated with the to-be-detected image is retrieved in the preset database. The to-be-detected image can be encoded by using two or more visual coding methods to obtain visual coding features of the to-be-detected image corresponding to different visual coding methods, and the visual coding features are integrated to retrieve each associated image associated with the to-be-detected image in the preset database, which can improve the accuracy of retrieving associated images and further improve the accuracy of cross-sample anomaly detection.

[0049] Specifically, the server can use each preset visual coding method to visually encode the to-be-detected image to obtain visual coding features of the to-be-detected image corresponding to different visual coding methods, and then for each preset visual coding method, based on the similarity between the visual coding features of each candidate image (i.e., each image already stored in the preset database) corresponding to the visual coding method and the visual coding features of the to-be-detected image corresponding to the visual coding method, select a to-be-determined image from the candidate images. Finally, based on the to-be-determined images selected for each preset visual coding method, a to-be-determined image set is determined, and each associated image associated with the to-be-detected image is determined in the to-be-determined image set.

[0050] The preset visual encoding methods can include: using a visual encoder ViT-B / 16 to visually encode the to-be-detected image, using a dense text vector encoder RoBERTa-wwm-Base to visually encode the to-be-detected image, and using a sparse text vector encoder HashingVectorizer to visually encode the to-be-detected image. Through the above three visual encoding methods, the server can obtain three kinds of visual encoding features of the to-be-detected image. Through the three kinds of visual encoding features, the server can retrieve three groups of to-be-determined images. After merging the three groups of to-be-determined images to obtain a to-be-determined image set, the server can determine each associated image associated with the to-be-detected image from the to-be-determined image set.

[0051] When determining each associated image associated with the to-be-detected image from the to-be-determined image set, the server can determine, for each to-be-determined image in the to-be-determined image set, a degree of association between the to-be-determined image and the to-be-detected image according to similarities between visual encoding features of the to-be-determined image corresponding to each preset visual encoding method and visual encoding features of the to-be-detected image corresponding to each preset visual encoding method. Then, the server can determine each associated image associated with the to-be-detected image from the to-be-determined image set according to the degrees of association between each to-be-determined image in the to-be-determined image set and the to-be-detected image. Specifically, for a to-be-determined image, a similarity Si between visual encoding features of the to-be-determined image corresponding to an i-th preset visual encoding method and visual encoding features of the to-be-detected image corresponding to the i-th preset visual encoding method can be determined. Then, similarities S i The weighted similarity ∑ i w i S i , where w i is a weight of the i-th preset visual encoding method. The weighted similarity ∑ i w i S i is the degree of association between the to-be-determined image and the to-be-detected image. After obtaining the degrees of association between each to-be-determined image in the to-be-determined image set and the to-be-detected image, n to-be-determined images can be selected from the to-be-determined image set in descending order of the degrees of association, as each associated image associated with the to-be-detected image.

[0052] In addition, the rule for cross-sample anomaly detection in the embodiment of the present specification described in detail in the above step S104 can also be generated by the large language model, and in order to ensure the rationality of the generated rule and improve the accuracy of subsequent cross-sample anomaly detection, after obtaining the cross-sample anomaly detection rule generated by the large language model, the rule can also be verified by fact-driven, and after verification, the rule is added to the cross-sample anomaly detection rule library for storage. Specifically, the server can first retrieve each image associated with the to-be-detected image from the pre-determined fact library as a sample image according to the to-be-detected image, and then perform cross-sample anomaly detection on each sample image according to the rule generated by the large language model in step S104 and each sample image to obtain a detection result, and finally verify the rule generated by the large language model in step S104 according to the detection result and the labeled detection result corresponding to the sample image. When the rule generated by the large language model passes the verification, the rule generated by the large language model is added to the cross-sample anomaly detection rule library for storage.

[0053] The pre-determined fact library can only include positive samples (i.e. images without anomalies), can only include negative samples (i.e. images with anomalies), or can include both positive samples and negative samples. Since in actual application scenarios, images with anomalies are indeed a minority, in order to facilitate the verification of the rule generated by the large language model, the fact library in the embodiment of the present specification can only include positive samples.

[0054] When retrieving the comparison image from the fact library, the same method as step S102 shown in FIG. 1 can be used for retrieval, which will not be described here. Figure 1

[0055] The detection result obtained by performing cross-sample anomaly detection on the sample image according to the rule generated in step S104 and each sample image includes two results: existence of anomaly and non-existence of anomaly, and can also include the reason for existence of anomaly and the reason for non-existence of anomaly.

[0056] If the detection result output by the large language model is consistent with the labeled detection result corresponding to the sample image, it is determined that the rule generated in step S104 passes the verification, and if it is not consistent, it is determined that the rule generated in step S104 does not pass the verification. Specifically, if the detection result output by the large language model for each sample image is consistent with the labeled detection result corresponding to each sample image, it is determined that the rule generated in step S104 passes the verification, and if the detection result output for at least one sample image is inconsistent with the labeled detection result corresponding to the sample image, it is determined that the rule generated in step S104 does not pass the verification.

[0057] ​For the case where the rule passes the fact-driven verification described above, the rule generated by the explanation step S104 has a high degree of credibility, and the rule can be used for subsequent anomaly detection, therefore, the server can add the rule to the cross-sample anomaly detection rule library for storage, and when a to-be-detected image is obtained, perform cross-sample anomaly detection on the to-be-detected image according to the rules in the cross-sample anomaly detection rule library. Of course, before adding the rule to the rule library, it can also be determined whether the rule is repeated with the rules already stored in the rule library. If it is repeated, there is no need to add the rule to the rule library. If it is not repeated, the rule is added to the rule library for storage.

[0058] For the case where the rule does not pass the fact-driven verification described above, the rule generated by the explanation step S104 has a low degree of credibility, and the rule cannot be used for subsequent anomaly detection, therefore, the server can input the fourth prompt information generated based on the rule into the large language model, and make the large language model adjust the rule based on the to-be-detected image and the associated images retrieved by the retrieval step S102 under the guidance of the fourth prompt information. Alternatively, the rule can also be discarded directly, and a cross-sample anomaly detection rule is re-generated by the generation step S104 until a rule that passes the verification is generated.

[0059] The above is an anomaly detection method provided by an embodiment of the present specification, based on the same idea, the present specification also provides a corresponding device, a storage medium and an electronic device.

[0060] Figure 2 An anomaly detection device provided by an embodiment of the present specification is shown in the schematic diagram, and the device comprises:

[0061] The acquisition module 201 is configured to acquire a to-be-detected image.

[0062] The retrieval module 202 is configured to retrieve, according to the to-be-detected image, each associated image associated with the to-be-detected image.

[0063] The input module 203 is configured to input the to-be-detected image and each associated image into a pre-trained large language model, and input first prompt information generated based on a stored rule into the large language model; the rule is used for cross-sample anomaly detection on the to-be-detected image.

[0064] The detection module 204 is configured to acquire a detection result of cross-sample anomaly detection on the to-be-detected image by the large language model based on each associated image and the rule under the guidance of the first prompt information.

[0065] Optionally, the searching module 202 is specifically configured to perform visual coding on the to-be-detected image to obtain visual coding features of the to-be-detected image; and search for the associated images associated with the to-be-detected image according to the visual coding features of the to-be-detected image.

[0066] Optionally, the searching module 202 is specifically configured to perform visual coding on the to-be-detected image to obtain visual coding features of the to-be-detected image; and search for the associated images associated with the to-be-detected image according to the visual coding features of the to-be-detected image.

[0067] Optionally, the searching module 202 is specifically configured to, for each of the to-be-detected images in the to-be-detected image set, determine an association degree between the to-be-detected image and the to-be-detected image according to the similarity between the visual coding features of the to-be-detected image corresponding to each of the preset visual coding methods and the visual coding features of the to-be-detected image corresponding to each of the preset visual coding methods; and determine the associated images associated with the to-be-detected image from the to-be-detected image set according to the association degree between each of the to-be-detected images in the to-be-detected image set and the to-be-detected image.

[0068] Optionally, the device further comprises:

[0069] The rule generation module 205 is configured to input second prompt information into the large language model before the input module 203 inputs the first prompt information into the large language model; obtain a rule generated by the large language model for cross-sample anomaly detection of the to-be-detected image based on the to-be-detected image and the associated images under the guidance of the second prompt information, and store the rule.

[0070] Optionally, the rule generation module 205 is further configured to input third prompt information into the large language model before inputting the second prompt information into the large language model; and obtain text content identified from the to-be-detected image and the associated images by the large language model under the guidance of the third prompt information, and a key field extracted from the text content.

[0071] The rule generation module 205 is specifically configured to input the second prompt information generated based on the key field into the large language model, and the second prompt information is used to guide the large language model to generate a rule for cross-sample anomaly detection of the to-be-detected image based on the key field.

[0072] Optionally, the rule generation module 205 is specifically configured to retrieve, from a pre-determined fact library, images associated with the to-be-detected image as sample images according to the to-be-detected image; perform cross-sample anomaly detection on each of the sample images according to the rule generated by the large language model and each of the sample images to obtain a detection result; verify the rule generated by the large language model according to the detection result and a labeled detection result corresponding to the sample image; and store the rule generated by the large language model in a cross-sample anomaly detection rule library when the rule generated by the large language model passes the verification.

[0073] Optionally, the images contained in the fact library are images without anomalies.

[0074] Optionally, the rule generation module 205 is specifically configured to determine that the rule generated by the large language model passes the verification if the detection result is consistent with the labeled detection result corresponding to the sample image, and otherwise determine that the rule generated by the large language model does not pass the verification.

[0075] Optionally, the rule generation module 205 is further configured to input fourth prompt information to the large language model when it is determined that the rule generated by the large language model does not pass the verification, so that the large language model adjusts the rule generated by the large language model based on the to-be-detected image and the associated image under the guidance of the fourth prompt information.

[0076] The specification also provides a computer-readable storage medium, which stores a computer program. The computer program is executed by a processor to implement the anomaly detection method provided above.

[0077] Based on the anomaly detection method shown in the specification, the embodiments of the specification also provide an electronic device as shown in the structure diagram. As shown in the structure diagram, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Figure 1 Figure 3 The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs to implement the anomaly detection method described above. Figure 3 At the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory, and of course can also include other hardware required by the business. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs to implement the anomaly detection method described above.

[0078] The above only describes the embodiments of the specification and is not intended to limit the specification. The specification can have various changes and modifications for those skilled in the art. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the specification shall be included in the scope of the claims of the specification.​

Claims

1. An anomaly detection method, comprising: obtaining a to-be-detected image; retrieving each associated image associated with the to-be-detected image according to the to-be-detected image; inputting the to-be-detected image and each associated image into a pre-trained large language model, and inputting first prompt information generated based on a stored rule into the large language model, the rule being used for cross-sample anomaly detection on the to-be-detected image; obtaining a detection result of cross-sample anomaly detection on the to-be-detected image by the large language model based on each associated image and the rule under the guidance of the first prompt information.

2. The method of claim 1, wherein retrieving each associated image associated with the to-be-detected image according to the to-be-detected image specifically comprises: performing visual coding on the to-be-detected image to obtain visual coding features of the to-be-detected image; retrieving each associated image associated with the to-be-detected image according to the visual coding features of the to-be-detected image.

3. The method of claim 2, wherein performing visual coding on the to-be-detected image to obtain visual coding features of the to-be-detected image specifically comprises: performing visual coding on the to-be-detected image using each pre-set visual coding method to obtain visual coding features of the to-be-detected image corresponding to different visual coding methods; retrieving each associated image associated with the to-be-detected image according to the visual coding features of the to-be-detected image specifically comprises: for each pre-set visual coding method, selecting a to-be-determined image from each candidate image according to the similarity between the visual coding features of each candidate image corresponding to the visual coding method and the visual coding features of the to-be-detected image corresponding to the visual coding method; determining a to-be-determined image set according to the to-be-determined images selected for each pre-set visual coding method, and determining each associated image associated with the to-be-detected image in the to-be-determined image set.

4. The method of claim 3, wherein determining each associated image associated with the to-be-detected image in the to-be-determined image set specifically comprises: for each to-be-determined image in the to-be-determined image set, determining the association degree between the to-be-determined image and the to-be-detected image according to the similarity between the visual coding features of the to-be-determined image corresponding to each pre-set visual coding method and the visual coding features of the to-be-detected image corresponding to each pre-set visual coding method; determining each associated image associated with the to-be-detected image in the to-be-determined image set according to the association degree between each to-be-determined image in the to-be-determined image set and the to-be-detected image.

5. The method of claim 1, wherein before inputting the first prompt information generated based on the stored rule into the large language model, the method further comprises: inputting second prompt information into the large language model; obtaining a rule generated by the large language model based on the to-be-detected image and each associated image for cross-sample anomaly detection on the to-be-detected image under the guidance of the second prompt information, and storing.

6. The method of claim 5, wherein before inputting the second prompt information into the large language model, the method further comprises: inputting third prompt information into the large language model; obtaining text content recognized by the large language model from the to-be-detected image and the associated images under the guidance of the third prompt information, and a key field extracted from the text content; inputting second prompt information into the large language model, specifically including: inputting second prompt information generated based on the key field into the large language model, the second prompt information being used to guide the large language model to generate a rule for cross-sample anomaly detection of the to-be-detected image based on the key field.

7. The method of claim 5, storing the rule generated by the large language model for cross-sample anomaly detection of the to-be-detected image, specifically including: retrieving, from a pre-determined fact library, each image associated with the to-be-detected image as a sample image according to the to-be-detected image; performing cross-sample anomaly detection on each sample image according to the rule generated by the large language model and each sample image to obtain a detection result; verifying the rule generated by the large language model according to the detection result and a labeled detection result corresponding to the sample image; storing the rule generated by the large language model in a cross-sample anomaly detection rule library when the rule generated by the large language model passes the verification.

8. The method of claim 7, wherein the images contained in the fact library are images without anomalies.

9. The method of claim 7, verifying the rule generated by the large language model according to the detection result and the labeled detection result corresponding to the sample image, specifically including: if the detection result is consistent with the labeled detection result corresponding to the sample image, determining that the rule generated by the large language model passes the verification, otherwise determining that the rule generated by the large language model does not pass the verification.

10. The method of claim 7, when it is determined that the rule generated by the large language model does not pass the verification, the method further includes: inputting fourth prompt information into the large language model to enable the large language model to adjust the rule generated by the large language model based on the to-be-detected image and the associated images under the guidance of the fourth prompt information.

11. An anomaly detection device, the device comprising: an acquisition module configured to acquire a to-be-detected image; a retrieval module configured to retrieve each associated image associated with the to-be-detected image according to the to-be-detected image; an input module configured to input the to-be-detected image and each associated image into a pre-trained large language model, and input first prompt information generated based on a stored rule into the large language model, the rule being used for cross-sample anomaly detection of the to-be-detected image; a detection module configured to obtain a detection result of cross-sample anomaly detection of the to-be-detected image by the large language model based on each associated image and the rule under the guidance of the first prompt information.

12. A computer-readable storage medium, the storage medium storing a computer program, the computer program being executed by a processor to implement the method of any one of claims 1-10.

13. An electronic device comprising a memory, a processor, and a computer program stored on the memory and loadable on the processor, the processor implementing the method of any of claims 1-10 when executing the program.