Abnormality detection method, device, electronic device and storage medium
The target detection model is used to preliminarily determine abnormal areas and categories, and combined with the large model to generate abnormal detection results, which solves the accuracy and real-time problems of information review on the Internet platform and realizes efficient and accurate information review.
Patent Information
- Application Number
- CN202411578450.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-06
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2044-11-06
AI Technical Summary
Existing technologies for reviewing user-uploaded information on Internet platforms have problems with low accuracy and low real-time performance, making it difficult to effectively detect inappropriate content.
The abnormal areas and categories are preliminarily determined through the target detection model, and then the large model is used in combination with the target reference information to generate the abnormal detection results, reducing the dependence on the large model and improving the detection accuracy and real-time performance.
It improves the accuracy and real-time performance of anomaly detection, reduces the false alarm rate, and ensures the efficiency and accuracy of information review.
Smart Images

Figure CN119399447B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to the fields of large models, generative models, multimodal models, text processing, image processing, etc. More specifically, the present disclosure provides an anomaly detection method, device, electronic device, storage medium, and computer program product. Background Art
[0002] Some Internet platforms need to review the information uploaded by users and take some countermeasures if the information contains inappropriate content to ensure the security of platform information. Summary of the Invention
[0003] The present disclosure provides an anomaly detection method, apparatus, electronic device, storage medium, and computer program product.
[0004] According to one aspect of the present disclosure, a method for anomaly detection is provided, comprising: in response to obtaining input information, performing target detection on an input image in the input information to obtain the position and anomaly category of an abnormal area in the input image; and generating an anomaly detection result for the input information based on a large model according to input text, input image, the position, anomaly category and target reference information in the input information.
[0005] According to another aspect of the present disclosure, an anomaly detection device is provided, comprising: an object detection module and a generation module. The object detection module is configured to, in response to receiving input information, perform object detection on an input image in the input information to obtain the location and anomaly category of an abnormal region in the input image. The generation module is configured to generate an anomaly detection result for the input information based on a large model, according to the input text, input image, location of the abnormal region, anomaly category, and target reference information in the input information.
[0006] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method provided by the present disclosure.
[0007] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause a computer to execute the method provided by the present disclosure.
[0008] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, which implements the method provided in the present disclosure when executed by a processor.
[0009] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.
[0011] Figure 1 is a schematic diagram of an application scenario of the anomaly detection method and device according to an embodiment of the present disclosure;
[0012] Figure 2 is a schematic flow chart of an anomaly detection method according to an embodiment of the present disclosure;
[0013] Figure 3 is a schematic flow chart of an anomaly detection method according to another embodiment of the present disclosure;
[0014] Figure 4 is a schematic diagram of an anomaly detection method according to an embodiment of the present disclosure;
[0015] Figure 5 is a schematic structural block diagram of an abnormality detection device according to an embodiment of the present disclosure; and
[0016] Figure 6 It is a structural block diagram of an electronic device used to implement the abnormality detection method of the embodiment of the present disclosure. DETAILED DESCRIPTION
[0017] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0018] In the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0019] In the technical solution disclosed herein, the user's authorization or consent is obtained before obtaining or collecting the user's personal information.
[0020] Some Internet platforms need to review information uploaded by users. For example, when users upload text or images, they need to review them to confirm whether the information contains violent, sensitive or other illegal and negative information.
[0021] In some related technologies, for example, specific categories of bad information can be detected through preset rules or models. However, rule-making is difficult to cope with the diversity of content and there is a problem of high false alarm rate. For another example, deep neural networks can be used for image classification or target detection, but the accuracy of this detection method is relatively low, and it is easy to intercept normal content, affecting the user experience. For another example, a preliminary review can be conducted through a deep neural network, and then manual review can be used to enhance the user experience, but the real-time performance of this method is low, and it is difficult to meet the security and real-time requirements in large model scenarios. In summary, the review method provided by the relevant technology has the problems of low accuracy and low real-time performance.
[0022] The present disclosure provides a method for detecting abnormalities. The method first obtains input information, including an input image and input text. Object detection is then performed on the input image to preliminarily determine the location and category of abnormal regions. Reference information for objects in the same domain as the input text is then retrieved. A large model is then used to process the input text, the input image, the location of the abnormal region, the category of the abnormality, and the reference information to generate an abnormality detection result.
[0023] As can be seen, the above solution can improve the accuracy of anomaly detection results by auditing both target detection and large models. In addition, since automatic detection can be performed without manual audit, the real-time audit can be ensured.
[0024] The technical solutions provided by the present disclosure will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0025] Figure 1 Schematic diagram of an application scenario of the anomaly detection method and device according to an embodiment of the present disclosure.
[0026] It should be noted that Figure 1 The examples shown are merely examples of system architectures to which the embodiments of the present disclosure may be applied, to help those skilled in the art understand the technical content of the present disclosure, but do not mean that the embodiments of the present disclosure may not be used in other devices, systems, environments or scenarios.
[0027] like Figure 1 As shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used as a medium for providing communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.
[0028] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Terminal devices 101, 102, and 103 can be various electronic devices with display screens and support web browsing, including but not limited to smartphones, tablet computers, laptop computers, and desktop computers, etc.
[0029] Server 105 may be a server that provides various services, such as a backend management server (for example only) that supports websites browsed by users using terminal devices 101, 102, and 103. The backend management server may analyze and process received data such as user requests, and provide feedback to the terminal device regarding the processing results (e.g., determining anomaly detection results based on user input and generating prompt information as a processing result when an anomaly is determined; the anomaly detection results and prompt information may be used as processing results).
[0030] It should be noted that the anomaly detection method provided in the embodiments of the present disclosure can generally be executed by the server 105. Accordingly, the anomaly detection device provided in the embodiments of the present disclosure can generally be set in the server 105. The anomaly detection method provided in the embodiments of the present disclosure can also be executed by a server or server cluster that is different from the server 105 and can communicate with the terminal devices 101, 102, 103 and / or the server 105. Accordingly, the anomaly detection device provided in the embodiments of the present disclosure can also be set in a server or server cluster that is different from the server 105 and can communicate with the terminal devices 101, 102, 103 and / or the server 105.
[0031] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.
[0032] Figure 2 is a schematic flowchart of an anomaly detection method according to an embodiment of the present disclosure.
[0033] like Figure 2 As shown, the abnormality detection method 200 may include operations S210 to S220.
[0034] In operation S210 , in response to obtaining input information, target detection is performed on an input image in the input information to obtain a position and an abnormality category of an abnormal region in the input image.
[0035] In one application scenario, the input information can be entered by the user through the front-end page. The input text can be user-written content such as stories and jokes, or it can also be questions raised by the user. In this application scenario, the anomaly detection method in this example can be used to detect whether the user-uploaded information is abnormal.
[0036] In another application scenario, the input information may be generated by a generative model. For example, the generative model generates graphic and text information, and the graphic and text information may be tested to determine whether the information generated by the generative model is abnormal.
[0037] For example, the input information includes input text and an input image. Object detection can be performed on the input image using YOLO (You Only Look Once, a target detection model) or other algorithms to determine whether the input image contains abnormal regions. If abnormal regions exist, the location and category of the abnormal regions can be determined. These abnormal regions may contain illegal content, and examples of abnormal categories include infringement of the legitimate rights and interests of others and inappropriate language. For example, for example, an input image may include a flag containing inappropriate language. Object detection can be used to determine the flag's location in the input image and to determine that the flag's abnormal category is "inappropriate language."
[0038] In operation S220 , an abnormality detection result for the input information is generated based on the large model according to the input text, the input image, the location of the abnormal area, the abnormality category, and the target reference information in the input information.
[0039] For example, the large model may be a large language model, a multimodal large model, etc. This embodiment does not limit the large model.
[0040] For example, the input text, the location of the abnormal area, the abnormality category, the target reference information, and the prompt information can be combined. The combined text and the input image are then input into the large model, which generates the abnormality detection results. For example, the prompt information is natural language text, and the prompt information may be, for example, "The area from the upper left corner coordinate (X1, Y1) to the lower right corner coordinate (X2, Y2) in the image is suspected to have a violation. The violation category is XX. Please determine whether it is a violation and explain the reason." In the prompt information, the violation category is abnormal and has not changed. The coordinates (X1, Y1) and (X2, Y2) are determined based on the location of the abnormal area.
[0041] The disclosed embodiment uses object detection to initially determine whether an input image contains an anomaly. If an anomaly is determined, the system then uses the large model to conduct a secondary review to determine whether the input information is abnormal. This improves the accuracy of anomaly detection results. Furthermore, since automatic detection eliminates the need for manual review, real-time review is ensured.
[0042] Figure 3 is a schematic flowchart of an anomaly detection method according to another embodiment of the present disclosure.
[0043] like Figure 3 As shown, the abnormality detection method 300 may include operation S310, operation S315~operation S316 and operation S320.
[0044] In operation S310 , in response to obtaining input information, target detection is performed on an input image in the input information to obtain a position and an abnormality category of an abnormal region in the input image.
[0045] In operation S315, based on the abnormal area, it is determined whether there is target data matching the abnormal area in the first database, wherein the data stored in the first database is determined based on the abnormal information. If yes, operation S316 is executed, otherwise, operation S320 is executed.
[0046] For example, a first database can be pre-built based on abnormal information. The abnormal information can be data in a modality such as text, images, or videos, and can contain abnormal content such as offensive language. The abnormal information can be added to the first database, or features can be extracted from the abnormal information and then added to the first database.
[0047] For example, a sub-image of the abnormal region can be cropped from the input image based on the location of the abnormal region. Features of the sub-image can then be extracted. The similarity between the sub-image features and the data in the first database can also be determined, and whether the similarity is greater than or equal to a similarity threshold can be determined. If so, the data can be determined to match the abnormal region and be identified as target data. If not, the data can be determined to not match the abnormal region and be identified as not target data.
[0048] For another example, a pre-trained classification model can be used to detect and record the category of abnormal information. If the similarity between a certain data in the first database and the abnormal region in the input image is greater than a threshold, and the category of the data is consistent with the abnormal category of the input image, the data is determined to match the abnormal region and is identified as the target data. If not, the data is determined not to match the abnormal region and is not the target data.
[0049] In operation S316 , determining the abnormality detection result includes: the input information is abnormal. In addition, the abnormality detection result may also include an abnormality category of the input information, and the abnormality category of the input information is consistent with the category of the matched target data.
[0050] In operation S320 , an abnormality detection result for the input information is generated based on the large model according to the input text, the input image, the location of the abnormal area, the abnormality category, and the target reference information in the input information.
[0051] In this embodiment, the first database is first searched to see if there is target data matching the abnormal region. If so, the input information is directly determined to be abnormal. If not, a second determination is made using the large model. This avoids requiring all input information to be tested against the large model, reducing the number of large model calls and the cost of using the large model.
[0052] According to another embodiment of the present disclosure, performing object detection on the input image in the input information may include determining a location and a class confidence of a region of interest in the input image. Furthermore, determining a target threshold based on a detection scenario. Then, in response to detecting that the class confidence is greater than or equal to the target threshold, determining the region of interest as an abnormal region.
[0053] For example, the location and confidence category of the region of interest can be determined by a target detection model. The target detection model can be YOLO or other models, such as PPYOLOE or other lightweight models suitable for rapid object detection in a real-time environment. This embodiment does not limit the target detection model. The target detection model needs to be trained in advance. During the training process, self-supervised training can be performed using a large-scale first data set, and supervised training can also be performed using a second data set. The data in the second data set may include data generated based on artificial intelligence, so that the target detection model can identify whether the information generated by artificial intelligence is abnormal.
[0054] For example, the correspondence between the detection scene and the target threshold may be pre-configured, and then the target threshold required for the current scene may be selected based on the correspondence, and then whether the region of interest is an abnormal region may be determined according to the target threshold.
[0055] This embodiment automatically adjusts the target threshold according to the detection scenario. Different detection scenarios may have different requirements for target detection. By setting an appropriate target threshold, it is possible to ensure that the target detection results meet the requirements of the scenario, thereby improving the accuracy of the detection results.
[0056] In the actual detection process, the target threshold for target detection can be set relatively low to avoid missed detections. At the same time, the input information of false detection can be re-detected through the large model to ensure the accuracy of detection.
[0057] According to another embodiment of the present disclosure, the input of the large model may further include target reference information. Next, the process of determining the target reference information is described.
[0058] In this embodiment, candidate reference information can be retrieved from the second database based on the input text based on predetermined search rules. The input text and the candidate reference information are then classified to obtain the fields covered by the input text and the fields covered by the candidate reference information. If the fields covered by the candidate reference information are the same as the fields covered by the input text, the candidate reference information is determined as the target reference information.
[0059] For example, a second database can be pre-built, and the reference information stored in the second database can include reference information with high confidence in multiple fields. For example, the confidence of the reference information is greater than a threshold, indicating that the reference information is basically correct information based on facts rather than erroneous information.
[0060] For example, the predetermined search rule is pre-configured according to actual needs, and may be, for example, a search based on text similarity, a search based on semantic similarity, etc., which is not limited in this embodiment.
[0061] For example, it can be determined whether the input text meets the conversion conditions. If the input text meets the conversion conditions, the input text can be subjected to text conversion processing to obtain the converted text. In one example, the conversion conditions include, for example: the number of characters in the input text is greater than or equal to a first quantity threshold. Accordingly, the text conversion processing may include extracting summaries, extracting keywords, etc. In one example, the conversion conditions include, for example: the number of characters in the input text is less than or equal to a second quantity threshold, and the second quantity threshold is less than the first quantity threshold. Accordingly, the text conversion processing may include expanding the input text. If the number of characters in the input text is between the first quantity threshold and the second quantity threshold, the input text may not be subjected to text conversion.
[0062] Retrieval may be performed based on the converted text, for example, by calculating the similarity between the converted text and each reference information in the second database, and determining the reference information as candidate reference information when the similarity is greater than a threshold.
[0063] Next, a pre-trained classification model can be used to detect the fields involved in the input text and the fields involved in the candidate reference information. If the fields involved in the candidate reference information are the same as the fields involved in the input text, the candidate reference information is determined as the target reference information; otherwise, the candidate reference information is not determined as the target reference information.
[0064] It should be noted that, in other embodiments, the input text may not be converted, but the input text before conversion may be directly used to search in the second database.
[0065] In this embodiment, candidate reference information related to the input text is first retrieved from the second database, and then the candidate reference information is screened by domain to obtain the target reference information. Then the target reference information is input into the large model, and more information is provided to the large model through the target reference information, so as to assist in the generation of the target detection result. This can alleviate the hallucination problem of the large model and further improve the accuracy of the detection result.
[0066] For example, the input text includes "Zhang San". The first candidate reference information obtained after retrieval includes: the personal introduction of a sports star named Zhang San, and the second candidate reference information also includes: the personal introduction of a doctor named Zhang San. After screening by domain, the first candidate reference information is filtered out, and the second candidate reference information is retained. In this way, the large model can generate an anomaly detection result based on the second candidate reference information, so as to avoid answering based on content related to the sports field.
[0067] Figure 4 It is a schematic principle diagram of the anomaly detection method according to an embodiment of the present disclosure.
[0068] In this embodiment, input information is first obtained, and the input information includes an input image 401 and an input text 406.
[0069] For the input image 401 in the input information, the position 403 and anomaly category 402 of the anomaly area can be determined through a target detection model. Then, according to the position 403 of the anomaly area, a sub-image 404 of the anomaly area is determined from the input image 401, and the similarity between the sub-image 404 and each data in the first database 405 is determined.
[0070] If the similarity between a certain data and the sub-image 404 is greater than or equal to the similarity threshold, it is determined that the data matches the sub-image 404, and the data is determined as the target data. At this time, it can be determined that the input information is abnormal, which is used as the anomaly detection result.
[0071] If the similarity between each data in the first database 405 and the sub-image 404 is less than the similarity threshold, it is determined that the first database 405 lacks target data that matches the sub-image 404. Next, based on a predetermined retrieval rule, candidate reference information can be retrieved from the second database 407 according to the input text 406. Then, the input text 406 and the candidate reference information are classified to obtain the fields involved in the input text 406 and the fields involved in the candidate reference information.
[0072] If the field involved in a certain candidate reference information is the same as the field involved in the input text 406, the candidate reference information is determined as the target reference information 408, otherwise the candidate reference information is filtered out, so that the target reference information 408 can be obtained.
[0073] Next, the input text 406, the location of the abnormal area 403, the abnormal category 402, the target reference information 408 and the prompt information can be combined, and then the combined information and the input image 401 are input into the large model 409, and the large model 409 generates the abnormality detection result 410.
[0074] In this embodiment, target detection is first used to determine whether the input image 401 has an abnormality. Then, when the target detection model detects an abnormal area, a second confirmation is performed by matching the input image 401 with the data in the first database 405, thereby reducing false detections. If the target detection result indicates that there is an abnormal area in the input image 401, and the target data is not matched in the first database 405, the large model 409 is used to review it again to confirm whether the input information is abnormal. In addition, in the process of generating the abnormality detection result 410 based on the large model 409, the target reference information 408 is also provided to the large model 409 to alleviate the hallucination problem of the large model 409 and further improve the accuracy of the detection results. Therefore, the abnormality detection method provided in this embodiment combines target detection, feature matching and the large model 409, which can improve detection accuracy, reduce false alarm rate, and improve real-time performance, thereby realizing efficient and real-time content review.
[0075] Figure 5 2 is a schematic structural block diagram of an abnormality detection device according to an embodiment of the present disclosure.
[0076] like Figure 5 As shown, the anomaly detection device 500 may include a target detection module 510 and a generation module 520.
[0077] The target detection module 510 is configured to perform target detection on an input image in the input information in response to obtaining input information, and obtain the location and abnormality category of an abnormal region in the input image.
[0078] The generation module 520 is used to generate anomaly detection results for the input information based on the large model according to the input text, input image, location of the abnormal area, abnormal category and target reference information in the input information.
[0079] According to another embodiment of the present disclosure, the apparatus further includes a determination module and a result determination module. The determination module is configured to determine, based on the abnormal region, whether target data matching the abnormal region exists in the first database; wherein the data stored in the first database is determined based on the abnormal information. The result determination module is configured to, in response to detecting the presence of the target data in the first database, determine that the abnormality detection result includes: the input information is abnormal.
[0080] According to another embodiment of the present disclosure, the determination module includes: a sub-image determination sub-module, a similarity determination sub-module, and a target data determination sub-module. The sub-image determination sub-module is configured to determine a sub-image of an abnormal region from an input image based on the location of the abnormal region. The similarity determination sub-module is configured to determine the similarity between the sub-image and data in a first database. The target data determination sub-module is configured to determine the data used for similarity determination as target data in response to detecting that the similarity is greater than or equal to a similarity threshold.
[0081] According to another embodiment of the present disclosure, the above-mentioned device further includes: a trigger module for triggering an operation of generating an anomaly detection result for the input information based on the large model in response to detecting that the target data is missing in the first database.
[0082] According to another embodiment of the present disclosure, a target detection module includes a first determination submodule, a second determination submodule, and a third determination submodule. The first determination submodule is configured to determine the location and category confidence of a region of interest in an input image. The second determination submodule is configured to determine a target threshold based on a detection scenario. The third determination submodule is configured to determine the region of interest as an abnormal region in response to detecting that the category confidence is greater than or equal to the target threshold.
[0083] According to another embodiment of the present disclosure, the apparatus further comprises: a retrieval module, a classification module, and a reference information determination module. The retrieval module is configured to retrieve candidate reference information from a second database based on the input text based on predetermined retrieval rules. The classification module is configured to classify the input text and the candidate reference information to obtain the fields covered by the input text and the fields covered by the candidate reference information. The reference information determination module is configured to determine the candidate reference information as target reference information in response to detecting that the fields covered by the candidate reference information are the same as the fields covered by the input text.
[0084] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, including at least one processor; and a memory communicatively connected to the at least one processor; the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the above-mentioned abnormality detection method.
[0085] According to an embodiment of the present disclosure, the present disclosure further provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute the above-mentioned abnormality detection method.
[0086] According to an embodiment of the present disclosure, the present disclosure further provides a computer program product, including a computer program, which implements the above-mentioned abnormality detection method when executed by a processor.
[0087] Figure 6is a block diagram of an electronic device used to implement the anomaly detection method of an embodiment of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided for example only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0088] like Figure 6 As shown, device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. RAM 603 may also store various programs and data required for the operation of device 600. Computing unit 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to bus 604.
[0089] Various components in device 600 are connected to I / O interface 605, including an input unit 606, such as a keyboard, mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a magnetic disk, optical disk, etc.; and a communication unit 609, such as a network card, modem, wireless communication transceiver, etc. The communication unit 609 allows device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0090] Computing unit 601 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Computing unit 601 performs the various methods and processes described above, such as the anomaly detection method. For example, in some embodiments, the anomaly detection method may be implemented as a computer software program tangibly embodied in a machine-readable medium, such as storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed onto device 600 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by computing unit 601, one or more steps of the anomaly detection method described above may be performed. Alternatively, in other embodiments, computing unit 601 may be configured to perform the anomaly detection method via any other suitable means (e.g., via firmware).
[0091] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-a-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0092] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0093] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0094] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0095] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0096] Computer systems may include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The client and server relationship arises through computer programs running on the respective computers and having a client-server relationship to each other.
[0097] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.
[0098] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A method for detecting anomalies, comprising: In response to obtaining input information, performing target detection on an input image in the input information to obtain a position and an abnormality category of an abnormal region in the input image; Determining, based on the abnormal area, whether there is target data matching the abnormal area in a first database; wherein the data stored in the first database is determined based on the abnormal information; In response to detecting that the target data exists in the first database, determining that the input information is abnormal; In response to detecting that the target data is missing from the first database, candidate reference information is retrieved from the second database based on the input text based on predetermined retrieval rules; the input text and the candidate reference information are classified to obtain the fields involved in the input text and the fields involved in the candidate reference information; in response to detecting that the fields involved in the candidate reference information are the same as the fields involved in the input text, the candidate reference information is determined as the target reference information; and based on the input text, the input image, the position of the abnormal area, the abnormal category and the target reference information in the input information, an abnormality detection result for the input information is generated based on a large model.
2. The method according to claim 1, wherein Determining, based on the abnormal area, whether there is target data matching the abnormal area in the first database includes: determining a sub-image of the abnormal area from the input image according to the position of the abnormal area; Determining the similarity between each data in the first database and the sub-image; and In response to detecting that the similarity between at least one data in the first database and the sub-image is greater than or equal to a similarity threshold, the at least one data is determined as the target data.
3. The method according to claim 1, wherein performing target detection on the input image in the input information comprises: Determining a location and a category confidence of a region of interest in the input image; Determine the target threshold based on the detection scenario; as well as In response to detecting that the category confidence is greater than or equal to the target threshold, the region of interest is determined as the abnormal region.
4. An anomaly detection device comprising: an object detection module, configured to, in response to obtaining input information, perform object detection on an input image in the input information to obtain a position and an abnormality category of an abnormal region in the input image; a determination module, configured to determine, based on the abnormal area, whether there is target data matching the abnormal area in a first database; wherein the data stored in the first database is determined based on the abnormal information; a result determination module, configured to determine that the input information is abnormal in response to detecting that the target data exists in the first database; A generation module is used to, in response to detecting that the target data is missing from the first database, retrieve candidate reference information in the second database based on predetermined retrieval rules and according to the input text; classify the input text and the candidate reference information to obtain the field involved in the input text and the field involved in the candidate reference information; in response to detecting that the field involved in the candidate reference information is the same as the field involved in the input text, determine the candidate reference information as the target reference information; and generate an anomaly detection result for the input information based on a large model according to the input text, the input image, the position of the abnormal area, the abnormal category and the target reference information in the input information.
5. The device according to claim 4, wherein The judgment module includes: a sub-image determination submodule, configured to determine a sub-image of the abnormal area from the input image according to the position of the abnormal area; a similarity determination submodule, configured to determine the similarity between the sub-image and each data in the first database; and The target data determination submodule is configured to, in response to detecting that the similarity between at least one data in the first database and the sub-image is greater than or equal to a similarity threshold, determine the at least one data as the target data.
6. The apparatus according to claim 4, wherein the target detection module comprises: A first determination submodule, configured to determine a position and a category confidence of a region of interest in the input image; The second determination submodule is used to determine the target threshold according to the detection scenario; as well as The third determining submodule is configured to determine the region of interest as the abnormal region in response to detecting that the category confidence is greater than or equal to the target threshold.
7. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 3.
8. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 3.
9. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 3.
Citation Information
Patent Citations
Target anomaly detection method and device
CN117173530A
Helical blade thickness wear detection method and system based on machine vision
CN117237367A