Abnormality recognition method and apparatus, and storage medium, electronic device and program product
Through the combination of two-stage feature extraction and target guidance information, the problem of low abnormal recognition efficiency in the prior art is solved, and the rapid and accurate recognition of past object objects is achieved, and the efficiency of abnormal recognition is improved.
Patent Information
- Application Number
- PCT/CN2024/117645
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-26
- Filing Date
- 2024-09-09
- Publication Date
- 2025-07-03
AI Technical Summary
In the prior art, the abnormal recognition efficiency is low, especially when it is necessary to identify past object objects, the model needs to be retrained, resulting in high resource consumption and low efficiency.
By obtaining the image to be identified, the two-stage feature extraction is performed. The basic visual features are first extracted, and then the second stage feature extraction is performed using the target guidance information, and finally comparing it with the standard visual features to determine the abnormal recognition result.
It realizes rapid identification of past objects, reduces the amount of calculation in the feature extraction process, and ensures the accuracy of recognition, thereby improving the efficiency of abnormal recognition.
Smart Images

Figure CN2024117645_03072025_PF_FP_ABST
Abstract
Description
Abnormality identification method, device, storage medium, electronic device and program product
[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on December 26, 2023, with application number 2023118209059 and application name “Abnormality Identification Method, Device, Storage Medium and Electronic Device”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present invention relates to the field of computers, and more particularly, to anomaly identification. Background Art
[0003] In anomaly recognition scenarios, a trained model is usually used to identify anomalies of objects. However, for past objects, the model still needs to be retrained to successfully identify anomalies of past objects. At the same time, model training itself requires a lot of resource support, which leads to low anomaly recognition efficiency. Therefore, there is a problem of low anomaly recognition efficiency.
[0004] To address the above-mentioned problems, no effective solutions have been proposed so far.
[0005] Summary of the Invention
[0006] The embodiments of the present application provide an anomaly identification method, apparatus, storage medium, electronic device, and program product to at least solve the technical problem of low anomaly identification efficiency.
[0007] According to one aspect of an embodiment of the present application, a method for identifying anomalies is provided, comprising: acquiring a first image to be identified, wherein the first image is an image obtained by capturing an object; performing a first stage of feature extraction on the first image to obtain at least one basic visual feature, wherein the basic visual feature is used to represent visual attributes presented by the surface of an object, and the object surface includes the surface of the object; determining, from at least one standard visual feature, a first visual feature whose feature similarity with the basic visual feature is greater than or equal to a preset threshold, wherein the standard visual feature is a feature obtained by extracting a plurality of sample images belonging to a plurality of object types, and the sample image corresponding to the first visual feature is a second image belonging to the target object type; obtaining target guidance information matched with the first visual feature, wherein the target guidance information is used to guide features related to the target object type in the image of interest; performing a second stage of feature extraction on the first image using the target guidance information to obtain at least one second visual feature related to the target object type; comparing the at least one second visual feature with the standard visual feature of the second image to obtain a comparison result; and determining an abnormality identification result of the first image based on the comparison result.
[0008] According to another aspect of an embodiment of the present application, an abnormality recognition device is also provided, including: a first acquisition unit, used to acquire a first image to be identified, wherein the above-mentioned first image is an image obtained by collecting an object object; a first extraction unit, used to perform a first stage of feature extraction on the above-mentioned first image to obtain at least one basic visual feature, wherein the above-mentioned basic visual feature is used to represent the visual attributes presented on the surface of the object, and the above-mentioned object surface includes the surface of the above-mentioned object object; a first determination unit, used to determine, from at least one standard visual feature, a first visual feature whose feature similarity with the above-mentioned basic visual feature is greater than or equal to a preset threshold, wherein the above-mentioned standard visual feature is extracted from multiple sample images belonging to multiple object types. The sample image corresponding to the first visual feature is a second image belonging to the target object type; the second acquisition unit is used to obtain target guidance information matched with the first visual feature, wherein the target guidance information is used to guide the features related to the target object type in the focus image; the second extraction unit is used to perform a second stage of feature extraction on the first image through the target guidance information to obtain at least one second visual feature related to the target object type; the first comparison unit is used to compare the at least one second visual feature with the standard visual feature of the second image to obtain a comparison result; the second determination unit is used to determine the abnormality recognition result of the first image through the comparison result.
[0009] According to another aspect of an embodiment of the present application, an electronic device is also provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the above-mentioned abnormality identification method through the computer program.
[0010] On the other hand, an embodiment of the present application provides a storage medium, which is used to store a computer program, and the computer program is used to execute the above-mentioned abnormality identification method.
[0011] On the other hand, an embodiment of the present application provides a computer program product including a computer program, which, when executed on a computer, enables the computer to execute the above-mentioned abnormality identification method.
[0012] In an embodiment of the present application, a first image to be identified is obtained, wherein the first image is an image obtained by capturing an object; a first stage of feature extraction is performed on the first image to obtain at least one basic visual feature, wherein the basic visual feature is used to represent the visual attributes presented on the surface of the object, and the object surface includes the surface of the object; a first visual feature whose feature similarity with the basic visual feature is greater than or equal to a preset threshold is determined from at least one standard visual feature, wherein the standard visual feature is a feature obtained by performing the first stage of extraction on multiple sample images belonging to multiple object types, and the first visual feature is a feature obtained by performing the first stage of extraction on a second image belonging to the target object type. The first stage of extraction obtains features; obtains target guidance information matching the first visual features, wherein the target guidance information is used to guide features related to the target object type in the focus image; performs a second stage of feature extraction on the first image using the target guidance information to obtain at least one second visual feature related to the target object type; when the standard visual features of the second image are obtained, compares the at least one second visual feature with the standard visual features of the second image to obtain a comparison result, wherein the standard visual feature is the feature obtained by performing the second stage of feature extraction on the second image; and determines the abnormal recognition result of the first image based on the comparison result. By extracting basic visual features and first visual features and comparing them with the standard visual features of the second image, rapid recognition of passing objects is achieved. Furthermore, by introducing target guidance information, this embodiment can not only quickly identify past objects, but also focus on features related to the target object type, which not only reduces the amount of calculation in the second-stage feature extraction process, but also ensures the accuracy of identifying past objects, thereby achieving the purpose of quickly identifying past objects while ensuring the accuracy of identification, thereby achieving the technical effect of improving the efficiency of abnormal recognition, and solving the technical problem of low abnormal recognition efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] FIG1 is a schematic diagram of an application environment of an optional anomaly identification method according to an embodiment of the present application;
[0014] FIG2 is a schematic diagram of a process of an optional abnormality identification method according to an embodiment of the present application;
[0015] FIG3 is a schematic diagram of an optional abnormality identification method according to an embodiment of the present application;
[0016] FIG4 is a schematic diagram of another optional abnormality identification method according to an embodiment of the present application;
[0017] FIG5 is a schematic diagram of another optional abnormality identification method according to an embodiment of the present application;
[0018] FIG6 is a schematic diagram of an optional abnormality identification device according to an embodiment of the present application;
[0019] FIG7 is a schematic structural diagram of an optional electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0020] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0021] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0022] For ease of understanding, the following terms are explained:
[0023] The solutions provided in the embodiments of this application involve technologies such as artificial intelligence computer vision technology, which are specifically described through the following embodiments:
[0024] According to one aspect of an embodiment of the present application, a method for identifying an anomaly is provided. Optionally, as an optional implementation, the method for identifying anomaly may be, but is not limited to, applied in an environment as shown in FIG1 . The environment may include, but is not limited to, a user device 102 and a server 112 . The user device 102 may include, but is not limited to, a display 104 , a processor 106 , and a memory 108 . The server 112 may include, but is not limited to, a database 114 and a processing engine 116 .
[0025] The specific process can be as follows:
[0026] Step S102: The user device 102 obtains a first image to be recognized, wherein the first image is an image obtained by capturing an object;
[0027] Step S104 , sending the first image to the server 112 via the network 110 ;
[0028] In steps S106-S114, the server 112 performs a first-stage feature extraction on the first image via the processing engine 116 to obtain at least one basic visual feature; determines, from the at least one standard visual feature, a first visual feature whose feature similarity with the basic visual feature is greater than or equal to a preset threshold; obtains target guidance information that matches the target object type; performs a second-stage feature extraction on the first image using the target guidance information to obtain at least one second visual feature related to the target object type; and, if the standard visual feature of the second image is obtained, compares the at least one second visual feature with the standard visual feature of the second image to obtain a comparison result.
[0029] In step S116 , the comparison result is sent to the user device 102 via the network 110 . The user device 102 determines the abnormality recognition result of the first image through the processor 106 through the comparison result, displays the abnormality recognition result on the display 104 , and stores the above comparison result in the memory 108 .
[0030] In addition to the example shown in Figure 1, the above-mentioned terminal device can be a terminal device configured with a target client, which can include but is not limited to at least one of the following: a mobile phone (such as an Android phone, an iOS phone, etc.), a laptop computer, a tablet computer, a PDA, an MID (Mobile Internet Devices), a PAD, a desktop computer, a smart TV, etc. The target client can be a video client, an instant messaging client, a browser client, an education client, etc. The above-mentioned network can include but is not limited to: a wired network, a wireless network, wherein the wired network includes: a local area network, a metropolitan area network and a wide area network, and the wireless network includes: Bluetooth, WIFI and other networks that realize wireless communication. The above-mentioned server can be a single server, or it can be a server cluster composed of multiple servers, or a cloud server. The above is only an example, and no limitation is made to this in this embodiment.
[0031] Optionally, as an optional implementation, as shown in FIG2 , the abnormality identification method may be performed by an electronic device, which may be, for example, a user device or a server as shown in FIG1 , and the specific steps include:
[0032] S202, obtaining a first image to be identified, wherein the first image is an image acquired by capturing an object;
[0033] S204, performing a first stage feature extraction on the first image to obtain at least one basic visual feature, wherein the basic visual feature is used to represent visual attributes presented by the surface of an object, where the surface of the object includes the surface of the object;
[0034] S206, determining, from the at least one standard visual feature, a first visual feature whose feature similarity to the basic visual feature is greater than or equal to a preset threshold, wherein the standard visual feature is a feature extracted from a plurality of sample images belonging to a plurality of object types, and the sample image corresponding to the first visual feature is a second image belonging to the target object type;
[0035] S208, obtaining target guidance information matched with the first visual feature, wherein the target guidance information is used to guide attention to features related to the target object type in the image;
[0036] S210, performing a second stage of feature extraction on the first image using the target guidance information to obtain at least one second visual feature related to the type of the target object;
[0037] S212, comparing the at least one second visual feature with the standard visual feature of the second image to obtain a comparison result;
[0038] S214: Determine an abnormality recognition result of the first image through comparison results.
[0039] Optionally, in this embodiment, the above-described anomaly identification method can be applied, but is not limited to, to continuous inspection scenarios for various industrial products, thereby achieving abnormal product location results in these scenarios. With model-based identification methods, each time a production line changes its manufactured items, the detection model must be retrained and redeployed, even if the replaced items have previously been inspected by the detection model. However, the method proposed in this embodiment does not require retraining and redeployment for previously encountered items. Additional training is only required when the product replaced on the production line is a new product. This significantly reduces the manpower, material resources, and time costs associated with retraining and redeploying for older products.
[0040] Optionally, in this embodiment, an object can be understood as, but is not limited to, an actual object that exists or needs to be identified, and can be any form of entity, such as a product, a part, or a living being. For example, on a production line, every part produced needs to be identified for anomalies. In this case, the part can be understood as an object. Whenever a new part is produced and captured by a camera, the image captured by the camera can be considered the first image.
[0041] Optionally, in this embodiment, basic visual features that can describe the visual properties of the surface of the object are extracted from the image to capture various properties of the object's surface, such as color, texture, shape, etc., to provide important information for subsequent abnormality identification. Among them, the surface visual properties of the object can refer to the characteristics presented by the surface of the object that can be perceived and recognized by the visual system. It is the basis for people to observe and recognize objects, and is also an important basis for computer vision systems to perform image analysis and understanding, such as gloss, color, texture, shape, transparency, etc.
[0042] In this embodiment, the basic visual features can be features extracted from the image to describe the surface visual properties of the object, covering multiple aspects and together constituting a comprehensive description of the visual characteristics of the object surface, such as color features, shape features, spatial relationship features, local features, etc.
[0043] Optionally, in this embodiment, a first visual feature that is similar to the basic visual feature extracted from the first image is determined from a series of standard visual features, such as by comparing the similarity between the basic visual feature and the standard visual features extracted from multiple sample images, and finding a standard visual feature whose similarity is greater than or equal to a preset threshold from these standard visual features as the first visual feature.
[0044] To further illustrate, it is optional to assume that there is a sample image library containing a variety of fruits, and each fruit has multiple sample images. Standard visual features are extracted from these sample images. There is an apple image to be identified (first image), from which basic visual features are extracted. Next, this embodiment compares these basic visual features with the standard visual features in the sample image library to find the first visual feature whose similarity with the apple image is greater than or equal to a preset threshold. The fruit included in the sample image corresponding to the first visual feature is very similar to the apple on the surface of the object and is likely to belong to the same object type, so these first visual features will be used for subsequent abnormality recognition of the apple image.
[0045] Optionally, in this embodiment, the standard visual features may be representative visual features extracted from multiple sample images of various object types. After processing and refinement, the standard visual features can accurately describe and represent the typical visual attributes of a particular object type. Extracting standard visual features typically involves analyzing and learning from a large number of sample images to extract the most representative and distinguishing features, such as color, texture, shape, spatial relationships, etc., used to describe the visual attributes of an object's surface.
[0046] Optionally, in this embodiment, feature similarity can be an indicator used to measure the degree of similarity between two visual features, typically obtained by calculating the distance or correlation between the two features. In anomaly recognition, feature similarity is used to compare the degree of similarity between basic visual features extracted from the image to be recognized and standard visual features extracted from the sample image. Feature similarity can be calculated using methods such as Euclidean distance, cosine similarity, and Pearson correlation coefficient. When the similarity between two features exceeds a preset threshold, the features can be considered similar.
[0047] Optionally, in this embodiment, object type refers to a classification or category of objects with similar visual attributes and features. In anomaly identification, object type can refer to a normal, non-anomalous object or a specific type of abnormal object. For example, on an industrial production line, object type may refer to type A for product A, type B for product B, or type C for product C. By identifying and classifying object types, anomalies can be detected and addressed promptly.
[0048] Optionally, in this embodiment, the sample images are image datasets used to extract standard visual features and first visual features. These datasets contain images of various object types, with each object type having multiple sample images. By analyzing and learning the sample images, standard visual features representing each object type can be extracted, including first visual features used for comparison and recognition.
[0049] Optionally, in this embodiment, target guidance information matching the target object type is obtained. The purpose of this target guidance information is to direct attention to features in the image related to the target object type, enabling more accurate identification and analysis of these features. Obtaining and using target guidance information matching the target object type can significantly improve the accuracy and efficiency of anomaly identification. By guiding the image processing algorithm to focus on features related to the target object type, the interference of irrelevant information on feature extraction can be reduced, improving the accuracy and efficiency of feature extraction. This will help to more accurately identify abnormal objects or phenomena and enable the timely implementation of appropriate treatment measures. Furthermore, the use of target guidance information can reduce the need for manual intervention, improve the degree of automated processing, and reduce the cost and complexity of anomaly identification.
[0050] To further illustrate, for example, a task of identifying anomalies in a specific type of fruit (e.g., apples) may be underway. In this case, the target guidance information may include information about the apple's color, shape, texture, and other visual features. This information will be used to guide the image processing algorithm to focus more on features associated with apples, such as red or green areas, round or oval shapes, specific texture patterns, and so on. By obtaining and using this target guidance information, the present embodiment can more accurately extract features associated with apples from the image, thereby performing anomaly identification.
[0051] Optionally, in this embodiment, the target guidance information can be information specifically designed to guide the image processing or analysis algorithm to pay more attention to features related to a specific object or task, and can be in various forms, such as pre-set rules, models or parameters learned from a large amount of data, or even guidance obtained through expert experience or user interaction. The target guidance information can be used to enhance the feature extraction process, making it more focused on features related to the task. For example, in a face recognition task, the target guidance information may guide the algorithm to pay more attention to key areas such as the eyes, nose, and mouth. In deep learning, the attention mechanism is a commonly used technique that dynamically adjusts the model's attention to different image areas through target guidance information. This can help the model better understand and interpret the image. At the same time, the target guidance information can also be used as a supervisory signal in the training process to help the model learn and converge to a good solution faster.
[0052] Optionally, in this embodiment, the second stage of feature extraction using target guidance information can significantly improve the accuracy and relevance of the second visual features. By guiding the algorithm to focus on features related to the target object type, the interference of irrelevant information can be reduced, thereby improving the accuracy and efficiency of feature extraction. This will help to more accurately identify objects in the image and distinguish subtle differences between different object types. At the same time, the extracted second visual features can also be used for further data processing tasks, such as object classification, scene understanding, 3D reconstruction, etc., providing strong support for improving the performance of computer vision systems.
[0053] To further illustrate, it is optional to assume that there is an image dataset containing multiple fruits, and target guidance information matching the target object type of apple has been obtained. This target guidance information may include features such as the color, shape, and texture of the apple. This embodiment can use this target guidance information to perform a second stage of feature extraction on a first image containing an apple, focusing on extracting second visual features related to the apple. These second visual features may include the smoothness of the apple's skin, the presence or absence of spots or stripes, etc., which will be used to more accurately identify the apple in the image and distinguish it from other types of fruit.
[0054] Optionally, in this embodiment, the second visual feature may be, but is not limited to, the same as the first visual feature, or may be, but is not limited to, different. For example, the second visual feature may be a feature used to describe the texture structure of the surface of an object in the image. Texture is a visual attribute of the surface of an object, typically consisting of many small, repetitive or periodic patterns. These patterns may be the result of the interaction of factors such as the microscopic geometry of the surface of the object, lighting conditions, and material properties, and may provide rich information about the surface structure and material properties of the object, thereby facilitating the identification and classification of the object. Compared with other visual features such as color and shape, the second visual feature is generally more robust to changes in lighting and viewing angle.
[0055] Optionally, in this embodiment, texture attributes refer to specific texture characteristics exhibited by the surface of an object. These attributes may include roughness, smoothness, directionality, regularity, etc. For example, the surface of a stone may exhibit rough and irregular texture attributes, while the surface of a glass may exhibit smooth and reflective texture attributes.
[0056] Optionally, in this embodiment, the at least one second visual feature related to the target object type refers to a second visual feature that is closely related to a specific object type. Each object type may have one or more unique second visual features, which are unique to that type of object or are the main differences between that type of object and other types of objects. Taking wood as an example, its typical second visual features may include annual rings, the direction and density of wood grain, etc. These features can be used to distinguish different types of wood, such as oak, pine, etc. In image processing, the extraction and analysis of second visual features for wood can help automatically identify and classify different types of wood.
[0057] Optionally, in this embodiment, by comparing at least one second visual feature extracted from the first image with a standard visual feature, a quantitative comparison result can be obtained, thereby accurately determining whether the first image contains an anomaly. This can also significantly improve the accuracy and efficiency of anomaly identification, reducing false positives and missed negatives. Furthermore, because the second visual feature is relatively robust to changes in lighting and viewing angle, this embodiment exhibits good stability and reliability in practical applications.
[0058] To further illustrate, it is optional to assume that there is a data set containing normal apple images, and standard visual features are extracted from it. Now, for an apple image to be identified (i.e., the first image), this embodiment extracts at least one second visual feature. Next, this embodiment compares this feature with the standard visual feature of the second image to obtain a comparison result. If the comparison result shows that the two are very similar, then this embodiment can consider that the apple image to be identified is normal; if the comparison result shows that there is a large difference between the two, then this embodiment can consider that there is an abnormality in this image.
[0059] Optionally, in this embodiment, the standard visual features are second visual features extracted from images that are considered to be "normal" or "normal-free". These features represent the texture properties that this type of object should have under normal circumstances. The standard visual features play the role of a "benchmark" in abnormality recognition. When a new image needs to be detected for abnormalities, its second visual features can be compared with the standard visual features to determine whether there are abnormalities. Typically, standard visual features are obtained by feature extraction and statistical analysis of a large number of normal sample images. This may require the use of specific algorithms or tools, such as deep learning models, image processing techniques, etc., to extract and describe these features.
[0060] Optionally, in this embodiment, the comparison result is obtained by comparing the second visual feature of the image to be tested with the standard visual feature. This result is typically a numerical value or a set of numerical values that quantifies the similarity or difference between the two. The comparison result may include a similarity score, a difference score, a distance metric, etc. For example, if Euclidean distance is used as the similarity metric, the comparison result will be a numerical value indicating the degree of difference between the second visual feature of the image to be tested and the standard image.
[0061] Optionally, in this embodiment, the abnormality identification result is a conclusion on whether the image to be detected is abnormal based on the comparison result. This result is usually a binary classification label (normal or abnormal) or a continuous value indicating the degree of abnormality. The abnormality identification result is usually determined by comparing the comparison result with a preset threshold or comparison standard. If the comparison result exceeds the threshold or does not meet the comparison standard, the image to be detected is determined to be abnormal; otherwise, it is determined to be normal.
[0062] It should be noted that this embodiment significantly improves the accuracy and efficiency of anomaly identification by combining a two-stage feature extraction and comparison process and using target guidance information to focus on features related to the target object type. This effectively reduces false positives and false negatives, thereby improving the accuracy and efficiency of anomaly identification. Furthermore, due to the use of multiple features and guidance information, this embodiment also has high flexibility and adaptability, and can be applied to different types of objects and anomaly identification tasks.
[0063] In an embodiment of the present application, the standard visual features used in the first stage and the standard visual features used in the second stage may be the same features. For example, for sample image a as the second image, after the first image obtains the basic visual features through the first stage, a similarity comparison will be performed based on the basic visual image and the standard visual features of the sample image a. After the first image obtains the second visual features through the second stage, a comparison will be performed based on the second visual features and the same standard visual features of the sample image a. That is, for the same sample image, feature extraction will not be repeated based on feature comparisons at different stages to obtain its standard visual features. In this case, the standard visual features can be extracted in advance.
[0064] In some cases, feature extraction may be performed twice on the same sample image as the second image in the two aforementioned stages to obtain standard visual features for feature comparison in different stages, respectively. This application does not impose any limitation on this.
[0065] To further illustrate, optionally, as shown in FIG3 , a first image 302 to be identified is obtained, wherein the first image 302 is an image obtained by collecting an object; a first stage of feature extraction is performed on the first image 302 to obtain at least one basic visual feature 304, wherein the basic visual feature 304 is used to represent the visual attributes presented on the surface of the object, and the surface of the object includes the surface of the object; a first visual feature 308 having a feature similarity with the basic visual feature 304 and being greater than or equal to a preset threshold is determined from at least one standard visual feature 306, wherein the standard visual feature 306 is a feature obtained by performing the first stage of extraction on multiple sample images belonging to each object type, and the first visual feature 308 is a feature obtained by performing the first stage of extraction on a second image belonging to the target object type 310, and the multiple sample images include the second image, and the second image is a first visual feature of the object type. an image obtained by capturing a standard object of the target object type 310; obtaining target guidance information 312 matching the target object type 310, wherein the target guidance information 312 is used to guide the features related to the target object type 310 in the focus image; performing a second stage of feature extraction on the first image 302 through the target guidance information 312 to obtain at least one second visual feature 314 related to the target object type 310, wherein the second visual feature is used to represent the texture attribute presented on the surface of the object; when the standard visual feature 316 of the second image is obtained, comparing the at least one second visual feature 314 with the standard visual feature 316 to obtain a comparison result 318, wherein the standard visual feature 316 is a feature obtained by performing the second stage of feature extraction on the second image; and determining an abnormality recognition result 320 of the first image 302 through the comparison result 318.
[0066] In the embodiment corresponding to FIG. 3 , the standard visual feature 306 and the standard visual feature 316 for the same sample image may be the same standard visual feature.
[0067] Through the embodiments provided by this application, basic visual features and first visual features are extracted and compared with the standard visual features of the second image, thereby achieving rapid recognition of past objects. Furthermore, by introducing target guidance information, this embodiment can not only quickly identify past objects, but also focus on features related to the target object type, not only reducing the amount of computation in the second-stage feature extraction process, but also ensuring the accuracy of recognition of past objects, thereby achieving the purpose of quickly identifying past objects while ensuring recognition accuracy, thereby achieving the technical effect of improving the efficiency of abnormal recognition.
[0068] As an optional solution, before obtaining target guidance information matching the target object type, the method further includes:
[0069] S1-1, performing region segmentation on the second image to obtain a first region and a second region, wherein a feature similarity between each second visual feature in the first region or the second region is greater than or equal to a first threshold;
[0070] S1-2, guiding the second stage feature extraction process of the image belonging to the target object type by the initial guidance information, extracting a first sample feature and a second sample feature, wherein the first sample feature is a feature having a first characteristic attribute, the second sample feature is a feature having a second characteristic attribute, the first characteristic attribute is a characteristic attribute within the first area, and the second characteristic attribute is a characteristic attribute within the second area;
[0071] S1-3, comparing the first sample feature with the second visual feature in the first area to obtain a first result, and comparing the second sample feature with the second visual feature in the second area to obtain a second result;
[0072] S1-4, adjusting the initial guidance information according to the first result and the second result to obtain target guidance information.
[0073] Optionally, in this embodiment, region segmentation can be a key step in image processing, involving dividing an image into multiple regions with similar visual attributes. These regions can be divided based on color, texture, shape or other image features. In region segmentation, the algorithm attempts to identify and aggregate visually similar parts of the image while preserving the boundaries between different parts. The main purpose of region segmentation is to simplify image representation, highlight important image structures, and reduce the complexity of subsequent processing tasks. By dividing the image into regions with similar attributes, it is easier to analyze and understand the image content and extract useful information. Region segmentation can be achieved by a variety of methods, including but not limited to threshold processing, edge detection, region growing, level set methods, clustering algorithms (such as K-means), graph cut methods, etc.
[0074] Optionally, in this embodiment, the characteristic attributes shared by the second visual features in a region refer to the common or similar visual characteristics exhibited by the texture in a specific region. These characteristic attributes can be descriptions based on aspects such as the shape, size, direction, density, periodicity, and roughness of the texture. Understanding the characteristic attributes shared by the second visual features in a region is very important for image processing and computer vision tasks. By identifying and analyzing these characteristic attributes, this embodiment can better understand the image content and extract useful information from it. For example, in an object recognition task, different objects may have different second visual feature attributes. By identifying and comparing these attributes, the classification and recognition of objects can be achieved.
[0075] Optionally, the initial guidance information provides this embodiment with basic information about the target object type or preliminary recognition results. This information can be basic features such as the object's shape, color, and texture, or the results obtained from the first-stage feature extraction. Based on this initial guidance information, this embodiment can perform a more targeted second-stage feature extraction on the image.
[0076] Optionally, in this embodiment, the image is divided into a plurality of image blocks. Each image block represents a local region of the image and contains information about the pixels within the region. By analyzing and processing these image blocks, this embodiment can extract more detailed and specific feature information.
[0077] Specifically, this embodiment utilizes initial guidance information to guide the feature extraction process for image blocks. It may focus on image blocks related to the initial guidance information and extract features from these blocks as first sample features. Simultaneously, this embodiment also considers relationships between image blocks and contextual information, extracting higher-level features related to the target object type as second sample features.
[0078] In this way, the image block serves as a carrier of local information. Under the guidance of the initial guidance information, this embodiment more accurately extracts feature information related to the target object type, thereby improving the efficiency and accuracy of feature extraction and providing strong support for subsequent classification, recognition or reconstruction tasks.
[0079] It should be noted that this embodiment firstly performs region segmentation and feature extraction on the second image to more accurately identify and extract the second visual features associated with the target object type. Secondly, by guiding the algorithm to focus on textures with specific characteristic attributes, the pertinence and accuracy of feature extraction can be further improved. Finally, by comparing and verifying the sample features and the second visual features within the region, the validity and accuracy of the final determined feature attributes of interest can be ensured, thereby significantly improving the accuracy and efficiency of subsequent anomaly identification.
[0080] To further illustrate, let's assume the goal is to identify a fabric with a unique texture. First, this embodiment acquires a sample image (second image) of this fabric and performs region segmentation, resulting in two regions with similar second visual features: a first region and a second region. Next, this embodiment analyzes the second visual features within these two regions, extracting their shared first and second characteristic attributes, such as texture density and orientation. This embodiment then directs the algorithm to focus on textures with these characteristic attributes, extracting first and second sample features from other images of this fabric. Next, this embodiment compares these sample features with the second visual features within the original region, obtaining first and second results. If both results indicate that the similarity between the sample features and the second visual features within the region exceeds a set threshold (the second and third thresholds), this embodiment confirms that the first and second characteristic attributes are valid characteristic attributes of interest. Otherwise, this embodiment adjusts these characteristic attributes and repeats the above process until the correct characteristic attributes of interest are found.
[0081] Through the embodiment provided by the present application, the second image is segmented into regions to obtain a first region and a second region, wherein the feature similarity between each second visual feature in the first region or the second region is greater than or equal to a first threshold; the initial guidance information is used to guide the second stage of feature extraction of the image belonging to the target object type, and the first sample feature and the second sample feature are extracted, wherein the first sample feature is a feature with a first feature attribute, the second sample feature is a feature with a second feature attribute, the first feature attribute is a feature attribute in the first region, and the second feature attribute is a feature attribute in the second region; the first sample feature and the second visual feature in the first region are compared to obtain a first result, and the second sample feature and the second visual feature in the second region are compared to obtain a second result; the initial guidance information is adjusted according to the first result and the second result to obtain target guidance information, thereby achieving the purpose of improving the pertinence and accuracy of feature extraction, thereby achieving the technical effect of providing effective support for subsequent efficient abnormality recognition.
[0082] As an optional solution, when the at least one second visual feature is a plurality of second visual features, the at least one second visual feature is compared with a standard visual feature of the second image to obtain a comparison result, including:
[0083] S2-1, obtaining first distances in a feature space between a plurality of second visual features and a standard visual feature of a second image, wherein the distance in the feature space is inversely related to the similarity between the features;
[0084] S2-2, determining, from the plurality of second visual features, a first feature with the farthest first distance and a second feature with the closest first distance;
[0085] S2-3, obtaining a second distance between the first feature and the second feature in the feature space;
[0086] S2-4, obtaining a comparison result based on the second distance.
[0087] Optionally, in this embodiment, a feature space is a high-dimensional space used to represent and compare image features. In this space, each point represents a feature, and the coordinates of the point are composed of the dimensions of the feature. The feature space provides a method for visualizing and quantifying the relationships between features, making it easier to understand and compare the similarities and differences between different features.
[0088] Optionally, in this embodiment, in the feature space, distance is a metric used to quantify the similarity or difference between two features. This distance is usually calculated based on a specific distance measurement function, such as Euclidean distance, Manhattan distance, cosine similarity, etc. Different distance measurement functions may emphasize different feature attributes or characteristics. In the feature space, the distance between two points reflects the similarity between the features they represent. The closer the distance, the higher the similarity; the farther the distance, the lower the similarity. This similarity measurement can help us understand the relationship between different features and which features are more important or useful in a specific task.
[0089] It should be noted that this embodiment can obtain a quantitative assessment of the similarity between at least one second visual feature and the standard visual feature, which can help identify the second visual features that are most and least similar to the standard visual feature and understand the degree of difference between them. This information is very valuable for subsequent image processing and analysis tasks (such as object recognition and scene understanding) because it can help more accurately understand and interpret image content.
[0090] To further illustrate, consider a set of second visual features A, B, and C and a standard visual feature S. In feature space, the distances between A, B, and C and S are measured, and it is found that A is the farthest from S, while C is the closest to S. Next, the distance between A and C (the second distance) is calculated. If this distance is large, it indicates that the differences between the second visual features are significant, and the comparison result may indicate that the second visual features do not match the standard visual features well. If this distance is small, it indicates that the differences between the second visual features are small, and the comparison result may be more positive.
[0091] Through the embodiments provided by the present application, each second visual feature of at least one second visual feature and a first distance between the second visual feature and the standard visual feature in the feature space are obtained, wherein the distance in the feature space is inversely related to the similarity between the features; the first feature with the farthest first distance and the second feature with the closest first distance are determined from at least one second visual feature; the second distance between the first feature and the second feature in the feature space is obtained; based on the second distance, a comparison result is obtained, thereby achieving the purpose of more accurately understanding and interpreting the image content, thereby realizing the technical effect of improving the accuracy of the comparison result.
[0092] As an optional solution, based on the second distance, a comparison result is obtained, including:
[0093] S3-1, calculating a third distance between a third feature and the first feature in the feature space, where the third feature is a second visual feature adjacent to the second feature in the feature space;
[0094] S3-2, performing robust optimization on the second distance using the third distance to obtain a target distance, wherein the comparison result includes the target distance;
[0095] As an optional solution, when the at least one basic visual feature is a first number of basic visual features, performing a first stage of feature extraction on the first image to obtain at least one basic visual feature includes:
[0096] S4-1, dividing the first image into a first number of image blocks, wherein the image blocks are used to represent local image regions of the first image;
[0097] S4-2, performing a first stage of feature extraction on the first number of image blocks to obtain a first image feature corresponding to each image block as a basic visual feature;
[0098] When the at least one second visual feature is a second number of second visual features, performing a second stage of feature extraction on the first image to obtain at least one second visual feature related to the type of the target object includes:
[0099] S5-1, dividing the first image into a second number of image blocks, wherein the image blocks are used to represent local image regions of the first image;
[0100] S5-2, performing a second stage of feature extraction on the second number of image blocks to obtain a second image feature corresponding to each image block as a second visual feature.
[0101] Optionally, in this embodiment, an image block can be a group of adjacent pixel regions in an image that have the same characteristics (such as grayscale value). In image processing, these adjacent pixel regions are typically treated as a whole for processing and analysis. The concept of image blocks is applied to various image processing tasks, such as feature extraction, image compression, and image enhancement. By processing image blocks, information in the image can be more efficiently extracted and computational complexity can be reduced.
[0102] Optionally, in this embodiment, the first number and the second number can be different. For example, considering the features of the image blocks extracted in the first stage, these features constitute the basic visual features used to determine the object type, and the accuracy requirements are not as strict, then the smaller first number can be used to coarsen the granularity of the image blocks to improve the efficiency of feature extraction in the first stage. For feature extraction in the second stage, which needs to be more relevant to the target object type and has more stringent accuracy requirements, the larger second number can be used to refine the granularity of the image blocks to improve the accuracy of feature extraction in the second stage. In other words, in this embodiment, the first number can be smaller than the second number.
[0103] It should be noted that this embodiment uses a two-stage feature extraction method to first obtain basic feature information from each local area of the image. Then, in the second stage, through more targeted processing, it can further extract features closely related to the target object type, thereby improving the recognition accuracy of the target object, especially in complex backgrounds or when the target object is not much different from the background.
[0104] To further illustrate, let's assume there's a first image of a basket containing fruit. In the first stage, the image is segmented into multiple image blocks, each containing a portion of a piece of fruit. From these image blocks, basic visual features such as color and shape can be extracted as first image features. In the second stage, if the target object type is an apple, more attention may be paid to image blocks related to apple features, and more specific features, such as the apple's red color, specific shape, or texture, may be extracted from these blocks as second image features.
[0105] Through the embodiments provided by the present application, a first image is divided into a first number of image blocks, wherein the image blocks are used to represent local image areas of the first image; a first stage of feature extraction is performed on the first number of image blocks to obtain first image features corresponding to each image block, wherein the basic visual features include the first image features; the first image is divided into a second number of image blocks, wherein the image blocks are used to represent local image areas of the first image; a second stage of feature extraction is performed on the second number of image blocks to obtain second image features corresponding to each image block, wherein the second visual features include the second image features, thereby achieving the purpose of performing two-stage feature extraction on the features at the granularity of the image blocks, thereby realizing the technical effect of improving the recognition accuracy of the target object.
[0106] As an optional solution, after performing the second stage of feature extraction on the second number of image blocks to obtain second image features corresponding to each image block, the method further includes:
[0107] Perform the following steps until the number of key image features stored according to the target object type is greater than or equal to the number threshold:
[0108] S6-1, determining a current image feature from the second image features corresponding to each image block;
[0109] S6-2 calculates, among the second image features corresponding to each image block, other image features except the current image feature, and a fourth distance between the current image feature in the feature space, wherein the distance in the feature space is inversely related to the similarity between the features;
[0110] S6-3 determines the second image feature with the largest fourth distance as the key image feature;
[0111] S6-4, storing the obtained key image features according to the target object type;
[0112] S6-5: When the number of obtained key image features is less than the number threshold, determine the next image feature from the second image features corresponding to each image block, and use the next image feature as the current image feature.
[0113] Optionally, in this embodiment, the key image features are used to update the standard visual features belonging to the target object type. For example, when the second stage of feature extraction is performed on the second number of image blocks to obtain the second image features corresponding to each of the image blocks, the standard visual features belonging to the target object type are the standard visual features of the second image. After obtaining the key image features, the standard visual features of the second image can be optimized by the key image features to be closer to the true expression of the target object type, thereby realizing the update of the target object type.
[0114] It should be noted that this embodiment allows for the gradual extraction of key image features most relevant to the target object type. These key features are not only highly correlated with the target object type but also exhibit significant differences from one another, enabling the pre-storage of key characteristics that more comprehensively describe the target object type. Once a sufficient number of key image features are obtained, they can be used in subsequent classification, recognition, or reconstruction tasks, ensuring both accuracy and efficiency in anomaly identification.
[0115] To further illustrate, it is optional to assume that there is an image containing multiple fruits, and the target object type is apple. After the second image features of multiple image blocks are extracted in the second stage, the iterative process begins. First, a feature of an image block is selected as the current image feature. Then, the fourth distance between the feature and the features of all other image blocks is calculated. Assuming that the feature of a certain image block is found to be the most different from the current image feature, then this feature is determined to be the key image feature and stored. If the number of stored key image features is insufficient, the feature of the next image block will continue to be selected as the current image feature and the above process will be repeated. If the number of stored key image features is sufficient, the repetition will end.
[0116] Through the embodiment provided by the present application, the following steps are performed until the key image features stored according to the target object type are obtained: the current image feature is determined from the second image features corresponding to each image block; the other image features except the current image feature are calculated from the second image features corresponding to each image block, and the fourth distance of the current image feature in the feature space is obtained, wherein the distance in the feature space is inversely related to the similarity between the features; the second image feature with the largest fourth distance is determined as the key image feature; when the number of key image features obtained is greater than or equal to the number threshold, the key image features obtained are stored according to the target object type; when the number of key image features obtained is less than the number threshold, the next image feature is determined from the second image features corresponding to each image block, and the next image feature is used as the current image feature, thereby achieving the purpose of pre-storing some key characteristics that can more comprehensively describe the target object type, thereby achieving the technical effect of taking into account both the accuracy and efficiency of abnormality recognition.
[0117] As an optional solution, after performing the first stage feature extraction on the first image to obtain at least one basic visual feature, the method further includes:
[0118] S7-1, when the feature similarity between each standard visual feature of the at least one standard visual feature and the basic visual feature is less than a preset threshold, setting the object type to which the object belongs to as an additional object type, wherein the additional object type is a type other than the rated object type to which the plurality of sample images belong;
[0119] S7-2, capturing standard objects of the additional object type to obtain a plurality of third images, and using the third images as sample images.
[0120] Optionally, in this embodiment, in the field of image recognition and processing, a database (multiple sample images) containing multiple known object types is typically constructed. This database is used to train algorithms to recognize these known types. However, in actual applications, the system may encounter objects that are not in this database. In this case, a mechanism is required to handle these unknown, additional object types. In other words, additional object types are any object types that do not fall within the known or predefined nominal object types.
[0121] To further illustrate, consider a recognition system that can identify two object types: "cat" and "dog." When the system encounters an image containing a "rabbit," it extracts basic visual features. It then compares these features to the standard visual features of "cat" and "dog." If the similarity with both falls below a preset threshold, the system considers the object to be an additional object type, namely, "rabbit." Next, the system collects an image of the standard object type for "rabbit," i.e., a third image, and adds it to the sample image.
[0122] It should be noted that these additional steps enhance anomaly recognition capabilities, enabling it to handle a wider variety of objects. By dynamically adding newly recognized object types to the sample image library, it can gradually adapt to and recognize a wider range of scenes and objects. Furthermore, this embodiment enhances the flexibility of anomaly recognition, enabling it to better cope with various challenges and changes in practical applications.
[0123] Through the embodiments provided by the present application, when the feature similarity between each standard visual feature of at least one standard visual feature and the basic visual feature is less than a preset threshold, the object type to which the object belongs is set as an additional object type, wherein the additional object type is a type other than the rated object type to which multiple sample images belong; the standard object objects of the additional object type are collected to obtain multiple third images, and the third images are used as sample images, thereby achieving the purpose of the additional step being able to enhance the ability of anomaly recognition so that it can handle more types of object objects, thereby achieving the technical effect of improving the flexibility of anomaly recognition.
[0124] As an optional solution, for ease of understanding, the above-mentioned anomaly identification method is applied to a continuous detection scenario, as shown in FIG4 , to achieve abnormal product positioning results in the continuous detection scenario.
[0125] Optionally, in this embodiment, as shown in FIG5 , the process of applying the above-mentioned anomaly recognition method includes a training process and a testing process, and is executed by a continuous prompting module (CPM) and a structure-based contrastive learning module (SCL).
[0126] Specifically, the continuous prompt module (CPM) is mainly divided into three parts: first, the key for one-to-one retrieval from image to module; second, the prompt for extracting category-specific features; and finally, the knowledge part for storing category-specific features.
[0127] The Structure-based Contrastive Learning (SCL) module aims to distill the knowledge of the SAM (Segment Everything Model) into the base ViT. By training Prompt, ViT's feature representation of specific items can be modified. This Prompt is an additional parameter of a small magnitude that can be directly superimposed on each layer of ViT features to fine-tune the model.
[0128] Optionally, during the training process, the training image is first subjected to feature extraction using ViT and stored in a key for use during testing. Next, the image is passed through SAM to generate a segmentation map. Different regions in the segmentation map have consistent secondary visual features and corresponding labels. This information is then passed to the feature map extracted by ViT, enabling comparative learning with the SAM segmentation map.
[0129] With the features extracted by ViT and the labels obtained by SAM, we can perform contrastive learning on this basis and obtain the gradient for backpropagation. The contrastive loss function used in Figure 5 is used to calculate the loss in the contrastive learning process. First, the feature map is normalized. Then, the similarity matrix of the patch features is calculated. A mask is created based on the consistency of the labels and used to calculate the loss, ensuring that features with the same label are close to each other, while features with different labels are far away from each other. It is defined as follows:
[0130] Furthermore, after 50 epochs of training, Prompt was optimized and, when combined with features extracted from ViT, was able to better characterize the specific secondary visual features of different items. The model then extracted features from all images of each category and used the Farthest Point Sampling (FPS) algorithm to select the most representative features and store them in the Knowledge section of the CPM.
[0131] In the test process, the input image (Inference Image) in Figure 5 is first subjected to feature extraction through ViT. These features are compared one by one with the features of the Key module in CPM to find the product category and the corresponding CPM corresponding to the input image. Once the corresponding CPM is found, the input image is processed again through ViT, but this time the corresponding CPM Prompt is additionally added to extract category-specific features. After obtaining these features, they are compared with the Knowledge in the corresponding CPM to calculate the anomaly score of the corresponding image block. The method of retrieving the Key corresponding to the CPM based on the feature and retrieving the knowledge of the CPM based on the feature for comparison is the same. In this way, the method can accurately identify and locate abnormal areas in the input image, providing an important basis for subsequent anomaly processing and analysis.
[0132] Assume that all features of an image are x test , a certain block feature m test is x test If the Key or Knowledge to be retrieved is K and m is any feature in K, then you want to retrieve m test The closest feature in K can be obtained by PatchCore using the following formulas (1), (2), and (3): s * =||m test,* -m * ||2 (2)
[0133] For any x test The anomaly score between x and K can be calculated by the above formulas (1), (2), and (3). Formula (1) can be used to find the anomaly score between x and K. test The representative feature m of the distance from K test,* ,m * . m test,* is x test The feature farthest from K, and m * is the distance m in K test,* The closest feature. As shown in formula (2), the distance between the two can be calculated to obtain the basic anomaly score s* , and formula (3) further calculates m in K * Neighbors and m test,* The distance can make the score s more robust.
[0134] The unique features of a specific category are extracted and compared with the knowledge in the CPM. The anomaly score for each tile in the test image is then determined. This completes the calculation of the overall anomaly score for the test image. To obtain a more detailed anomaly distribution, the test image is divided into different tiles, and the score for each tile is calculated using the above method, but only the features of the tile to be calculated are considered and processed.
[0135] In this way, an anomaly score is obtained for each tile in the test image. To obtain a continuous and more intuitive anomaly distribution map, these discrete tile scores are further processed using Gaussian smoothing. Gaussian smoothing ensures that the anomaly scores form a continuous, gradual distribution in the image space, thus generating the final anomaly score map. The anomaly score map not only clearly shows which areas in the image contain anomalies, but also reflects the severity of the anomaly through the gradual change in the score, providing an intuitive and powerful tool for subsequent anomaly processing and analysis.
[0136] The examples of this application demonstrate the high flexibility of the anomaly identification method. Specifically, the algorithm used to calculate the anomaly score map is not limited to the Gaussian smoothing mentioned above; it can be replaced with other suitable algorithms based on actual needs or application scenarios. This flexibility ensures that the method can adapt to different data distributions and anomaly types, thereby improving its practicality and accuracy.
[0137] In addition to the flexibility of the anomaly score map calculation algorithm, the training process in this embodiment also allows for adjustments. For example, the contrastive learning loss function can be modified or replaced based on specific needs to accommodate different training objectives and data characteristics. Similarly, the contrastive learning method during training can also be adjusted as needed, providing the possibility of further optimizing model performance.
[0138] Furthermore, in this embodiment, Prompt is used to enhance the model's ability to extract corresponding types of information without changing the structure of the backbone network. This strategy effectively improves the model's representational ability and accuracy. However, the embodiment also provides an alternative: adding an adaptation module after the backbone network. This adaptation module can have a similar effect to Prompt, that is, enhancing the model's sensitivity to specific information by fine-tuning or expanding the network structure. This flexibility ensures that the model can be optimized for performance without changing its core structure, thereby reducing the complexity and cost of model adjustment.
[0139] It is understandable that in the specific implementation of this application, related data such as user information is involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.
[0140] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.
[0141] According to another aspect of the embodiments of the present application, an abnormality identification device for implementing the above abnormality identification method is also provided. As shown in FIG6 , the device includes:
[0142] A first acquiring unit 602 is configured to acquire a first image to be identified, wherein the first image is an image acquired by capturing an object;
[0143] a first extraction unit 604 configured to perform a first stage of feature extraction on the first image to obtain at least one basic visual feature, wherein the basic visual feature is used to represent visual attributes presented by an object surface, where the object surface includes a surface of an object;
[0144] A first determining unit 606 is configured to determine, from at least one standard visual feature, a first visual feature whose feature similarity to the basic visual feature is greater than or equal to a preset threshold, wherein the standard visual feature is a feature obtained by extracting features from a plurality of sample images belonging to a plurality of object types, and the sample image corresponding to the first visual feature is a second image belonging to the target object type;
[0145] A second acquisition unit 608 is configured to acquire target guidance information matched with the first visual feature, wherein the target guidance information is used to guide attention to features related to the target object type in the image;
[0146] A second extraction unit 610 is configured to perform a second stage of feature extraction on the first image using the target guidance information to obtain at least one second visual feature related to the target object type;
[0147] A first comparison unit 612 is configured to compare the at least one second visual feature with a standard visual feature of a second image to obtain a comparison result;
[0148] The second determining unit 614 is configured to determine an abnormality recognition result of the first image based on the comparison result.
[0149] For specific embodiments, reference may be made to the examples shown in the above-mentioned abnormality identification method, which will not be described in detail in this example.
[0150] As an optional solution, the device further includes:
[0151] a segmentation unit, configured to, before obtaining target guidance information matching the target object type, perform region segmentation on the second image to obtain a first region and a second region, wherein a feature similarity between each second visual feature within the first region or the second region is greater than or equal to a first threshold;
[0152] a third extraction unit, configured to, before obtaining target guidance information matching the target object type, guide the image belonging to the target object type through the initial guidance information to perform a second stage of feature extraction, extracting a first sample feature and a second sample feature, wherein the first sample feature is a feature having a first characteristic attribute, the second sample feature is a feature having a second characteristic attribute, the first characteristic attribute is a characteristic attribute within the first region, and the second characteristic attribute is a characteristic attribute within the second region;
[0153] a second comparison unit, configured to compare the first sample feature with the second visual feature in the first area to obtain a first result, and to compare the second sample feature with the second visual feature in the second area to obtain a second result, before obtaining target guidance information matching the target object type;
[0154] The third determining unit is configured to adjust the initial guidance information according to the first result and the second result before acquiring target guidance information matching the target object type, so as to obtain target guidance information.
[0155] For specific embodiments, reference may be made to the examples shown in the above-mentioned abnormality identification method, which will not be described in detail in this example.
[0156] As an optional solution, when the at least one second visual feature is a plurality of second visual features, the first comparing unit 612 includes:
[0157] a first acquisition module, configured to acquire first distances between the plurality of second visual features and the standard visual feature of the second image in a feature space, wherein the distances in the feature space are inversely related to the similarity between the features;
[0158] A first determining module is configured to determine, from a plurality of second visual features, a first feature with the longest first distance and a second feature with the shortest first distance;
[0159] A second acquisition module, configured to acquire a second distance between the first feature and the second feature in the feature space;
[0160] The third acquisition module is used to obtain a comparison result based on the second distance.
[0161] For specific embodiments, reference may be made to the examples shown in the above-mentioned abnormality identification method, which will not be described in detail in this example.
[0162] As an optional solution, the third acquisition module includes:
[0163] a calculation submodule, configured to calculate a third distance between a third feature and the first feature in the feature space, wherein the third feature is a second visual feature adjacent to the second feature in the feature space;
[0164] an optimization submodule, configured to perform robust optimization on the second distance using the third distance to obtain a target distance, wherein the comparison result includes the target distance;
[0165] The second determining unit 614 includes:
[0166] a second determining module, configured to determine, when the target distance is greater than or equal to a preset distance threshold, that the abnormality recognition result is that the first image is in an abnormal state;
[0167] The third determining module is configured to determine, when the target distance is less than a preset distance threshold, that the abnormal recognition result is that the first image is in a normal state.
[0168] For specific embodiments, reference may be made to the examples shown in the above-mentioned abnormality identification method, which will not be described in detail in this example.
[0169] As an optional solution, when the at least one basic visual feature is a first number of basic visual features, the first extraction unit 604 includes:
[0170] A first segmentation module, configured to segment the first image into a first number of image blocks, wherein the image blocks are used to represent local image regions of the first image;
[0171] a first extraction module, configured to perform a first stage of feature extraction on a first number of image blocks, and obtain a first image feature corresponding to each image block as a basic visual feature;
[0172] When the at least one second visual feature is a second number of second visual features, the second extraction unit 610 includes:
[0173] A second segmentation module, configured to segment the first image into a second number of image blocks, wherein the image blocks are used to represent local image regions of the first image;
[0174] The second extraction module is used to perform a second stage of feature extraction on a second number of image blocks to obtain a second image feature corresponding to each image block as a second visual feature.
[0175] For specific embodiments, reference may be made to the examples shown in the above-mentioned abnormality identification method, which will not be described in detail in this example.
[0176] As an optional solution, the device further includes:
[0177] After performing the second stage of feature extraction on the second number of image blocks to obtain second image features corresponding to each image block, an execution module is configured to perform the following steps until the number of key image features stored according to the target object type is greater than or equal to a quantity threshold:
[0178] Determining a current image feature from the second image features corresponding to each image block;
[0179] Calculating a fourth distance between the second image features corresponding to each image block, excluding the current image feature, and the current image feature in the feature space, wherein the distance in the feature space is inversely related to the similarity between the features;
[0180] determining the second image feature with the largest fourth distance as the key image feature;
[0181] When the number of obtained key image features is greater than or equal to the number threshold, the obtained key image features are stored according to the target object type;
[0182] When the number of the obtained key image features is less than the number threshold, the next image feature is determined from the second image features corresponding to each image block, and the next image feature is used as the current image feature.
[0183] For specific embodiments, reference may be made to the examples shown in the above-mentioned abnormality identification method, which will not be described in detail in this example.
[0184] As an optional solution, the device further includes:
[0185] a setting unit configured to, after performing a first-stage feature extraction on the first image to obtain at least one basic visual feature, set the object type to which the object belongs as an additional object type if the feature similarity between each standard visual feature from the at least one standard visual feature and the basic visual feature is less than a preset threshold, wherein the additional object type is a type other than the rated object type to which the plurality of sample images belong;
[0186] The acquisition unit is used to perform a first-stage feature extraction on the first image to obtain at least one basic visual feature, and then acquire standard object objects of additional object types to obtain multiple third images, and use the third images as sample images.
[0187] For specific embodiments, reference may be made to the examples shown in the above-mentioned abnormality identification method, which will not be described in detail in this example.
[0188] According to another aspect of an embodiment of the present application, an electronic device for implementing the above-mentioned abnormality identification method is also provided. The electronic device may be, but is not limited to, the user device 102 or the server 112 shown in Figure 1. This embodiment takes the electronic device as the user device 102 as an example. Further, as shown in Figure 7, the electronic device includes a memory 702 and a processor 704. A computer program is stored in the memory 702, and the processor 704 is configured to execute the steps in any of the above-mentioned method embodiments through the computer program.
[0189] Optionally, in this embodiment, the electronic device may be located in at least one network device among a plurality of network devices of a computer network.
[0190] Optionally, in this embodiment, the processor may be configured to execute the following steps through a computer program:
[0191] S1, obtaining a first image to be identified, wherein the first image is an image acquired by capturing an object;
[0192] S2, performing a first stage feature extraction on the first image to obtain at least one basic visual feature, wherein the basic visual feature is used to represent visual attributes presented by the surface of an object, where the surface of the object includes the surface of the object;
[0193] S3, determining, from at least one standard visual feature, a first visual feature whose feature similarity to the basic visual feature is greater than or equal to a preset threshold, wherein the standard visual feature is a feature extracted from a plurality of sample images belonging to a plurality of object types, and the sample image corresponding to the first visual feature is a second image belonging to the target object type;
[0194] S4, obtaining target guidance information matched with the first visual feature, wherein the target guidance information is used to guide attention to features related to the target object type in the image;
[0195] S5, performing a second stage of feature extraction on the first image using the target guidance information to obtain at least one second visual feature related to the type of the target object;
[0196] S6, comparing the at least one second visual feature with the standard visual feature of the second image to obtain a comparison result;
[0197] S7, determining the abnormality recognition result of the first image through the comparison result.
[0198] Alternatively, those skilled in the art will appreciate that the structure shown in FIG7 is merely illustrative and does not limit the structure of the electronic device. For example, the electronic device may include more or fewer components (such as network interfaces) than those shown in FIG7 , or may have a configuration different from that shown in FIG7 .
[0199] Among them, the memory 702 can be used to store software programs and modules, such as the program instructions / modules corresponding to the abnormality identification method and device in the embodiments of the present application. The processor 704 executes various functional applications and data processing by running the software programs and modules stored in the memory 702, that is, implementing the above-mentioned abnormality identification method. The memory 702 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 702 may further include a memory remotely located relative to the processor 704, and these remote memories can be connected to the electronic device via a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof. Among them, the memory 702 can specifically be used, but is not limited to, to store information such as the first image, the second image, and the abnormality identification result. As an example, as shown in FIG7 , the memory 702 may include, but is not limited to, the first acquisition unit 602, the first extraction unit 604, the first determination unit 606, the second acquisition unit 608, the second extraction unit 610, the first comparison unit 612, and the second determination unit 614 of the abnormality identification device. Furthermore, the memory 702 may also include, but is not limited to, other module units of the abnormality identification device, which will not be described in detail in this example.
[0200] Optionally, the transmission device 706 is configured to receive or send data via a network. Specific examples of the network may include a wired network and a wireless network. In one embodiment, the transmission device 706 includes a network interface controller (NIC), which can be connected to other network devices and a router via a network cable to communicate with the Internet or a local area network. In one embodiment, the transmission device 706 is a radio frequency (RF) module, which is configured to communicate with the Internet wirelessly.
[0201] In addition, the electronic device further includes: a display 708 for displaying the first image, the second image, abnormality recognition results and other information; and a connection bus 710 for connecting various module components in the electronic device.
[0202] In other embodiments, the user device or server may be a node in a distributed system, wherein the distributed system may be a blockchain system, and the blockchain system may be a distributed system formed by connecting multiple nodes via network communication. The nodes may form a peer-to-peer network, and any computing device, such as a server, user device, or other electronic device, may become a node in the blockchain system by joining the peer-to-peer network.
[0203] According to one aspect of the present application, a computer program product is provided, comprising a computer program / instructions containing program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication component and / or installed from a removable medium. When the computer program is executed by a central processing unit, the various functions provided in the embodiments of the present application are performed.
[0204] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0205] It should be noted that the computer system of the electronic device is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0206] A computer system includes a central processing unit (CPU), which can perform various appropriate actions and processes based on programs stored in read-only memory (ROM) or programs loaded from the storage unit into random access memory (RAM). The RAM also stores various programs and data required for system operation. The CPU, the read-only memory, and the RAM are connected to each other via a bus. Input / output interfaces (I / O interfaces) are also connected to the bus.
[0207] The following components are connected to the input / output interface: an input section including a keyboard, mouse, etc.; an output section including a cathode ray tube (CRT), a liquid crystal display (LCD), and a speaker; a storage section including a hard disk; and a communication section including a network interface card such as a local area network card and a modem. The communication section performs communication processing via a network such as the Internet. A drive is also connected to the input / output interface as needed. Removable media such as magnetic disks, optical disks, magneto-optical disks, semiconductor memories, etc. are installed in the drive as needed so that computer programs read from them can be installed into the storage section as needed.
[0208] In particular, according to an embodiment of the present application, the processes described in the various method flow charts can be implemented as computer software programs. For example, an embodiment of the present application includes a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for executing the methods shown in the flow charts. In such an embodiment, the computer program can be downloaded and installed from a network via a communication portion, and / or installed from a removable medium. When the computer program is executed by a central processing unit, the various functions defined in the system of the present application are performed.
[0209] According to one aspect of the present application, a computer-readable storage medium is provided, and a processor of a computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the methods provided in the various optional implementations described above.
[0210] Optionally, in this embodiment, the computer-readable storage medium may be configured to store a computer program for performing the following steps:
[0211] S1, obtaining a first image to be identified, wherein the first image is an image acquired by capturing an object;
[0212] S2, performing a first stage feature extraction on the first image to obtain at least one basic visual feature, wherein the basic visual feature is used to represent visual attributes presented by the surface of an object, where the surface of the object includes the surface of the object;
[0213] S3, determining, from at least one standard visual feature, a first visual feature whose feature similarity to the basic visual feature is greater than or equal to a preset threshold, wherein the standard visual feature is a feature extracted from a plurality of sample images belonging to a plurality of object types, and the sample image corresponding to the first visual feature is a second image belonging to the target object type;
[0214] S4, obtaining target guidance information matched with the first visual feature, wherein the target guidance information is used to guide attention to features related to the target object type in the image;
[0215] S5, performing a second stage of feature extraction on the first image using the target guidance information to obtain at least one second visual feature related to the type of the target object;
[0216] S6, comparing the at least one second visual feature with the standard visual feature of the second image to obtain a comparison result;
[0217] S7, determining the abnormality recognition result of the first image through the comparison result.
[0218] Alternatively, in the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0219] Optionally, in this embodiment, a person of ordinary skill in the art may understand that all or part of the steps in the various methods of the above embodiments may be completed by instructing hardware related to the electronic device through a program, and the program may be stored in a computer-readable storage medium, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0220] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0221] If the integrated units in the above embodiments are implemented in the form of software functional units and sold or used as independent products, they can be stored in the above-mentioned computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling one or more computer devices (which can be personal computers, servers, or network devices, etc.) to execute all or part of the steps of the methods of each embodiment of the present application.
[0222] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0223] In the several embodiments provided in this application, it should be understood that the disclosed user equipment can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is merely a logical functional division. In actual implementation, other division methods may be used. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not implemented. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interface, or indirect coupling or communication connection between units or modules may be electrical or other forms.
[0224] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0225] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0226] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. An anomaly recognition method, which is executed by an electronic device, and the method includes: Obtain a first image to be recognized, where the first image is an image collected for an object; Perform feature extraction in a first stage on the first image to obtain at least one basic visual feature, where the basic visual feature is used to represent the visual attributes presented on the surface of the object, and the object surface includes the surface of the object; Determine, from at least one standard visual feature, a first visual feature whose feature similarity with the basic visual feature is greater than or equal to a preset threshold, where the standard visual feature is a feature obtained by performing feature extraction on multiple sample images belonging to multiple object types, and the sample image corresponding to the first visual feature is a second image belonging to the target object type; Obtain target guidance information matching the first visual feature, where the target guidance information is used to guide the attention to features related to the target object type in the image; Perform feature extraction in a second stage on the first image through the target guidance information to obtain at least one second visual feature related to the target object type; Compare the at least one second visual feature with the standard visual feature of the second image to obtain a comparison result; Determine the anomaly recognition result of the first image through the comparison result.
2. The method according to claim 1, before obtaining the target guidance information matching the target object type, the method further includes: Perform region segmentation on the second image to obtain a first region and a second region, where the feature similarity within the first region or the second region is greater than or equal to a first threshold; During the process of performing feature extraction in the second stage on the image belonging to the target object type guided by the initial guidance information, extract a first sample feature and a second sample feature, where the first sample feature is a feature with a first feature attribute, the second sample feature is a feature with a second feature attribute, the first feature attribute is the feature attribute within the first region, and the second feature attribute is the feature attribute within the second region; Compare the first sample feature with the second visual feature within the first region to obtain a first result, and compare the second sample feature with the second visual feature within the second region to obtain a second result; Adjust the initial guidance information according to the first result and the second result to obtain the target guidance information.
3. The method according to claim 1, when the at least one second visual feature is multiple second visual features, the comparing the at least one second visual feature with the standard visual feature of the second image to obtain a comparison result includes: Obtain a first distance between each of the multiple second visual features and the standard visual feature of the second image in a feature space, where the distance in the feature space has an inverse relationship with the similarity between features; Determine a first feature with the farthest first distance and a second feature with the closest first distance from the multiple second visual features; Obtain a second distance between the first feature and the second feature in the feature space; Based on the second distance, obtain the comparison result.
4. The method according to claim 3, wherein the obtaining the comparison result based on the second distance includes: Calculate a third distance between a third feature and the first feature in the feature space, where the third feature is a second visual feature adjacent to the second feature in the feature space; Robustly optimize the second distance through the third distance to obtain a target distance, where the comparison result includes the target distance; The determining the anomaly recognition result of the first image through the comparison result includes: When the target distance is greater than or equal to a preset distance threshold, determine that the anomaly recognition result is that the first image is in an abnormal state; When the target distance is less than the preset distance threshold, determine that the anomaly recognition result is that the first image is in a normal state.
5. The method according to claim 1, when the at least one basic visual feature is a first number of basic visual features, the performing a first-stage feature extraction on the first image to obtain at least one basic visual feature includes: Divide the first image into a first number of image blocks, where the image blocks are used to represent local image regions of the first image; Perform the first-stage feature extraction on the first number of image blocks to obtain first image features corresponding to the respective image blocks as the basic visual features; When the at least one second visual feature is a second number of second visual features, the performing a second-stage feature extraction on the first image through the target guidance information to obtain at least one second visual feature related to the target object type includes: Divide the first image into a second number of image blocks, where the image blocks are used to represent local image regions of the first image; Perform the second-stage feature extraction on the second number of image blocks through the target guidance information to obtain second image features corresponding to the respective image blocks as the second visual features.
6. The method according to claim 5, after the performing the second-stage feature extraction on the second number of image blocks to obtain second image features corresponding to the respective image blocks as the second visual features, the method further includes: Execute the following steps until the number of key image features stored according to the target object type is greater than or equal to a quantity threshold, where the key image features are used to update the standard visual features belonging to the target object type: Determine a current image feature from the second image features corresponding to the respective image blocks; Calculate a fourth distance between other image features except the current image feature and the current image feature among the second image features corresponding to the respective image blocks in the feature space, where the distance in the feature space is inversely related to the similarity between features; Determine the second image feature with the largest fourth distance as the key image feature; Store the obtained key image features according to the target object type; When the number of the obtained key image features is less than the number threshold, determine a next image feature from the second image features corresponding to the respective image blocks, and use the next image feature as the current image feature.
7. The method according to any one of claims 1 to 6, after performing first-stage feature extraction on the first image to obtain at least one basic visual feature, the method further includes: When the feature similarity between each standard visual feature among at least one standard visual feature and the basic visual feature is less than the preset threshold, set the object type to which the object belongs to an additional object type, where the additional object type is a type other than the rated object types to which the multiple sample images belong; Collect standard object objects of the additional object type to obtain a plurality of third images, and use the third images as the sample images.
8. An anomaly recognition device, comprising: A first acquisition unit configured to acquire a first image to be recognized, where the first image is an image acquired of an object; A first extraction unit configured to perform first-stage feature extraction on the first image to obtain at least one basic visual feature, where the basic visual feature is used to represent the visual attributes presented on the object surface, and the object surface includes the surface of the object; A first determination unit configured to determine, from at least one standard visual feature, a first visual feature whose feature similarity with the basic visual feature is greater than or equal to a preset threshold, where the standard visual feature is a feature obtained by performing feature extraction on multiple sample images belonging to multiple object types, and the sample image corresponding to the first visual feature is a second image belonging to the target object type; A second acquisition unit configured to acquire target guidance information matching the first visual feature, where the target guidance information is used to guide attention to features related to the target object type in the image; A second extraction unit configured to perform second-stage feature extraction on the first image through the target guidance information to obtain at least one second visual feature related to the target object type; A first comparison unit configured to compare the at least one second visual feature with the standard visual feature of the second image to obtain a comparison result; A second determination unit configured to determine an anomaly recognition result of the first image through the comparison result.
9. A computer-readable storage medium, the computer-readable storage medium including a stored computer program, wherein, The computer program, when run on an electronic device, executes the method according to any one of claims 1 to 7.
10. A computer program product, comprising a computer program, which when executed by a processor, implements the steps of the method according to any one of claims 1 to 7.
11. An electronic device, comprising a memory and a processor, where a computer program is stored in the memory, and the processor is configured to execute the method according to any one of claims 1 to 7 through the computer program.
Citation Information
Patent Citations
Abnormal identification method, device, storage medium and electronic device
CN117788918B
Visual attribute recognition method and device and storage medium
CN109447095A
Target anomaly recognition method and device, electronic equipment and storage medium
CN111652940A
Data classification method, computer and readable storage medium
CN113516146A
Abnormality determination method and device, electronic equipment and storage medium
CN116580301A
Cited By
Building decorative plate production quality rapid detection method based on machine vision
CN120707554A