Abnormal image recognition method and device, computer equipment and readable storage medium
By locally counting the number of key point features in images on the learning machine terminal and performing preliminary screening, combined with a multi-dimensional anomaly recognition strategy, the problem of server computing pressure and network bandwidth consumption caused by abnormal images during image upload to the learning machine was solved, thereby improving the stability and efficiency of the system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING CENTURY TAL EDUCATION TECH CO LTD
- Filing Date
- 2025-12-25
- Publication Date
- 2026-04-17
AI Technical Summary
In existing technologies, the large number of abnormal images containing non-required content during the image upload process increases the computational load of the homework grading system, leading to increased server computational pressure, excessive network bandwidth consumption, and affecting system operating efficiency and stability.
By counting the number of key features in images locally on the learning machine terminal, abnormal images are initially screened out. When necessary, multi-dimensional anomaly recognition or text quantity recognition strategies are adopted to ensure that only high-quality images are processed in the cloud.
It effectively reduced the upload of abnormal images, lowered the computing pressure on the cloud server, stabilized system performance, reduced network bandwidth usage, avoided system response delays or crashes, and ensured the reliability and efficiency of the service.
Smart Images

Figure CN121884083A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to an abnormal image recognition method, apparatus, computer device, and readable storage medium. Background Technology
[0002] In the field of modern education, learning machines have become an important tool to assist students in their learning, and are widely used in various aspects such as homework completion and knowledge consolidation. Students often use learning machines to take pictures of the contents of book pages and upload the images to the homework grading system to achieve automated homework grading. At the same time, in scenarios such as page turning detection, it is also necessary to take pictures and upload images to complete the detection process, thereby improving learning efficiency and promoting the digital and intelligent application of educational resources.
[0003] In related technologies, when learning machines capture and upload images to the cloud system, all images—whether book content, desktop environment, portraits, blank paper, covers, drafts, hand-drawn pictures, or images that are blurry, truncated, or severely obscured—are uploaded to the homework grading system or used for page-turning detection. The applicant recognizes that the upload of a large number of abnormal images containing irrelevant content significantly increases the computational load on the homework grading system, especially in large-scale usage scenarios, such as when many students are learning online simultaneously. This indiscriminate uploading puts significant computational pressure on the cloud server, severely impacting system efficiency. Furthermore, in page-turning detection scenarios, the upload of abnormal images can lead to non-standard algorithm inputs and instability, resulting in frequent image uploads. This not only exacerbates the server's computational burden but also easily saturates network bandwidth, affecting normal system service and even causing system response delays or crashes. Summary of the Invention
[0004] In view of this, this application provides an abnormal image recognition method, apparatus, computer device and readable storage medium, the main purpose of which is to solve the problems that currently increase the computing burden on the server side, easily cause network bandwidth to be fully occupied, affect the normal service of the system, and even lead to system response delay or crash.
[0005] According to a first aspect of this application, an abnormal image recognition method is provided, the method comprising: The target image to be identified is determined, the number of key point features included in the target image is counted, and the target image is determined to be an abnormal image based on the number of features. If it is determined that the target image is not an abnormal image, a target recognition strategy to be executed is determined, and the target recognition strategy indicates to perform multi-dimensional anomaly recognition or image text quantity recognition. The target image is identified according to the target recognition strategy to determine whether the target image is an abnormal image.
[0006] According to a second aspect of this application, an abnormal image recognition device is provided, the device comprising: The feature point filtering module is used to determine the target image to be identified, count the number of key point features included in the target image, and determine whether the target image is an abnormal image based on the number of features. The identification strategy determination module is used to determine the target identification strategy to be executed when it is determined that the target image is not an abnormal image. The target identification strategy indicates to perform multi-dimensional anomaly identification or image text quantity identification. The image recognition module is used to recognize the target image according to the target recognition strategy in order to determine whether the target image is an abnormal image.
[0007] According to a third aspect of this application, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the method described in any of the first aspects above.
[0008] According to a fourth aspect of this application, a readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described in any one of the first aspects above.
[0009] By employing the above technical solutions, this application provides an abnormal image recognition method, apparatus, computer equipment, and readable storage medium. This application effectively filters out irrelevant abnormal images by statistically analyzing the number of key point features in the target image and making preliminary anomaly judgments based on this. This reduces the upload of a large number of abnormal images at the source. Furthermore, based on this, targeted recognition strategies, such as multi-dimensional anomaly or text quantity recognition, are applied only to the preliminarily determined non-abnormal images. This prevents problems such as non-standard algorithm input and frequent uploads caused by abnormal images, avoids the waste of computing resources caused by indiscriminately processing all images, significantly reduces the computing pressure on cloud servers, stabilizes system performance, reduces server burden and network bandwidth usage, reduces the risk of system response delays or crashes, and ensures the reliability and efficiency of the service.
[0010] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description
[0011] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of this application. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings: Figure 1 This paper illustrates a flowchart of an abnormal image recognition method provided in an embodiment of this application. Figure 2 This paper illustrates a schematic flowchart of another abnormal image recognition method provided in an embodiment of this application. Figure 3 This paper shows a schematic diagram of the structure of an abnormal image recognition device provided in an embodiment of this application; Figure 4 A schematic diagram of the device structure of a computer device provided in an embodiment of this application is shown. Detailed Implementation
[0012] Exemplary embodiments of the present application will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this application will be thorough and complete, and will fully convey the scope of the present application to those skilled in the art.
[0013] This application provides an abnormal image recognition method, such as... Figure 1 As shown, the method includes: S10: Determine the target image to be identified, count the number of key point features included in the target image, and determine whether the target image is an abnormal image based on the number of features.
[0014] The technical solutions in this application can be applied to terminals, which can be auxiliary learning tools such as learning machines or learning computers, integrating a processor and a display screen, and supporting multiple interaction methods, including touch and voice input, to adapt to the usage habits of students of different ages. The terminal in this application has a built-in image processing algorithm that can interface with educational applications such as homework grading systems and e-book reading systems, enabling real-time analysis and processing of captured images. For example, when grading homework, students only need to take a picture of the homework page with the learning machine, and the terminal can automatically identify and filter out images that do not meet the requirements, ensuring that only high-quality images enter the grading process; when reading e-books, the terminal can also capture page changes in real time through its camera, accurately determine whether page turning has occurred, and optimize the displayed content to provide a seamless reading experience.
[0015] In this embodiment, when the camera of a terminal such as a learning machine or learning computer captures an image taken by the user, the image is identified as the target image to be recognized. The terminal first extracts key point features from the target image. Key point features refer to local feature points in the image with significant identifiability, such as corner points, edge intersections, and areas with drastic texture changes. Specifically, algorithms such as SIFT (Scale Invariant Feature Transform), SURF (Speed-Up Robust Feature Transform), or ORB (Oriented Fast and Rotated BRIEF) can be used to detect and extract these key point features. After extraction, the terminal counts the number of these key point features to obtain a specific numerical value, i.e., the number of features.
[0016] It's important to note that the process of counting key points is completed locally on the terminal, eliminating the need to upload the original target image to a cloud server, thus reducing data transmission at the source. This simple and efficient key point counting allows for the rapid identification of images that clearly do not meet the requirements, such as images with solid color backgrounds, completely blurred images, or severely occluded images. These images typically have very few key points or abnormally distributed key points, thereby filtering out a large number of low-quality images in the earliest stage, reducing the burden of subsequent processing, and providing initial screening for later steps.
[0017] For example, in a homework grading scenario, when a student takes a picture of a page of math homework, the terminal first extracts the key features of the image. Assuming that a normal homework image usually contains a lot of text, formulas and graphics, it will generate 500-2000 key points; however, if the student accidentally takes a picture of a blank sheet of paper or a desktop, the number of key points may be less than 1. Based on this, the terminal can directly identify this abnormal situation and avoid further processing or uploading such meaningless images to the cloud server.
[0018] After obtaining the number of features, the terminal determines whether the target image is an abnormal image based on this value. Specifically, the number of features can be compared with a preset minimum threshold to determine whether the target image is abnormal. This minimum threshold is an empirical value derived from training with a large number of samples, used to indicate the minimum number of keypoint features a normal image should include. For example, in a homework grading scenario, the minimum threshold can be set based on the distribution of keypoint features on a normal homework page. In practical applications, the statistically obtained number of keypoint features can be compared with the set minimum threshold. If the number of features is less than the minimum threshold, the image content is considered too simple or invalid, and can be judged as an abnormal image; conversely, if the number of features is greater than or equal to the minimum threshold, the image is considered to potentially contain valid information and requires further verification.
[0019] In this way, a lightweight yet efficient preliminary filtering mechanism is established based on the number of features. This mechanism can quickly eliminate obviously non-compliant images, preventing them from consuming subsequent computing resources and network bandwidth, and improving the accuracy and adaptability of the filtering. For example, in page-turning detection applications, when students are reading on a learning machine, the terminal continuously captures page images. If the page is severely reflective or the image is extremely blurry, the number of key points will be below the minimum threshold. The terminal can identify these situations and avoid sending these abnormal images to the page-turning detection algorithm. This prevents frequent false triggering and repeated uploading of the algorithm due to non-standard input, maintaining the stability and response speed of the page-turning detection function.
[0020] S20: If it is determined that the target image is not an abnormal image, determine the target recognition strategy to be executed.
[0021] In this embodiment, if the target image is determined not to be an abnormal image, the terminal will further analyze its content quality. Specifically, the terminal will dynamically select the most suitable recognition strategy based on the preliminary feature analysis results of the image. The target recognition strategy indicates two strategies: one is to perform multi-dimensional anomaly recognition, and the other is to indicate an image text quantity recognition strategy. Multi-dimensional anomaly recognition is mainly used to perform multi-dimensional detection on the image, such as using multiple methods to identify whether the image belongs to an anomaly category, whether there are fingers obscuring the image, whether there are shadows covering it, whether it is blurry, etc. The image text quantity recognition strategy mainly counts the number and distribution density of identifiable characters in the image to determine whether it is a valid book page or homework image.
[0022] In this way, the image quality can be further evaluated through the above process, avoiding the waste of computing resources caused by indiscriminately processing all images. At the same time, the processing efficiency and accuracy under different conditions are optimized by dynamically selecting strategies, reducing unnecessary computation and transmission.
[0023] S30: According to the target recognition strategy, the target image is identified to determine whether the target image is an abnormal image.
[0024] In this embodiment, the terminal analyzes the target image to make a final judgment based on the determined target recognition strategy. Specifically, if a multi-dimensional anomaly recognition strategy is selected, the terminal will use a pre-trained model or a preset algorithm to recognize the image to determine whether the target image is an abnormal image. If an image text quantity recognition strategy is selected, the terminal will recognize the number of characters included in the target image to determine whether the target image is an abnormal image.
[0025] When an image is identified as abnormal, the terminal will discard it locally and prompt the user to retake the image. If the image is determined not to be abnormal, it will be allowed to enter the subsequent job grading or page-turning detection process and uploaded to the cloud server. The cloud server will then perform the corresponding operation process for the target image, ensuring that only high-quality, compliant images enter the subsequent processing or are uploaded to the cloud. This effectively prevents non-standard algorithm input and frequent uploads caused by abnormal images, significantly reduces the computing pressure on the cloud server, stabilizes overall performance, avoids the risk of system response delays or service interruptions due to resource overload, and protects user privacy, as most sensitive image data does not need to leave the local device.
[0026] For example, during online homework grading, a student took a picture of their English homework where part of it was obscured by a finger. Although the number of key points met the requirements, the system detected that the finger was covering most of the answer area, classifying the image as abnormal. The student was then prompted to retake the picture. This effectively prevents the low-quality image from being uploaded to the cloud, avoids invalid processing of the image by the grading system and potential erroneous feedback, reduces the pressure on the cloud server, ensures the stability and response speed of the overall service, and allows other students' homework to be graded in a timely manner.
[0027] The method provided in this application embodiment can effectively filter out irrelevant abnormal images by statistically analyzing the number of key point features in the target image and making preliminary anomaly judgments based on this. This reduces the upload of a large number of abnormal images at the source. Furthermore, based on this, targeted recognition strategies such as multi-dimensional anomaly recognition or text quantity recognition are applied only to the preliminarily determined non-abnormal images. This prevents problems such as non-standard algorithm input and frequent uploads caused by abnormal images, avoids the waste of computing resources caused by indiscriminate processing of all images, significantly reduces the computing pressure on the cloud server, stabilizes system performance, reduces server burden and network bandwidth occupation, reduces the risk of system response delays or crashes, and ensures the reliability and efficiency of the service.
[0028] In this embodiment of the application, optionally, determining the target image to be identified, counting the number of key point features included in the target image, and determining whether the target image is an abnormal image based on the number of features, includes: determining a preset feature type, identifying the features included in the target image to extract multiple key point features of the same type as the feature type from the target image; counting the number of multiple key point features to obtain the number of features; obtaining a preset minimum number threshold, comparing the number of features with the minimum number threshold; when the comparison determines that the number of features is greater than or equal to the minimum number threshold, determining that the target image is not an abnormal image; when the comparison determines that the number of features is less than the minimum number threshold, determining that the target image is an abnormal image; correspondingly, the method further includes: generating and outputting an abnormal image alert when the target image is determined to be an abnormal image, and ending the current image recognition process.
[0029] In this embodiment, firstly, the terminal presets feature types, which can be any of SIFT (Scale Invariant Feature Transform), SURF (Speed Robust Feature Transform), or ORB (Oriented Fast and Rotated BRIEF). These features are commonly used in image processing to describe the local characteristics of an image. Specifically, the terminal scans the target image to be identified, identifies and extracts multiple key point features that match the preset feature types. For example, in a homework grading scenario, the image uploaded by the student should contain clear text or question content, which will be represented by a series of stable key point features in the image.
[0030] Subsequently, the terminal counts the number of extracted keypoint features to obtain the feature count. Simultaneously, the terminal presets a minimum feature count threshold, which is set based on extensive experimental data and real-world application scenarios to distinguish between normal and abnormal images. During the comparison phase, the terminal compares the feature count with the minimum feature count threshold. If the feature count is greater than or equal to the minimum feature count threshold, the target image is determined not to be abnormal, and subsequent recognition processes can continue; otherwise, it is determined to be an abnormal image.
[0031] In addition, for images identified as abnormal, the terminal will generate and output an abnormal image alert, informing the user that the uploaded image has a problem and cannot be effectively processed, and will terminate the current image recognition process. This not only improves the level of automation but also enhances the user experience and avoids the waste of resources caused by ineffective processing.
[0032] Taking homework grading as an example, suppose a student uploads an image containing numerous handwritten notes, but key text areas are obscured. Key point feature recognition reveals that the number of key point features in the image is far below a preset threshold, thus classifying it as an abnormal image and immediately prompting the student to re-upload a clear and complete image. Similarly, in page-turning detection, if a student uploads a blurry or blank page, a similar mechanism can be used to quickly identify and process it, ensuring the accuracy and efficiency of page-turning detection.
[0033] In practical applications, a feature point filtering module can be set up in the terminal to realize the above-mentioned key point feature recognition, feature quantity statistics and threshold comparison.
[0034] In this embodiment of the application, optionally, determining the target recognition strategy to be executed includes: determining a preset standard quantity threshold, comparing the number of features with the standard quantity threshold, wherein the standard quantity threshold is greater than a minimum quantity threshold; when the comparison determines that the number of features is greater than or equal to the standard quantity threshold, determining the target recognition strategy that indicates multi-dimensional anomaly recognition; when the comparison determines that the number of features is less than the standard quantity threshold, determining the target recognition strategy that indicates image text quantity recognition.
[0035] In this embodiment, a preset standard number threshold is first determined. This standard number threshold is set based on a comprehensive analysis of a large amount of experimental data and real-world application scenarios. In practical applications, the standard number threshold will be greater than the minimum number threshold required to ensure basic recognition accuracy, so as to provide sufficient sample support for more complex recognition tasks. Next, the terminal compares the number of features with the standard number threshold. When the comparison result shows that the number of features is greater than or equal to the standard number threshold, it only indicates that a large number of key point features have been identified, meaning that the target image is likely to contain a large amount of text information. Therefore, no additional text recognition is needed, and multi-dimensional anomaly recognition is performed.
[0036] When the number of features determined by comparison is less than the standard threshold, it means that the target image does not contain much text content. Further determination is needed to confirm whether there is really no large amount of text information in the target image. Therefore, in this case, the terminal will switch to the target recognition strategy that indicates the text quantity recognition of the image, thereby focusing on the text region in the image. By accurately recognizing the number of texts or characters, it can determine whether the image meets the expected text content requirements, so as to improve the recognition speed, reduce unnecessary computing resource consumption, and improve the overall operating efficiency.
[0037] Optionally, in this embodiment, the target image is identified according to a target recognition strategy to determine whether the target image is an anomalous image. This includes: when the target recognition strategy indicates that multi-dimensional anomaly recognition is to be performed, determining a multi-dimensional anomaly recognition node for multi-dimensional anomaly recognition, wherein the multi-dimensional anomaly recognition node is used to call a multi-classification filtering model, a binary classification filtering model, a fuzzy region recognition algorithm, and an occlusion detection model to perform multi-dimensional anomaly recognition on the target image; based on the multi-dimensional anomaly recognition node, the target image is identified to obtain classification recognition results, fuzzy region recognition results, and occlusion detection results; when the classification recognition results, fuzzy region recognition results, and occlusion detection results all indicate that the target image is not an anomalous image, the target image is determined to be not an anomalous image; when any of the classification recognition results, fuzzy region recognition results, and occlusion detection results indicate that the target image is an anomalous image, the target image is determined to be an anomalous image.
[0038] In this embodiment, when the target recognition strategy indicates the need for multi-dimensional anomaly recognition, the terminal identifies and invokes a multi-dimensional anomaly recognition node. This node integrates a multi-classification filtering model, a binary filtering model, a blurred region recognition algorithm, and an occlusion detection model to perform comprehensive and detailed anomaly recognition on the target image. In practical applications, the multi-dimensional anomaly recognition node can be a physical machine within the same local area network as the terminal. Due to the limited computing power of the terminal, some computationally intensive models or algorithms cannot run on mobile devices because the terminal's relatively poor computing power causes these algorithms to run slowly or fail. Therefore, a physical machine can be set up as the multi-dimensional anomaly recognition node. Most anomaly images can be filtered on the physical machine. For example, a classification filtering model can be invoked on the physical machine to filter out anomalies such as book covers, scribbled drafts, hand-drawn pictures, blurred images, severely truncated pages or homework images, and severely occluded pages or homework images. In practical applications, an abnormal image judgment module can be set up on the physical machine. This module integrates multi-classification filtering models and binary classification filtering models to classify and filter target images. Furthermore, a fuzzy image judgment module can be set up on the physical machine, integrating a fuzzy region recognition algorithm to identify fuzzy regions in the target image. Additionally, an occlusion judgment module can be set up on the physical machine, integrating an occlusion detection model to determine whether occlusions exist in the target image. This allows the terminal to call relevant models or algorithms to recognize target images solely through the physical machine, reducing the terminal's workload. Moreover, the fact that the terminal and physical machine operate on the same local area network ensures rapid data exchange.
[0039] Specifically, the multi-class filtering model is responsible for dividing images into multiple predefined categories (such as normal book pages, desktops, portraits, blank paper, etc.), using a trained deep learning model to perform high-level understanding and classification of image content. The binary filtering model further focuses on distinguishing between normal and abnormal images, providing more direct judgment results. Meanwhile, the blurred region recognition algorithm can be the Laplacian blur judgment algorithm to assess image sharpness and effectively identify image regions that are difficult to discern due to blurry images. The occlusion detection model can be the YOLOv8 model, which features fast detection speed and high accuracy, and can quickly and accurately detect the position of hands in images.
[0040] In practical applications, the terminal performs parallel processing on the target image based on multi-dimensional anomaly detection nodes, obtaining classification recognition results, blurred region recognition results, and occlusion detection results. Only when all recognition results indicate that the target image is not an anomaly will the terminal ultimately determine that the image is not an anomaly and allow it to enter the subsequent job grading or page-turning detection process. Conversely, if any recognition result indicates that the image is abnormal, it will be immediately marked as an anomaly image, and the corresponding anomaly handling mechanism will be triggered, such as prompting the user to re-upload the image or skipping the processing of the image.
[0041] Taking homework grading as an example, suppose a student uploads a blurry image where part of the content is obscured by a hand. Through a multi-dimensional anomaly detection mechanism, the obstruction and blurry areas in the image can be identified. The obstruction detection model and blurry area recognition algorithm provide clear identification results respectively. Since at least one of these results indicates the image is abnormal, it will be determined that the image cannot be used for homework grading, thus avoiding grading errors or omissions caused by image quality issues and significantly improving the accuracy and efficiency of homework grading. Similarly, in page-turning detection scenarios, this mechanism can effectively filter out images that do not meet the requirements, ensuring the smooth progress of page-turning detection.
[0042] It should be noted that this application embodiment adds a blur judgment module and an occlusion judgment module. The reason for adding these two separate modules is that these two types of abnormal images account for the largest proportion of all abnormal images and are the most difficult to process. In homework grading scenarios, students may accidentally obscure part of the content with their fingers while taking pictures of their homework, or the image may be blurry due to poor shooting conditions. These situations can seriously affect the accuracy and efficiency of homework grading. For example, if the image is blurry, the grading system may not be able to accurately recognize the text written by the student, thus giving incorrect grading results; if the image has severe hand occlusion, the grading system may not be able to fully obtain the homework information, resulting in incomplete grading. The same applies to page-turning detection scenarios. Blurry images may cause the system to misjudge the page content, and hand occlusion may prevent the system from accurately recognizing page-turning actions. Therefore, through the separate setting and accurate judgment of these two modules, these two types of abnormal images can be identified in a timely manner, reminding the user to retake a clear image, or the system can automatically take measures such as image enhancement and repair to improve image quality, ensuring the accuracy and stability of intelligent education applications such as homework grading and page-turning detection, and providing reliable technical support for intelligent education.
[0043] Optionally, in this embodiment, the method further includes: collecting and constructing an image dataset for model training, the image dataset including multiple types of abnormal sample images and invalid sample images; performing image size standardization and image data standardization on each sample image included in the image dataset according to a first preset resolution to obtain a first training dataset; and performing image size standardization and image data standardization on each sample image included in the image dataset according to a second preset resolution to obtain a second training dataset, wherein the second preset resolution is greater than the first preset resolution; determining a multi-classification model architecture, training the multi-classification model architecture using the first training dataset to obtain and store a multi-classification filtering model, so as to obtain and call the multi-classification filtering model when performing abnormal image recognition; determining a binary classification model architecture, training the binary classification model architecture using the second training dataset to obtain and store a binary classification filtering model, so as to obtain and call the binary classification filtering model when performing abnormal image recognition.
[0044] In this embodiment, firstly, an image dataset containing various types of abnormal and invalid sample images needs to be collected and constructed. The sample images in the image dataset cover various abnormal situations that students may encounter during actual learning, such as desktops, portraits, blank paper, covers, scribbled drafts, hand-drawn pictures, blurry images, severely truncated pages or homework images, and severely obscured pages, ensuring the diversity and representativeness of the image dataset. Next, to adapt to the input image resolution requirements of different model architectures, this embodiment performs detailed image size standardization and image data standardization processing on each sample image in the dataset according to a first preset resolution, forming a first training dataset. Specifically, considering the balance between model inference speed and image quality, in practical applications, the sample images can be uniformly scaled to a fixed size of 240*240 to obtain the first training dataset. This maintains image clarity while effectively compressing image data, thereby improving the model's inference efficiency. Furthermore, standardizing the image data ensures that all input data has a consistent data distribution, which is beneficial for the model to learn stable image features.
[0045] Simultaneously, the dataset was processed in the same way according to the second preset resolution, resulting in a second training dataset specifically used for training the binary classification model architecture. The second preset resolution is greater than the first preset resolution; higher-resolution image data helps the binary classification model capture more subtle image features, thereby improving its ability to distinguish between normal and abnormal images. In practical applications, the second preset resolution can be 320*320, that is, uniformly scaling the sample images to a fixed size of 320*320. This ensures image clarity while effectively compressing the data volume, thereby improving the inference speed of the trained binary classification filtering model. Furthermore, standardizing the image data eliminates differences in brightness, contrast, etc., between different images, allowing the model to focus more on the essential features of the image content.
[0046] During the model training phase, this embodiment first determines the multi-classification model architecture. Using a first training dataset, the multi-classification model is fully trained. By continuously adjusting the model parameters and optimizing the algorithm, a high-performance multi-classification filtering model is finally obtained and stored for later use. Similarly, for the binary classification model architecture, a similar training process is performed using a second training dataset to obtain a binary classification filtering model that can efficiently distinguish between normal and abnormal images, which is also stored.
[0047] For example, during homework grading, the terminal first uses a multi-class filtering model to identify various abnormal features in the image, such as whether it is a blank page or whether it is occluded. Then, it combines the judgment results of a binary filtering model to comprehensively determine whether the image is a valid homework image, thus improving the accuracy of homework grading. Similarly, in page-turning detection scenarios, the terminal can use these two models to quickly and accurately identify page images and eliminate interference from other irrelevant images, thereby ensuring the stability and reliability of page-turning detection.
[0048] Optionally, in this embodiment, determining the multi-class classification model architecture includes: obtaining a basic multi-class classification model architecture, which includes multiple neurons; adding a first segmented activation function to the fully connected layer in the basic multi-class classification model architecture, and setting the first activation threshold of the first segmented activation function to 1 to obtain the multi-class classification model architecture. The fully connected layer in the multi-class classification model architecture includes a regular activation function and a first segmented activation function. During model operation, the regular activation function is used to calculate the activation value of the received input value once. If the calculation result obtained from the first activation value calculation is less than the first activation threshold, the first segmented activation function is used to calculate the activation value a second time.
[0049] In this embodiment, to improve the accuracy and generalization ability of the multi-classification model in the abnormal image recognition task, it is necessary to improve and optimize the basic multi-classification model architecture and its activation mechanism. Specifically, it is first necessary to select an efficient basic model architecture suitable for image classification tasks. In this embodiment, MobileNetV3 Small is selected as the basic model architecture. This model performs well on mobile and embedded devices due to its lightweight and efficient performance. In addition, in practical applications, other model architectures can be flexibly selected according to specific needs and resource constraints. This application does not impose specific limitations on this.
[0050] Before model training, this application primarily improves the activation mechanism of neurons in the fully connected layer of the basic multi-class classification model architecture. Traditionally, classification models often use a single activation function, such as RelU, to introduce non-linear factors and enhance the model's expressive power. However, a single activation function may have limitations when processing complex image features, failing to fully capture all key features. Therefore, this application proposes a dual activation function mechanism to improve the neurons in the penultimate layer of the fully connected layer. Specifically, the fully connected layer includes a conventional activation function. During actual model operation, the conventional activation function (such as RelU) is first used to calculate the neurons in the penultimate layer to obtain preliminary activation values. In this application embodiment, a first segmented activation function is added to the fully connected layer for secondary activation, and the first activation threshold of the first segmented activation function is set to 1 to obtain the multi-class classification model architecture. In this multi-classification model architecture, the first segmented activation function is set to perform a second activation value calculation on the calculation result obtained by the first activation value calculation when the activation value of the neuron is less than a certain first activation threshold (i.e., 1), and set the value of the neuron to 0, that is, suppress the contribution of the neuron to the final classification result; while when the activation value is greater than or equal to the first activation threshold, the original activation value remains unchanged, that is, the effective features extracted by the neuron are fully utilized.
[0051] In this way, the dual activation function mechanism can highlight key features in the image and suppress noise and irrelevant features, thereby improving the accuracy and generalization ability of classification. Just as the human eye mainly relies on salient features such as wings when recognizing birds, while ignoring features that are not obvious on birds, such as fins, this multi-classification model can accurately identify non-assignment content images such as desktops and portraits, as well as abnormal images such as blank paper images, covers, scribbles, hand-drawn pictures, blurry images, severely truncated pages or assignment images, and severely occluded pages in practical applications such as homework grading. This effectively avoids these illegal images from interfering with the homework grading system and improves the accuracy and efficiency of grading. Similarly, in page-turning detection scenarios, this model can accurately identify blurry, occluded, or severely truncated page images that appear during page turning, ensuring the stability and reliability of page-turning detection.
[0052] Optionally, in this embodiment, determining the binary classification model architecture includes: obtaining a basic binary classification model architecture, which includes two neurons; adding a second segmented activation function to the fully connected layer in the basic binary classification model architecture, and setting the second activation threshold of the second segmented activation function to 2 to obtain the binary classification model architecture. The fully connected layer in the binary classification model architecture includes a regular activation function and a second segmented activation function. During model operation, the regular activation function is used to calculate the activation value of the received input value once. If the calculation result obtained from the first activation value calculation is less than the second activation threshold, the second segmented activation function is used to calculate the activation value a second time.
[0053] In this embodiment, in order to train the binary classification filtering model, during the process of generating the second training dataset, non-book page images such as desktop or portrait images, blank paper images, covers, scribbled drafts, hand-drawn drawings, blurry images, severely truncated book pages or homework images, and severely occluded book pages or homework images can be classified into one category, namely the abnormal image category; while normal, clear and complete book page or homework images are classified into another category, namely the normal image category, thereby constructing a binary classification task.
[0054] In this embodiment, the ShuffleNetV2 model is selected as the basic binary classification model architecture. ShuffleNetV2 is lightweight and high-performance, making it suitable for operation in resource-constrained environments. Furthermore, in practical applications, other model architectures can be flexibly selected based on specific needs and resource availability; this application does not impose specific limitations on this. It should be noted that the binary classification model and the multi-class classification model in this embodiment employ different architectures to increase the richness of features extracted by the model, enabling it to determine whether the current user image is normal from different dimensions and perspectives. Moreover, as mentioned above, the image size input to the binary classification model (320*320) is larger than that of the multi-class classification model, which helps the binary classification model capture more subtle feature differences, thereby improving classification accuracy.
[0055] Similar to multi-class classification models, this application embodiment also improves the activation function of the penultimate layer neurons in the output layer for calculating the confidence scores of various classes in the basic binary classification model architecture. Traditionally, classification models often use a single activation function, such as RelU, to introduce non-linear factors. However, in order to further improve the accuracy and generalization ability of the classification model, this application embodiment adopts a dual activation function mechanism. Specifically, for neurons in the penultimate layer of the basic binary classification model architecture, a conventional activation function is preset so that a preliminary activation value is calculated using a conventional activation function (such as RelU) during model operation. In this embodiment, a second segmented activation function is added to the fully connected layer in the basic binary classification model architecture, and the second activation threshold of the second segmented activation function is set to 2 to obtain the binary classification model architecture. In this way, after obtaining the preliminary activation value during model operation, the second segmented activation function can be used to calculate the activation value of these neurons again. The second activation threshold of the second segmented activation function is set to 2, that is, when the value of the neuron is greater than or equal to the threshold 2, the original value remains unchanged; when it is less than the threshold 2, the value of the neuron is set to 0. This preserves the feature value distribution extracted by the original activation function and suppresses smaller feature values from participating in the final classification calculation, thereby making full use of neurons with larger feature values for classification and improving the classification accuracy and generalization ability.
[0056] Thus, employing a dual activation function leverages the advantages of both. On one hand, the conventional activation function extracts the basic feature distribution in the image; on the other hand, the piecewise activation function highlights strong features and suppresses weak features, allowing the model to focus more on features crucial to classification decisions. This is similar to how the human eye, when recognizing objects, only needs to focus on salient features to make a judgment, without considering all subtle details. Therefore, the dual activation function mechanism enables the model to complete classification tasks more efficiently and accurately. Taking homework grading as an example, this binary classification model can accurately identify various illegal images, such as desktop backgrounds and human figures, ensuring the accuracy and efficiency of the grading system. In page-turning detection scenarios, the model can also stably distinguish between normal and abnormal pages, providing reliable support for intelligent education applications.
[0057] Optionally, based on a multi-dimensional anomaly detection node, the target image is identified to obtain classification results, fuzzy region detection results, and occlusion detection results. This includes: using a multi-classification filtering model and a binary filtering model to classify and filter the target image based on the multi-dimensional anomaly detection node; determining the classification results based on the first classification filtering result output by the multi-classification filtering model and the second classification filtering result output by the binary filtering model; simultaneously, using a fuzzy region detection algorithm to identify fuzzy regions in the target image based on the multi-dimensional anomaly detection node to obtain fuzzy region detection results; and simultaneously using an occlusion detection model to identify occlusions in the target image based on the multi-dimensional anomaly detection node to obtain occlusion detection results.
[0058] In this embodiment, firstly, based on the multi-dimensional anomaly detection node, both a multi-classification filtering model and a binary filtering model are simultaneously invoked to classify and filter the target image. The multi-classification filtering model uses deep learning technology to divide the image into multiple specific categories, such as normal book pages, covers, and scribbled drafts, and outputs the first classification filtering result; while the binary filtering model focuses on distinguishing whether the image is legal and outputs the second classification filtering result. The multi-dimensional anomaly detection node combines the outputs of these two models to determine the final classification and recognition result, thereby effectively identifying possible category errors or illegal content in the image.
[0059] Meanwhile, the multi-dimensional anomaly identification node also employs a fuzzy region identification algorithm to detect fuzzy regions in the target image. This algorithm can be the Laplacian fuzziness judgment algorithm, which is a method to evaluate image sharpness by calculating the second derivative (or variance) of an image using the Laplacian operator. By performing a Laplacian transform on the image, the rate of change of pixel values can be measured. For sharp images, the edges have large gradient changes, and the variance of the second derivative (i.e., the response of the Laplacian operator) will be large. Conversely, for blurry images, the edges are not obvious, the gradient changes are small, and the variance of the Laplacian operator response will be low. This helps to filter out images that cannot be clearly identified due to poor shooting quality, thereby improving the accuracy of subsequent assignment grading or page turning detection.
[0060] In addition, the multi-dimensional anomaly recognition node will also call the occlusion detection model to identify occlusions in the target image. The occlusion detection model can be the YOLOv8 model, which has the characteristics of fast detection speed and high accuracy. It can quickly and accurately detect the position of the hand in the image and determine whether it has caused substantial occlusion of the image content. This allows images with incomplete content due to occlusion to be excluded, ensuring that only clear and complete images can enter the subsequent processing flow.
[0061] Taking homework grading as an example, suppose a student uploads an image containing blurry text and a hand obscuring the work. Through multi-dimensional anomaly detection nodes, the system can simultaneously obtain classification results (e.g., identifying the image as a "scribbled draft"), blurry area recognition results (e.g., detecting blurry text areas in the image), and obstruction detection results (e.g., locating a hand obscuring part of the homework content). Based on these combined results, the image can be accurately determined to be an anomaly, prompting the student to re-upload a clear and complete image. This effectively avoids grading errors or omissions caused by image quality issues, significantly improving the accuracy and efficiency of homework grading. Similarly, in page-turning detection scenarios, this mechanism ensures that only images meeting the requirements can be used for page-turning judgment, guaranteeing the smooth progress of page-turning detection.
[0062] Optionally, in this embodiment, the classification and identification result is determined based on the first classification filtering result output by the multi-class filtering model and the second classification filtering result output by the binary filtering model. This includes: generating a classification and identification result indicating that the target image is an abnormal image when both the first classification filtering result and the second classification filtering result indicate that the target image belongs to the abnormal image category; and generating a classification and identification result indicating that the target image is not an abnormal image when either the first classification filtering result or the second classification filtering result indicates that the target image does not belong to the abnormal image category.
[0063] In this embodiment, the first classification filtering results from the multi-class filtering model and the second classification filtering results from the binary filtering model are processed in parallel. The multi-class filtering model, with its feature extraction and classification capabilities, can meticulously distinguish the specific category to which an image belongs, such as a normal book page, a cover, or a scribbled draft, while the binary filtering model focuses on quickly distinguishing between the legality and illegality of an image.
[0064] During the judgment process, the first step is to check whether both the first and second classification filtering results indicate that the target image belongs to the abnormal image category. If both indicate that the image is abnormal, a classification recognition result indicating that the target image is an abnormal image will be generated. This ensures a rapid response to potential problems in the image and avoids interference with subsequent job grading or page-turning detection processes due to image quality issues. Conversely, if either the first or second classification filtering result indicates that the target image does not belong to the abnormal image category, a classification recognition result indicating that the target image is not an abnormal image will be generated. Thus, a comprehensive judgment is made based on the output results of the two models, effectively improving the accuracy and reliability of classification recognition.
[0065] Taking homework grading as an example, suppose a student uploads an image that contains both normal homework content and slight doodles. A multi-class filtering model might identify the image as containing "scribbled drafts," while a binary filtering model might determine it as a legitimate image based on the overall image content. In this case, by combining the outputs of both models, an indication image is generated as an anomaly, disallowing it from the subsequent homework grading process. This not only improves the flexibility of image recognition but also ensures the efficiency and accuracy of homework grading. Similarly, in page-turning detection scenarios, this mechanism can effectively distinguish between legitimate and illegitimate images, ensuring the smooth progress of page-turning detection.
[0066] In this embodiment of the application, optionally, a fuzzy region recognition algorithm is used to perform fuzzy region recognition on the target image to obtain a fuzzy region recognition result, including: adjusting the size of the target image to a preset fuzzy region recognition size; using a preset fuzzy region recognition algorithm to calculate the image response variance of the target image after size adjustment; generating a fuzzy region recognition result indicating that the target image is not an abnormal image when the image response variance is greater than or equal to a preset variance threshold; and generating a fuzzy region recognition result indicating that the target image is an abnormal image when the image response variance is less than the preset variance threshold.
[0067] In this embodiment, firstly, considering the potential differences in processing and analysis of images of different sizes, to unify standards and improve the adaptability and accuracy of the recognition algorithm, the size of the target image is first adjusted to a preset fuzzy region recognition size. In this embodiment, the fuzzy region recognition size can be 512*512 to maximize the highlighting of fuzzy features in the image, providing a good foundation for subsequent fuzzy region recognition. For example, in homework grading scenarios, the homework images uploaded by students may vary in size due to factors such as shooting equipment and shooting distance. If fuzzy region recognition is directly performed on the original size image, the recognition effect may be affected by the image size being too large or too small. By adjusting it to the preset size of 512*512, it can be ensured that the recognition algorithm runs in a relatively stable image environment, improving the reliability of recognition.
[0068] After image resizing, a preset blur region identification algorithm is used to calculate the image response variance of the resized target image. In practical applications, the preset blur region identification algorithm can be a Laplacian blur judgment algorithm. Laplacian blur judgment is a method that uses the Laplacian operator to calculate the second derivative (or variance) in an image to evaluate image sharpness. By performing a Laplacian transform on the image, the rate of change of pixel values can be measured. For sharp images, the edges have large gradient changes, and the variance of the second derivative (i.e., the response of the Laplacian operator) will be large; conversely, for blurry images, the edges are not obvious, the gradient changes are small, and the variance of the Laplacian operator response will be low. Therefore, in this embodiment, when calculating the image response variance, a Laplacian transform can be performed on the image first to obtain the Laplacian operator response value of each pixel in the image. These response values reflect the rate of change of pixel values at that pixel. Then, the variance of all response values is calculated, and the magnitude of the variance directly reflects the overall sharpness of the image. Thus, the fuzzy region recognition method based on the Laplacian algorithm combines the advantages of the Laplacian operator in image edge detection, and can accurately evaluate the image sharpness from the perspective of pixel value changes.
[0069] Next, based on the calculated image response variance, it is compared with a preset variance threshold. This threshold, determined through extensive sample training and actual testing, serves as a critical value for distinguishing between sharp and blurry images. If the image response variance is greater than or equal to the preset variance threshold, it indicates that the target image is generally sharp and free of blurry areas. Therefore, a blurry region identification result is generated to indicate that the target image is not an abnormal image, confirming that the image meets the sharpness requirements and can proceed to subsequent image analysis tasks. Conversely, if the image response variance is less than the preset variance threshold, it indicates the presence of blurry areas in the target image. These blurry areas are identified as interfering factors, and the target image is determined to be an abnormal image with blurry interference. A blurry region identification result is generated to indicate that the target image is an abnormal image, prompting a re-capture of a clear image.
[0070] In this embodiment, optionally, identifying whether there are interfering factors in the target image includes: adjusting the size of the target image to a preset occlusion recognition size, and standardizing the size-adjusted target image; using a pre-trained occlusion detection model to detect occlusions in the standardized target image to obtain the occlusion area; calculating the total image area of the standardized target image, and using the occlusion area and the total image area to calculate the occlusion coefficient; generating an occlusion detection result indicating that the target image is not an abnormal image when the occlusion coefficient is less than or equal to a preset occlusion coefficient threshold; and generating an occlusion detection result indicating that the target image is an abnormal image when the occlusion coefficient is greater than the preset occlusion coefficient threshold.
[0071] In this embodiment, to comprehensively and accurately identify whether a target image is abnormal, in addition to the blurred region identification method described above, this application embodiment can also identify occlusions, a common interference factor. First, considering that images of different sizes may lead to inaccurate detection results due to size differences, the target image size needs to be adjusted to a preset occlusion identification size. This occlusion identification size is determined through extensive experiments and real-world testing, ensuring that occlusions exhibit relatively stable and easily detectable feature proportions in the image, providing a good foundation for subsequent occlusion detection. For example, in homework grading scenarios, student-uploaded homework images may vary in size due to factors such as shooting distance and angle. If occlusion detection is performed directly, misjudgments or omissions may occur due to the different relative sizes of occlusions in the image. Therefore, adjusting to a preset size unifies the detection standard and improves detection accuracy.
[0072] After resizing the image, the resized target image undergoes standardization. This involves normalizing the pixel values to a specific range, such as 0-1, to eliminate differences in pixel values caused by lighting conditions, shooting equipment, and other factors. This allows the subsequent occlusion detection model to operate with relatively uniform input data, improving its stability and generalization ability. For example, images taken under different lighting conditions may have significantly different pixel value distributions. After standardization, the pixel value distribution becomes more consistent, which helps the occlusion detection model accurately identify occlusions.
[0073] Next, a pre-trained occlusion detection model is used to detect occlusions in the standardized target image. In this embodiment, the occlusion detection model can be the YOLOv8 model, which features fast detection speed and high accuracy, and can quickly and accurately detect the position of a hand in the image. After inputting the standardized target image into the occlusion detection model, the model outputs the position information of the detected hand, and further determines the area covered by the hand in the image based on the output hand position information, and calculates the area of the area. After obtaining the occlusion area, the total area of the standardized target image is first calculated, and then the occlusion area and the total image area are used to calculate the occlusion coefficient. The occlusion coefficient is an indicator used to measure the degree of occlusion in an image, reflecting the proportion of the occluder in the image. Based on the calculated occlusion coefficient, it is compared with a preset occlusion coefficient threshold, which is determined through a large number of experiments and analysis of real-world application scenarios, and represents a critical value for distinguishing whether an image has occlusion interference. When the occlusion coefficient is less than or equal to the preset occlusion coefficient threshold, it indicates that the proportion of occlusions in the target image is small, and their impact on the display of image content and subsequent analysis is minimal. Therefore, an occlusion detection result is generated to indicate that the target image is not an abnormal image. In this case, subsequent image analysis tasks can continue, such as accurately identifying homework content in homework grading and accurately determining page status in page turning detection. Conversely, when the occlusion coefficient is greater than the preset occlusion coefficient threshold, it indicates that there are a large proportion of occlusions in the target image. These occlusions will seriously affect the integrity and clarity of the image content. Therefore, an occlusion detection result is generated to indicate that the target image is an abnormal image, and a reminder is given to retake a clear image.
[0074] Optionally, in this embodiment, the target image is identified according to a target recognition strategy to determine whether the target image is an abnormal image. This includes: when the target recognition strategy indicates that image text quantity recognition is to be performed, text block recognition is performed on the target image and the number of recognized text blocks is counted; when the counted number of blocks is greater than or equal to a preset text quantity threshold, multi-dimensional anomaly recognition is performed on the target image, and the result of multi-dimensional anomaly recognition is used to determine whether the target image is an abnormal image; when the counted number of blocks is less than the text quantity threshold, the target image is determined to be an abnormal image.
[0075] In this embodiment, when the target recognition strategy instructs the image to recognize the number of text blocks, the target image is first identified by text block recognition. In the homework grading scenario, student-submitted homework images typically contain a large amount of text content, and the identified text blocks may originate from questions, answers, solution processes, etc. After completing the text block recognition, the identified text blocks are counted to obtain the number of blocks, recording the total number of regions identified as text blocks. Next, the counted number of blocks is compared with a preset text number threshold. The preset text number threshold is set based on the actual application scenario and requirements, representing a reasonable range of text numbers. For example, in the homework grading scenario, a complete math homework assignment may contain more than 10 text blocks, so the preset text number threshold can be set to 10.
[0076] When the number of blocks obtained from the statistics is greater than or equal to the preset text quantity threshold, it indicates that the target image contains enough text content, which meets the expectations of normal work or documents. Therefore, multi-dimensional anomaly identification is performed on the target image, and the result of multi-dimensional anomaly identification is used to determine whether the target image is an anomalous image. That is, the text content in the image is further identified and analyzed to further determine whether the target image is an anomalous image. The process of multi-dimensional anomaly identification can be found in the description in the above embodiment, and will not be repeated here. For example, for a math homework image containing detailed solution steps and answers, if the number of text blocks identified is 12, which is greater than the preset threshold of 10, multi-dimensional anomaly identification can be performed on the target image, and the result of multi-dimensional anomaly identification can be used to determine whether the target image is an anomalous image. Conversely, when the number of text blocks obtained from the statistics is less than the text quantity threshold, it indicates that the text content in the target image is too small, which may indicate an anomaly. In homework grading scenarios, this could be due to reasons such as students only submitting part of the homework, incomplete image capture, or severe occlusion in the image. For example, a student might only capture the question portion of the homework and omit the answer portion, resulting in only 5 text blocks in the image, less than the preset threshold of 10. In this case, the target image is determined to be an abnormal image, and the student is promptly reminded to resubmit the complete homework image or undergo appropriate processing. Furthermore, in page-turning detection scenarios, if the number of text blocks in the image is too small, it may mean that the page content is not fully displayed, resulting in unsuccessful page turning or an incorrect image capture. The page-turning operation needs to be repeated or the shooting parameters adjusted to ensure that complete and accurate page information is obtained. Thus, by using an image text quantity recognition strategy, the accuracy and reliability of applications such as homework grading and page-turning detection can be effectively improved, providing users with a higher quality service.
[0077] In this embodiment of the application, optionally, the process of recognizing text blocks in the target image and counting the number of recognized text blocks includes: converting the target image into a grayscale image, and performing a binary transformation on the grayscale image to obtain a binary image; performing connected component calculation on the binary image to determine multiple connected component blocks in the binary image; evaluating the aspect ratio and / or block area of each connected component block in the multiple connected component blocks to determine at least one text block in the multiple connected component blocks; and counting the number of the at least one determined text block as the number of blocks.
[0078] In this embodiment, when identifying text blocks and counting the number of blocks, considering that color images contain rich color information, this information is not necessary for the identification of text blocks and will instead increase the computational complexity. Grayscale images, on the other hand, are composed of pixels of different gray levels, retaining the brightness information of the image and can more concisely present the outline and structure of the image. Therefore, in this embodiment, the target image is first converted into a grayscale image to remove color interference and focus on the shape and position features of the text, providing a clearer data foundation for subsequent processing.
[0079] After obtaining the grayscale image, it is further binarized to obtain a binary image. A binary image has only two pixel values, typically 0 and 1, representing black and white respectively. Specifically, a threshold can be set, assigning pixel values greater than the threshold as 1 (white) and those less than or equal to the threshold as 0 (black). This clearly separates the text and background in the image, creating a strong contrast and facilitating subsequent connected component calculations.
[0080] Next, connected component calculation is performed on the binary image. A connected component refers to a set of pixels in a binary image that have the same pixel value and are interconnected. Through the connected component calculation algorithm, multiple connected component blocks can be found in the binary image. These connected component blocks may contain various elements such as text, graphics, and noise, providing a basis for subsequent selection of text blocks. Considering that not all connected component blocks are text blocks, in this embodiment, it is necessary to evaluate the aspect ratio and / or block area of each connected component block. For example, text usually has certain shape characteristics; the aspect ratio of Chinese characters is generally within a certain range, and the area of text blocks is neither too large nor too small. Therefore, by setting reasonable aspect ratio and area thresholds, at least one text block that meets the characteristics of text can be selected from multiple connected component blocks. Finally, the number of at least one text block is counted as the block count. This is used to assess whether the target image is abnormal. For example, in homework grading, if the count of text blocks is too low, it may mean that the student's submitted homework is incomplete, with missing images or incomplete work. Students can be reminded to complete the work in a timely manner. In page-turning detection, when turning pages in an electronic document, the number of text blocks in the current page image can be counted to determine whether the page content is displayed completely. If the number of text blocks is much smaller than the number of normal pages, it indicates that the page turning may not have been successful or the image was captured incorrectly. Automatic page turning or adjustment of shooting parameters can be performed to ensure that complete and accurate page information is obtained, thus improving the accuracy and reliability of page-turning detection.
[0081] In this embodiment of the application, optionally, the aspect ratio and / or block area of each connected component in a plurality of connected component blocks are evaluated to determine at least one text block in the plurality of connected component blocks, including: determining a preset aspect ratio threshold range, calculating the aspect ratio corresponding to each connected component block respectively, and extracting a plurality of first connected component blocks whose aspect ratio is within the aspect ratio threshold range from the plurality of connected component blocks; determining a preset block area threshold range, calculating the block area corresponding to each first connected component block respectively, and extracting at least one first connected component block whose block area is within the block area threshold range from the plurality of first connected component blocks as at least one text block.
[0082] In this embodiment, to remove noise blocks, two indicators, aspect ratio and block area, are used to extract text blocks. First, a preset aspect ratio threshold range is determined, including a maximum and a minimum aspect ratio. In practical applications, the maximum aspect ratio can be set to 5, and the minimum aspect ratio to 0.2. The aspect ratio of each connected component block is calculated, and multiple first connected component blocks with aspect ratios within the threshold range are extracted from these blocks. In other words, blocks with aspect ratios greater than the maximum or less than the minimum are classified as non-text blocks. This is because different characters exhibit certain regularities in their aspect ratios due to differences in their glyph structure. For example, Chinese characters are usually relatively regular square characters, and their aspect ratios generally fall within a relatively stable range; while English words, composed of multiple letters, also have a certain range of aspect ratios. Therefore, by setting such an aspect ratio threshold range, it is possible to effectively exclude connected component blocks whose aspect ratios clearly do not conform to the characteristics of text, such as long strips of noise or thin lines in images. In homework grading, if an image contains thin, elongated noise caused by paper wrinkles, its aspect ratio may be much larger than that of normal text. By filtering by aspect ratio, the connected component blocks formed by these noise points will be excluded, thereby reducing interference in subsequent processing.
[0083] Regarding block area, a preset block area threshold range is determined. This range includes a maximum block area and a minimum block area. In practical applications, the maximum block area can be set to 10,000 pixels, and the minimum block area can be set to 100 pixels. The block area corresponding to each first connected component block is calculated separately, and at least one first connected component block whose area falls within the block area threshold range is extracted as at least one text block. In other words, blocks with areas smaller than the minimum block area are classified as noise blocks, blocks with areas larger than the maximum block area are classified as non-text blocks, and only blocks with areas within the block area threshold range are identified as text blocks. This is because the size of text blocks also has a certain degree of reasonableness; blocks that are too small may be noise or dust in the image, while blocks that are too large may be graphics, tables, or other non-text content in the image. Therefore, by analyzing the area range occupied by normal text in the image and setting such an area threshold, it is possible to further ensure that the selected connected component blocks have a reasonable area size. For example, in a page-turning detection scenario, if there is a connected component block in the page image of an electronic document that is too large but has an aspect ratio that matches the characteristics of text, it may be a large graphic in the image. By filtering by the area of the block, this graphic will be excluded, thus accurately identifying the real text block.
[0084] It should be noted that in practical applications, the aspect ratio threshold range and the block area threshold range can be set to dynamic values. The threshold for noisy blocks can be adjusted based on the average block width, height, and area, thus better adapting to the characteristics of images in different scenarios and improving the accuracy and flexibility of text block detection. For example, in images with different font sizes and layout styles, dynamically adjusting the threshold can more accurately identify text blocks.
[0085] In this way, the image text block detection method based on connected components can effectively filter images containing debris on blank paper, or illegal images such as desktops, floors, and walls, providing accurate and reliable image analysis results for applications such as homework grading and page turning detection, thereby improving the performance and user experience of the entire system.
[0086] Optionally, in this embodiment of the application, the method further includes: uploading the target image to a cloud server when it is determined that the target image is not an abnormal image according to the target recognition strategy; receiving and displaying information fed back by the cloud server, wherein the information fed back by the cloud server includes text evaluation results or text content to be displayed.
[0087] In this embodiment, if the target image is determined not to be an abnormal image according to the target recognition strategy, for example, in the homework grading scenario, the image is a homework image containing complete and clear text content, and in the page turning detection scenario, it is an image that displays the correct page content after normal page turning, then the target image can be uploaded to the cloud server.
[0088] After receiving the target image, the cloud server processes it using text recognition and analysis technology. In homework grading scenarios, the cloud server identifies the text content within the image and, combined with pre-set grading rules and standard answers, provides a comprehensive and detailed evaluation of the assignment, generating detailed text evaluation results, including accuracy, error types, and knowledge point mastery. For example, in a Chinese composition assignment, the cloud server can not only identify the text but also analyze the composition's grammatical structure, vocabulary usage, and thematic expression, providing targeted evaluations and suggestions. In page-turning detection scenarios, the cloud server identifies the text content on the page to determine what content should be displayed on the next page.
[0089] Subsequently, the terminal receives information from the cloud server and displays it to the user in an intuitive and clear manner. In homework grading scenarios, teachers or students can see detailed text assessment results on the terminal. Teachers can quickly understand students' learning progress based on these results and provide targeted guidance; students can also clearly understand their homework problems and make timely improvements. In page-turning detection scenarios, accurate page-turning detection and text display enable smooth browsing of document content without affecting the reading experience due to page-turning errors, improving the accuracy and efficiency of homework grading and page-turning detection, and providing users with a higher quality and more convenient service.
[0090] Optionally, in this embodiment, the cloud server's processing of the target image includes: performing text recognition on the target image to obtain text recognition results, and recognizing the scene environment in which the terminal obtains the target image; when the scene environment indicates that text evaluation is to be performed, obtaining an evaluation standard that matches the target image, evaluating the text recognition results using the evaluation standard, obtaining text evaluation results, and feeding them back to the terminal; when the scene environment indicates that the text content is to be changed, querying the text content to be displayed using the text recognition results and feeding them back to the terminal.
[0091] In this embodiment, once the target image is uploaded to the cloud server, the cloud server initiates a series of processing steps for the target image. First, text recognition technology is used to process the target image. This technology integrates deep learning algorithms and Optical Character Recognition (OCR) technology, enabling accurate identification of various textual information within the image. Whether printed or handwritten, it can achieve high accuracy, yielding detailed text recognition results. For example, in a homework grading scenario, student-submitted homework images may contain text in different fonts and sizes. The cloud server's text recognition technology can accurately identify each character, laying the foundation for subsequent evaluation. Similarly, in a page-turning detection scenario, text on electronic document pages can be clearly recognized, ensuring the acquisition of complete textual information.
[0092] Simultaneously, the cloud server accurately identifies the scene environment in which the target image is acquired by the terminal. For example, it identifies the function under which the target image was uploaded. If the target image was uploaded during homework grading, it can be determined that this is a homework submission scenario; if the image was uploaded during reading, it can be determined that this is a page-turning scenario. When the scene environment indicates that text evaluation is required, the cloud server obtains the evaluation criteria that match the target image. The evaluation criteria are pre-set according to different subjects, different homework types, and different teaching requirements. For example, for math homework, the evaluation criteria will cover whether the formula is used correctly, whether the solution steps are complete, and whether the logical reasoning is rigorous; for Chinese composition, the evaluation criteria will include whether the theme is clear, whether the content is rich, and whether the language expression is fluent. After obtaining the evaluation criteria, the cloud server uses these criteria to evaluate the text recognition results. Specifically, it uses relevant algorithms to compare and analyze each piece of content in the text recognition results with the evaluation criteria to obtain the text evaluation result, which is then promptly fed back to the terminal. In homework grading scenarios, teachers can quickly obtain students' homework assessment results through the terminal, understand students' mastery of various knowledge points, and thus provide targeted teaching and guidance; students can also identify their shortcomings based on the assessment results, make timely improvements, and improve learning outcomes.
[0093] When the scene environment indicates a change in text content, the cloud server uses the text recognition results to query its database. This database stores a large amount of text information, including various documents and knowledge bases. Based on keywords and themes in the text recognition results, the cloud server checks the database to see if the image is a legitimate page image and determines the content of the next page to be displayed. This content is then sent back to the terminal. In other words, in page-turning detection scenarios, when a user turns a page, the server can use the recognized page text content to determine the content to be displayed on the next page and send it to the terminal device, eliminating the need for manual page turning and improving the user experience.
[0094] In this way, through the above process, the cloud server can provide accurate and efficient services according to different scenario requirements, improving the functionality and practicality of applications such as homework grading and page-turning detection. In practical applications, a cloud search module can be set up in the cloud server, and the processes mentioned above, such as generating text evaluation results and determining the text content to be displayed, can all be completed by the cloud search module.
[0095] In summary, the logical process of the technical solution in this application is summarized as follows: See Figure 2First, the input image is acquired and feature points are extracted. If the number of extracted feature points is not greater than a minimum threshold, it is directly identified as an abnormal image. If it is greater than the minimum threshold, it is further determined whether the number of feature points is greater than a standard threshold. If not, text count recognition is performed. After text count recognition, if the number of recognized characters is less than the text count threshold, it is identified as an abnormal image. If it is greater than or equal to the text count threshold, it needs to be transmitted to a physical machine for multi-dimensional anomaly recognition. For the physical machine, it calls a multi-class filtering model, a binary filtering model, a fuzzy region recognition algorithm, and an occlusion detection model to perform multi-dimensional anomaly recognition on the target image. The judgment results of the multi-class filtering model and the binary filtering model are combined. If both models belong to the same dimension and both models determine that the target image is an abnormal image, then the target image is determined to be an abnormal image. Therefore, the multi-class filtering model and the binary filtering model will combine to obtain a classification filtering result, while the fuzzy region recognition algorithm and the occlusion detection model will each determine a judgment result, resulting in three judgment results. If any one of the three results determines that the target image is an abnormal image, the physical machine will determine that the image is abnormal. Only when all dimensions determine that the current image is a normal image will the image be uploaded to the cloud server, where the cloud server will perform subsequent processing such as text recognition, text evaluation based on the scene environment, or modification of text content, such as text evaluation related to homework correction or text content modification and display during page turning detection.
[0096] Furthermore, as Figure 1 To specifically implement the method, this application provides an abnormal image recognition device, such as... Figure 3 As shown, the device includes: a feature point filtering module 301, a recognition strategy determination module 302, and an image recognition module 303.
[0097] The feature point filtering module 301 is used to determine the target image to be identified, count the number of key point features included in the target image, and determine whether the target image is an abnormal image based on the number of features. The identification strategy determination module 302 is used to determine the target identification strategy to be executed when it is determined that the target image is not an abnormal image. The target identification strategy indicates to perform multi-dimensional anomaly identification or image text quantity identification. The image recognition module 303 is used to recognize the target image according to the target recognition strategy in order to determine whether the target image is an abnormal image.
[0098] In a specific application scenario, the feature point filtering module 301 is used to determine a preset feature type, identify the features included in the target image, and extract multiple key point features of the same type as the feature type from the target image; count the number of the multiple key point features to obtain the feature count; obtain a preset minimum number threshold, and compare the feature count with the minimum number threshold; when the comparison determines that the feature count is greater than or equal to the minimum number threshold, determine that the target image is not the abnormal image; when the comparison determines that the feature count is less than the minimum number threshold, determine that the target image is the abnormal image. Accordingly, the feature point filtering module 301 is also used to generate and output an abnormal image alert and terminate the current image recognition process when it is determined that the target image is the abnormal image.
[0099] In a specific application scenario, the recognition strategy determination module 302 is used to determine a preset standard quantity threshold, compare the number of features with the standard quantity threshold, wherein the standard quantity threshold is greater than a minimum quantity threshold; when the comparison determines that the number of features is greater than or equal to the standard quantity threshold, the target recognition strategy indicating multi-dimensional anomaly recognition is determined; when the comparison determines that the number of features is less than the standard quantity threshold, the target recognition strategy indicating image text quantity recognition is determined.
[0100] In a specific application scenario, the image recognition module 303 is used to determine a multi-dimensional anomaly recognition node for performing multi-dimensional anomaly recognition when the target recognition strategy indicates that multi-dimensional anomaly recognition should be performed. The multi-dimensional anomaly recognition node is used to call a multi-classification filtering model, a binary classification filtering model, a fuzzy region recognition algorithm, and an occlusion detection model to perform multi-dimensional anomaly recognition on the target image. Based on the multi-dimensional anomaly recognition node, the target image is recognized to obtain classification recognition results, fuzzy region recognition results, and occlusion detection results. When the classification recognition results, the fuzzy region recognition results, and the occlusion detection results all indicate that the target image is not the anomaly image, the target image is determined not to be the anomaly image. When any one of the classification recognition results, the fuzzy region recognition results, and the occlusion detection results indicates that the target image is the anomaly image, the target image is determined to be the anomaly image.
[0101] In specific application scenarios, the image recognition module 303 is further configured to collect and construct an image dataset for model training, the image dataset including multiple types of abnormal sample images and invalid sample images; perform image size standardization and image data standardization on each sample image included in the image dataset according to a first preset resolution to obtain a first training dataset, and perform image size standardization and image data standardization on each sample image included in the image dataset according to a second preset resolution to obtain a second training dataset, wherein the second preset resolution is greater than the first preset resolution; determine a multi-classification model architecture, train the multi-classification model architecture using the first training dataset to obtain and store the multi-classification filtering model, so as to retrieve and call the multi-classification filtering model when performing abnormal image recognition; determine a binary classification model architecture, train the binary classification model architecture using the second training dataset to obtain and store the binary classification filtering model, so as to retrieve and call the binary classification filtering model when performing abnormal image recognition.
[0102] In a specific application scenario, the image recognition module 303 is used to obtain a basic multi-classification model architecture, which includes multiple neurons; a first segmented activation function is added to the fully connected layer in the basic multi-classification model architecture, and the first activation threshold of the first segmented activation function is set to 1 to obtain the multi-classification model architecture. The fully connected layer in the multi-classification model architecture includes a regular activation function and the first segmented activation function. During the model operation, the regular activation function is used to perform an activation value calculation on the received input value. If the calculation result obtained from the first activation value calculation is less than the first activation threshold, the first segmented activation function is used to perform a second activation value calculation on the calculation result obtained from the first activation value calculation.
[0103] In a specific application scenario, the image recognition module 303 is used to obtain a basic binary classification model architecture, which includes two neurons; a second segmented activation function is added to the fully connected layer in the basic binary classification model architecture, and the second activation threshold of the second segmented activation function is set to 2 to obtain the binary classification model architecture. The fully connected layer in the binary classification model architecture includes a regular activation function and the second segmented activation function. During the model operation, the regular activation function is used to perform an activation value calculation on the received input value. If the calculation result obtained from the first activation value calculation is less than the second activation threshold, the second segmented activation function is used to perform a second activation value calculation on the calculation result obtained from the first activation value calculation.
[0104] In specific application scenarios, the image recognition module 303 is used to classify and filter the target image by calling the multi-class filtering model and the binary filtering model based on the multi-dimensional anomaly recognition node, and to determine the classification recognition result based on the first classification filtering result output by the multi-class filtering model and the second classification filtering result output by the binary filtering model; simultaneously, based on the multi-dimensional anomaly recognition node, the fuzzy region recognition algorithm is used to perform fuzzy region recognition on the target image to obtain the fuzzy region recognition result; and simultaneously, based on the multi-dimensional anomaly recognition node, the occlusion detection model is called to perform occlusion detection on the target image to obtain the occlusion detection result.
[0105] In a specific application scenario, the image recognition module 303 is used to generate a classification recognition result indicating that the target image is an abnormal image when both the first classification filtering result and the second classification filtering result indicate that the target image belongs to the abnormal image category; and to generate a classification recognition result indicating that the target image is not an abnormal image when either the first classification filtering result or the second classification filtering result indicates that the target image does not belong to the abnormal image category.
[0106] In a specific application scenario, the image recognition module 303 is used to adjust the size of the target image to a preset fuzzy region recognition size, and use the fuzzy region recognition algorithm to calculate the image response variance of the target image after size adjustment; if the image response variance is greater than or equal to a preset variance threshold, a fuzzy region recognition result indicating that the target image is not the abnormal image is generated; if the image response variance is less than the preset variance threshold, a fuzzy region recognition result indicating that the target image is the abnormal image is generated.
[0107] In a specific application scenario, the image recognition module 303 is used to adjust the size of the target image to a preset occlusion recognition size, and to perform standardization processing on the size-adjusted target image; using the occlusion detection model, to detect occlusions in the standardized target image and obtain the occlusion area; to calculate the total area of the standardized target image, and to calculate the occlusion coefficient using the occlusion area and the total image area; if the occlusion coefficient is less than or equal to a preset occlusion coefficient threshold, to generate an occlusion detection result indicating that the target image is not the abnormal image; if the occlusion coefficient is greater than the preset occlusion coefficient threshold, to generate an occlusion detection result indicating that the target image is the abnormal image.
[0108] In a specific application scenario, the image recognition module 303 is used to perform text block recognition on the target image and count the number of recognized text blocks when the target recognition strategy indicates that image text quantity recognition is to be performed; when the counted number of blocks is greater than or equal to a preset text quantity threshold, the multi-dimensional anomaly recognition is performed on the target image, and the target image is determined to be an abnormal image based on the result of the multi-dimensional anomaly recognition; when the counted number of blocks is less than the text quantity threshold, the target image is determined to be an abnormal image.
[0109] In a specific application scenario, the image recognition module 303 is used to convert the target image into a grayscale image, and to perform a binary conversion on the grayscale image to obtain a binary image; to perform connected component calculation on the binary image to determine multiple connected component blocks in the binary image; to evaluate the aspect ratio and / or block area of each of the multiple connected component blocks to determine at least one text block in the multiple connected component blocks; and to count the number of the at least one text block as the number of blocks.
[0110] In specific application scenarios, the image recognition module 303 is used to determine a preset aspect ratio threshold range, calculate the aspect ratio corresponding to each of the connected component blocks, and extract a plurality of first connected component blocks whose aspect ratio is within the aspect ratio threshold range from the plurality of connected component blocks; determine a preset block area threshold range, calculate the block area corresponding to each of the first connected component blocks, and extract at least one first connected component block whose block area is within the block area threshold range from the plurality of first connected component blocks as the at least one text block.
[0111] In specific application scenarios, the device further includes: An upload module is used to upload the target image to a cloud server when the target image is determined not to be the abnormal image according to the target recognition strategy. The display module is used to receive and display the information fed back by the cloud server, wherein the information fed back by the cloud server includes text evaluation results or text content to be displayed.
[0112] In a specific application scenario, the cloud server's processing of the target image includes: performing text recognition on the target image to obtain a text recognition result, and recognizing the scene environment of the target image obtained by the terminal; when the scene environment indicates that text evaluation is to be performed, obtaining an evaluation standard matching the target image, evaluating the text recognition result using the evaluation standard, obtaining the text evaluation result, and feeding it back to the terminal; when the scene environment indicates that the text content should be changed, querying the text content to be displayed using the text recognition result and feeding it back to the terminal.
[0113] The apparatus provided in this application can effectively filter out irrelevant abnormal images by statistically analyzing the number of key point features in the target image and making preliminary anomaly judgments accordingly. This reduces the upload of a large number of abnormal images at the source. Furthermore, it employs targeted recognition strategies, such as multi-dimensional anomaly or text quantity recognition, only for the preliminarily determined non-abnormal images. This prevents problems such as non-standard algorithm input and frequent uploads caused by abnormal images, avoids the waste of computing resources caused by indiscriminately processing all images, significantly reduces the computing pressure on the cloud server, stabilizes system performance, reduces server burden and network bandwidth usage, reduces the risk of system response delays or crashes, and ensures the reliability and efficiency of the service.
[0114] It should be noted that other corresponding descriptions of the functional units involved in the abnormal image recognition device provided in this application embodiment can be found in the following references. Figure 1 and Figure 2 The corresponding description in [the document] will not be repeated here.
[0115] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.
[0116] The above embodiments and the technical features in the embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0117] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
[0118] In an exemplary embodiment, see Figure 4 The invention also provides a computer device including a bus, a processor, a memory, and a communication interface. It may also include an input / output interface and a display device, wherein the various functional units can communicate with each other via the bus. The memory stores a computer program, and the processor executes the program stored in the memory to perform the abnormal image recognition method described in the above embodiments.
[0119] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the abnormal image recognition method.
[0120] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented in hardware or by using software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solution of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) and includes several instructions to cause a computer device (such as a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0121] Those skilled in the art will understand that the accompanying drawings are merely schematic diagrams of a preferred embodiment, and the modules or processes shown in the drawings are not necessarily essential for implementing this application.
[0122] Those skilled in the art will understand that the modules in the apparatus of the implementation scenario can be distributed within the apparatus of the implementation scenario as described, or they can be located in one or more apparatuses different from this implementation scenario, with corresponding changes. The modules of the above-described implementation scenario can be combined into one module, or they can be further divided into multiple sub-modules.
[0123] The serial numbers in this application are for descriptive purposes only and do not represent the superiority or inferiority of the implementation scenario.
[0124] The above disclosures are only a few specific implementation scenarios of this application. However, this application is not limited to these. Any variations that can be conceived by those skilled in the art should fall within the protection scope of this application.
Claims
1. An abnormal image recognition method, characterized in that, include: The target image to be identified is determined, the number of key point features included in the target image is counted, and the target image is determined to be an abnormal image based on the number of features. If it is determined that the target image is not an abnormal image, a target recognition strategy to be executed is determined, and the target recognition strategy indicates to perform multi-dimensional anomaly recognition or image text quantity recognition. The target image is identified according to the target recognition strategy to determine whether the target image is an abnormal image.
2. The method of claim 1, wherein, The process of determining the target image to be identified, counting the number of key point features included in the target image, and determining whether the target image is an abnormal image based on the number of features includes: A preset feature type is determined, and the features included in the target image are identified in order to extract multiple key point features of the same type as the feature type from the target image; The number of features at the multiple key points is counted to obtain the total number of features; Obtain a preset minimum quantity threshold, and compare the number of features with the minimum quantity threshold; When the comparison determines that the number of features is greater than or equal to the minimum number threshold, it is determined that the target image is not the abnormal image; When the comparison determines that the number of features is less than the minimum number threshold, the target image is determined to be the abnormal image; Accordingly, the method further includes: If the target image is determined to be an abnormal image, an abnormal image alert is generated and output, and the current image recognition process is terminated.
3. The method of claim 1, wherein, The determination of the target identification strategy to be executed includes: A preset standard quantity threshold is determined, and the number of features is compared with the standard quantity threshold, wherein the standard quantity threshold is greater than the minimum quantity threshold; When the comparison determines that the number of features is greater than or equal to the standard number threshold, the target recognition strategy that indicates multi-dimensional anomaly recognition is determined. When the comparison determines that the number of features is less than the standard number threshold, the target recognition strategy that indicates the text quantity recognition in the image is determined.
4. The method of claim 1, wherein, The step of identifying the target image according to the target recognition strategy to determine whether the target image is an abnormal image includes: When the target recognition strategy indicates that multi-dimensional anomaly recognition is to be performed, a multi-dimensional anomaly recognition node is determined for multi-dimensional anomaly recognition. The multi-dimensional anomaly recognition node is used to call a multi-class filtering model, a binary filtering model, a fuzzy region recognition algorithm, and an occlusion detection model to perform multi-dimensional anomaly recognition on the target image. Based on the multi-dimensional anomaly recognition node, the target image is recognized to obtain classification recognition results, blurred region recognition results, and occlusion detection results; When the classification recognition result, the blurred region recognition result, and the occlusion detection result all indicate that the target image is not the abnormal image, the target image is determined to be not the abnormal image. When any one of the classification recognition result, the blurred region recognition result, and the occlusion detection result indicates that the target image is the abnormal image, the target image is determined to be the abnormal image.
5. The method of claim 4, wherein, The method further includes: Collect and construct an image dataset for model training, the image dataset including multiple types of abnormal sample images and invalid sample images; Based on a first preset resolution, each sample image included in the image dataset is subjected to image size standardization and image data standardization to obtain a first training dataset; and based on a second preset resolution, each sample image included in the image dataset is subjected to image size standardization and image data standardization to obtain a second training dataset, wherein the second preset resolution is greater than the first preset resolution. A multi-classification model architecture is determined, and the first training dataset is used to train the multi-classification model architecture to obtain and store the multi-classification filtering model, so as to obtain and call the multi-classification filtering model when performing abnormal image recognition. A binary classification model architecture is determined, and the binary classification model architecture is trained using the second training dataset to obtain and store the binary classification filtering model, so as to retrieve and call the binary classification filtering model when performing abnormal image recognition.
6. The method of claim 5, wherein, The determination of the multi-classification model architecture includes: Obtain the basic multi-class classification model architecture, which includes multiple neurons; A first segmented activation function is added to the fully connected layer in the basic multi-class classification model architecture, and the first activation threshold of the first segmented activation function is set to 1 to obtain the multi-class classification model architecture. The fully connected layer in the multi-class classification model architecture includes a regular activation function and the first segmented activation function. During the model operation, the regular activation function is used to calculate the activation value of the received input value once. If the calculation result of the first activation value is less than the first activation threshold, the first segmented activation function is used to calculate the activation value of the calculation result of the first activation value a second time.
7. The method of claim 5, wherein, The determination of the binary classification model architecture includes: Obtain the basic binary classification model architecture, which includes two neurons; A second segmented activation function is added to the fully connected layer in the basic binary classification model architecture, and the second activation threshold of the second segmented activation function is set to 2 to obtain the binary classification model architecture. The fully connected layer in the binary classification model architecture includes a regular activation function and the second segmented activation function. During the model operation, the regular activation function is used to calculate the activation value of the received input value once. If the calculation result of the first activation value is less than the second activation threshold, the second segmented activation function is used to calculate the activation value of the first activation value a second time.
8. The method of claim 4, wherein, The process of identifying the target image based on the multi-dimensional anomaly detection nodes, and obtaining classification results, blurred region identification results, and occlusion detection results, includes: Based on the multi-dimensional anomaly identification node, the multi-class filtering model and the binary filtering model are invoked to classify and filter the target image, and the classification and identification result is determined based on the first classification filtering result output by the multi-class filtering model and the second classification filtering result output by the binary filtering model. Simultaneously, based on the multi-dimensional anomaly identification node, the fuzzy region identification algorithm is used to identify the fuzzy region of the target image to obtain the fuzzy region identification result. Simultaneously, based on the multi-dimensional anomaly identification node, the occlusion detection model is invoked to identify occlusions in the target image, thereby obtaining the occlusion detection result.
9. The method of claim 8, wherein, The step of determining the classification recognition result based on the first classification filtering result output by the multi-class filtering model and the second classification filtering result output by the binary filtering model includes: When both the first classification filtering result and the second classification filtering result indicate that the target image belongs to the abnormal image category, the classification recognition result used to indicate that the target image is the abnormal image is generated; When the first classification filtering result or the second classification filtering result indicates that the target image does not belong to the abnormal image category, the classification recognition result indicating that the target image is not the abnormal image is generated.
10. The method of claim 8, wherein, The step of using the fuzzy region recognition algorithm to identify fuzzy regions in the target image and obtaining the fuzzy region recognition result includes: The size of the target image is adjusted to a preset size for blurry region recognition, and the image response variance of the target image after size adjustment is calculated using the blurry region recognition algorithm. If the image response variance is greater than or equal to a preset variance threshold, a blurred region identification result is generated to indicate that the target image is not the abnormal image; If the image response variance is less than the preset variance threshold, a blurred region identification result is generated to indicate that the target image is the abnormal image.
11. The method of claim 8, wherein, The step of calling the occlusion detection model to identify occlusions in the target image and obtaining the occlusion detection result includes: The size of the target image is adjusted to a preset occlusion recognition size, and the adjusted target image is then standardized. Using the aforementioned occlusion detection model, occlusion detection is performed on the standardized target image to obtain the occlusion area. The total area of the target image after statistical standardization is used to calculate the occlusion coefficient using the occlusion area and the total image area. If the occlusion coefficient is less than or equal to a preset occlusion coefficient threshold, an occlusion detection result is generated to indicate that the target image is not the abnormal image; If the occlusion coefficient is greater than the preset occlusion coefficient threshold, an occlusion detection result is generated to indicate that the target image is the abnormal image.
12. The method of claim 1, wherein, The step of identifying the target image according to the target recognition strategy to determine whether the target image is an abnormal image includes: When the target recognition strategy indicates that image text quantity recognition is to be performed, text block recognition is performed on the target image and the number of recognized text blocks is counted; When the number of blocks obtained by statistics is greater than or equal to the preset text number threshold, the target image is subjected to the multi-dimensional anomaly recognition, and the target image is determined to be an abnormal image based on the result of the multi-dimensional anomaly recognition. When the number of blocks obtained from the statistics is less than the number of characters, the target image is determined to be the abnormal image.
13. The method according to claim 12, characterized in that, The step of performing text block recognition on the target image and counting the number of recognized text blocks includes: The target image is converted into a grayscale image, and the grayscale image is then subjected to a binary transformation to obtain a binary image. Connectivity component calculation is performed on the binary image to determine multiple connected component blocks in the binary image; The aspect ratio and / or block area of each of the plurality of connected component blocks are evaluated to determine at least one text block among the plurality of connected component blocks; The number of the at least one text block determined by statistics is taken as the number of blocks.
14. The method of claim 13, wherein, The step of evaluating the aspect ratio and / or block area of each of the plurality of connected component blocks to determine at least one text block among the plurality of connected component blocks includes: A preset aspect ratio threshold range is determined, and the aspect ratio corresponding to each connected component block is calculated. A plurality of first connected component blocks whose aspect ratios are within the aspect ratio threshold range are extracted from the plurality of connected component blocks. A preset block area threshold range is determined, the block area corresponding to each first connected component block is calculated, and at least one first connected component block whose block area is within the block area threshold range is extracted from the plurality of first connected component blocks as the at least one text block.
15. The method of claim 1, wherein, The method further includes: If the target image is determined not to be the abnormal image according to the target recognition strategy, the target image is uploaded to the cloud server; Receive and display information fed back by the cloud server, wherein the information fed back by the cloud server includes text evaluation results or text content to be displayed.
16. The method of claim 15, wherein, The cloud server's processing of the target image includes: The target image is subjected to text recognition to obtain the text recognition result, and the scene environment of the target image is obtained by the recognition terminal; When the scene environment indicates that text evaluation is to be performed, an evaluation standard matching the target image is obtained, the text recognition result is evaluated using the evaluation standard, the text evaluation result is obtained, and the result is fed back to the terminal. When the scene environment indicates a change in text content, the text recognition result is used to query the text content to be displayed and feedback is sent to the terminal.
17. An abnormal image recognition device characterized by comprising: include: The feature point filtering module is used to determine the target image to be identified, count the number of key point features included in the target image, and determine whether the target image is an abnormal image based on the number of features. The identification strategy determination module is used to determine the target identification strategy to be executed when it is determined that the target image is not an abnormal image. The target identification strategy indicates to perform multi-dimensional anomaly identification or image text quantity identification. The image recognition module is used to recognize the target image according to the target recognition strategy in order to determine whether the target image is an abnormal image. 18.A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the computer device is configured to perform the method according to any one of claims 1-17. When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 16.
19. A readable storage medium, having stored thereon a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 16.