Image detection method and apparatus, and computer device and storage medium

By obtaining detection task description information and automatically processing images using semantic recognition models and image detection models, the problem of low image detection efficiency is solved, and efficient and accurate image detection and abnormal behavior recognition are achieved.

WO2025140712A1PCT designated stage expired Publication Date: 2025-07-03SF TECH CO LTD

Patent Information

Application Number
PCT/CN2024/143841
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-30
Filing Date
2024-12-30
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

In the prior art, image detection efficiency is low, and the need for manual labeling of images is time-consuming and affecting efficiency.

Method used

By obtaining detection task description information, using the semantic recognition model to convert it into a computer-recognizable detection task, and processing the image to be detected in combination with the image detection model to automatically determine the detection area and detection results.

Benefits of technology

It improves the efficiency and accuracy of image detection, can automatically identify and feedback abnormal behaviors, and reduces manual intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024143841_03072025_PF_FP_ABST
    Figure CN2024143841_03072025_PF_FP_ABST
Patent Text Reader

Abstract

Provided are an image detection method and apparatus, and a computer device, a storage medium and a computer program product. The method comprises: acquiring detection task description information for an image to be subjected to detection (S101); inputting the detection task description information into a semantic recognition model, so as to obtain a detection task corresponding to the detection task description information (S102); and on the basis of the detection task, performing corresponding detection processing on said image by means of an image detection model, so as to obtain a detection result for said image (S103).
Need to check novelty before this filing date? Find Prior Art

Description

Image detection method, device, computer equipment and storage medium

[0001] Related applications

[0002] This application claims priority to Chinese patent application number 2023118723339, filed on December 30, 2023, entitled “Image Detection Method, Device, Computer Equipment and Storage Medium,” the entire text of which is hereby incorporated by reference. Technical Field

[0003] The present application relates to the field of computer technology, and in particular to an image detection method, apparatus, computer equipment, storage medium, and computer program product. Background Art

[0004] With the development of computer technology, image detection has important applications in many fields. By analyzing traditional image detection methods, we can see that image detection is inefficient. Therefore, how to perform image detection efficiently has become an important research direction.

[0005] Traditional technologies usually require manual image annotation before image detection; however, this method of image detection requires a lot of manual processing time, resulting in low image detection efficiency. Summary of the Invention

[0006] According to various embodiments provided in the present application, an image detection method, apparatus, computer device, computer-readable storage medium, and computer program product are provided.

[0007] In a first aspect, the present application provides an image detection method, which is applied to a terminal. The method comprises:

[0008] Obtaining detection task description information for the image to be detected;

[0009] Inputting the detection task description information into a semantic recognition model to obtain a detection task corresponding to the detection task description information;

[0010] Through the image detection model, according to the detection task, corresponding detection processing is performed on the image to be detected to obtain a detection result for the image to be detected.

[0011] In one embodiment, the image detection model is used to perform corresponding detection processing on the image to be detected according to the detection task to obtain a detection result for the image to be detected, including:

[0012] Performing corresponding detection processing on the image to be detected according to the detection task through the image detection model to determine the detection area in the image to be detected and the detection result of the detection area;

[0013] A detection result of the image to be detected is determined according to the detection area and the detection result of the detection area.

[0014] In one embodiment, performing corresponding detection processing on the image to be detected according to the detection task by the image detection model to determine the detection area in the image to be detected and the detection result of the detection area includes:

[0015] Determining, by means of the image detection model and according to the detection task, a detection object of the detection task and a detection behavior of the detection task;

[0016] Performing corresponding detection processing on the image to be detected according to the detection object to determine the detection area;

[0017] According to the detection behavior, corresponding detection processing is performed on the detection area to determine the detection result of the detection area.

[0018] In one embodiment, performing corresponding detection processing on the detection area according to the detection behavior to determine the detection result of the detection area includes:

[0019] Performing corresponding detection processing on the detection area according to the detection behavior to determine a matching value of the detection behavior corresponding to the detection area;

[0020] The detection result of the detection area is determined according to a preset matching threshold and a matching value of the detection behavior; the detection result of the detection area is used to indicate whether the detection behavior exists in the detection area.

[0021] In one embodiment, determining the detection result of the image to be detected based on the detection area and the detection result of the detection area includes:

[0022] When the detection result of the detection area indicates that the detection behavior exists in the detection area, determining the detection result of the image to be detected as the presence of abnormal behavior in the image to be detected;

[0023] In a case where the detection result of the detection area indicates that the detection behavior does not exist in the detection area, the detection result of the image to be detected is determined to be that no abnormal behavior exists in the image to be detected.

[0024] In one embodiment, inputting the detection task description information into a semantic recognition model to obtain the detection task corresponding to the detection task description information includes:

[0025] Performing semantic recognition processing on the detection task description information through the semantic recognition model to obtain semantic feature information of the detection task description information;

[0026] The detection task is determined according to the semantic feature information.

[0027] In one embodiment, after performing corresponding detection processing on the image to be detected according to the detection task using the image detection model to obtain a detection result for the image to be detected, the method further includes:

[0028] When the detection result of the image to be detected is that abnormal behavior exists in the image to be detected, generating corresponding warning information according to the detection result of the image to be detected;

[0029] According to the warning information, corresponding warning processing is performed on the detection object of the detection task.

[0030] In a second aspect, the present application further provides an image detection device. The device comprises:

[0031] An information acquisition module is used to obtain detection task description information for the image to be detected;

[0032] An information input module is used to input the detection task description information into a semantic recognition model to obtain a detection task corresponding to the detection task description information;

[0033] The image processing module is used to perform corresponding detection processing on the image to be detected according to the detection task through the image detection model to obtain a detection result for the image to be detected.

[0034] In a third aspect, the present application further provides a computer device. The computer device includes a memory and one or more processors, wherein the memory stores computer-readable instructions, and when the one or more processors execute the computer-readable instructions, the following steps are implemented:

[0035] Obtaining detection task description information for the image to be detected;

[0036] Inputting the detection task description information into a semantic recognition model to obtain a detection task corresponding to the detection task description information;

[0037] Through the image detection model, according to the detection task, corresponding detection processing is performed on the image to be detected to obtain a detection result for the image to be detected.

[0038] In a fourth aspect, the present application further provides one or more computer-readable storage media. The computer-readable storage media stores computer-readable instructions thereon, and when the computer-readable instructions are executed by one or more processors, the following steps are implemented:

[0039] Obtaining detection task description information for the image to be detected;

[0040] Inputting the detection task description information into a semantic recognition model to obtain a detection task corresponding to the detection task description information;

[0041] Through the image detection model, according to the detection task, corresponding detection processing is performed on the image to be detected to obtain a detection result for the image to be detected.

[0042] In a fifth aspect, the present application further provides a computer program product. The computer program product includes computer-readable instructions, which, when executed by one or more processors, implement the following steps:

[0043] Obtaining detection task description information for the image to be detected;

[0044] Inputting the detection task description information into a semantic recognition model to obtain a detection task corresponding to the detection task description information;

[0045] Through the image detection model, according to the detection task, corresponding detection processing is performed on the image to be detected to obtain a detection result for the image to be detected.

[0046] The details of one or more embodiments of the present application are set forth in the accompanying drawings and the description below. Other features, objects, and advantages of the present application will become apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0048] FIG1 is a schematic diagram of a flow chart of an image detection method in one embodiment;

[0049] FIG2 is a schematic flow chart of the steps of determining the detection result of an image to be detected in one embodiment;

[0050] FIG3 is a schematic flow chart of an image detection method in another embodiment;

[0051] FIG4 is a schematic diagram of a flow chart of model processing in one embodiment;

[0052] FIG5 is a schematic diagram of outputting a video in one embodiment;

[0053] FIG6 is a block diagram of an image detection device according to an embodiment;

[0054] FIG7 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0055] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0056] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0057] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0058] With the development of globalization, the scale of the logistics industry continues to expand, and the workload of logistics personnel is also increasing. However, due to high work pressure and complex working environments, the problem of illegal operations by logistics personnel has gradually emerged. These illegal operations not only affect logistics efficiency and cause damage to express parcels, but also pose a threat to logistics safety. Therefore, how to effectively identify and prevent illegal operations by logistics personnel has become a pressing issue in the logistics industry. To this end, this application provides a logistics personnel illegal operation identification system based on a large model. This system uses big data and artificial intelligence technologies to monitor and analyze logistics personnel's behavior in real time, thereby achieving timely detection and prevention of illegal operations. The advantages of this solution are as follows: 1. Establishing a unified system solution for detecting behaviors such as throwing express parcels, randomly stacking express parcels, and dangerous operations; 2. Strengthening the detection of operator violations, improving work safety and optimizing express delivery services. Difficulties of this solution include: 1. A large language model to improve task analysis capabilities; 2. A large detection model, which is used to detect tasks; 3. Combining language and large detection models to design a complete solution for real-time identification of personnel illegal operations. In response to the difficulties mentioned above, this application has made the following improvements: 1. Increase the diversity of training data, or use more advanced models to improve the model's language understanding ability; 2. Use richer data for the large detection model and adopt the latest technology to design the detection network; 3. Combine the large language model with the large detection model to make language understanding and detection stronger.

[0059] Based on this, the present application provides an image detection method, apparatus, computer device, storage medium and computer program product. First, the image detection method provided by the present application is described.

[0060] In an exemplary embodiment, as shown in FIG1 , an image detection method is provided. This embodiment uses the method applied to a terminal as an example for illustration; it is understood that the method can also be applied to a server, or to a system including a terminal and a server, and implemented through interaction between the terminal and the server. The terminal can be, but is not limited to, various personal computers, laptops, smartphones, tablet computers, etc.; the server can be implemented as an independent server or a server cluster consisting of multiple servers. In this embodiment, the method includes the following steps:

[0061] Step S101: Acquire detection task description information for the image to be detected.

[0062] The image to be detected may be a picture or video to be detected. For example, in a logistics scenario, the image to be detected may be a picture or video related to logistics, such as a picture or video of a logistics worker.

[0063] The detection task description information may be the content of the detection task described by the user, and is used to determine the detection content that needs to be detected.

[0064] Optionally, the terminal obtains the detection task description information described by the user in a natural language to obtain the detection task description information for the image to be detected.

[0065] Step S102: input the detection task description information into the semantic recognition model to obtain the detection task corresponding to the detection task description information.

[0066] Among them, the semantic recognition model can understand natural language input and can be used to convert the detection task description information into a detection task that can be recognized by a computer. For example, the semantic recognition model can be a large language model.

[0067] The detection task may be a task of detecting what content in an image needs to be detected, as identified by a semantic recognition model based on the detection task description information. For example, the detection task may be a task of monitoring whether a person throws or throws goods.

[0068] Optionally, the terminal inputs the acquired detection task description information into a pre-trained semantic recognition model, understands the detection task description information through the semantic recognition model, and outputs the corresponding specific detection task as the detection task corresponding to the detection task description information.

[0069] Step S103: performing corresponding detection processing on the image to be detected according to the detection task through the image detection model to obtain a detection result for the image to be detected.

[0070] Among them, the image detection model can be used to locate and identify corresponding content in the image to be detected according to the detection task, such as detecting the person who throws the goods or the behavior of throwing the goods from the image to be detected. For example, the image detection model can be a deep learning model.

[0071] The detection result may be the result of processing the detection image by the image detection model according to the detection task, such as the labeled and identified person throwing the goods or the behavior of throwing the goods.

[0072] Optionally, the terminal inputs the detection task output by the semantic recognition model into the image detection model, and through the image detection model, performs corresponding detection processing on the image to be detected according to the detection task, and outputs the detection result corresponding to the detection task as the detection result for the image to be detected.

[0073] In the above-mentioned image detection method, detection task description information for the image to be detected is obtained; the detection task description information is input into a semantic recognition model to obtain the detection task corresponding to the detection task description information; and the image detection model is used to perform corresponding detection processing on the image to be detected according to the detection task to obtain a detection result for the image to be detected. This solution obtains detection task description information for the image to be detected, inputs the detection task description information into a semantic recognition model to obtain the corresponding detection task, and performs detection processing on the image to be detected according to the detection task through the image detection model to obtain a detection result; the semantic recognition model automatically determines the detection task corresponding to the detection task description information, and the image detection model automatically determines the detection result of the image to be detected according to the detection task, thereby improving the efficiency and accuracy of image detection.

[0074] In an exemplary embodiment, as shown in FIG2 , in step S103 , the image detection model performs corresponding detection processing on the image to be detected according to the detection task, and obtains a detection result for the image to be detected, which specifically includes the following:

[0075] Step S201: Perform corresponding detection processing on the image to be detected based on the detection task through the image detection model to determine the detection area and the detection result of the detection area in the image to be detected;

[0076] Step S202: Determine the detection result of the image to be detected based on the detection area and the detection result of the detection area.

[0077] The detection area may be an area of ​​corresponding content located in the image to be detected by the image detection model according to the detection task.

[0078] Among them, the detection result of the detection area can be the recognition result of the image detection model on the content in the detection area. For example, the detection result of the detection area can indicate whether the relevant behavior of the detection task (such as throwing goods) occurs.

[0079] Optionally, the terminal uses the image detection model to perform corresponding detection processing on the image to be detected according to the detection task, and determines the detection area in the image to be detected; performs corresponding detection processing on the detection area in the image to be detected according to the detection task, and obtains the detection result of the detection area; determines the detection result of the image to be detected based on all detection areas and the detection results corresponding to all detection areas.

[0080] The technical solution provided in this embodiment uses an image detection model to locate the detection area in the image according to the detection task, and gives the detection results of the content in the detection area. Finally, the final detection results of the image to be detected are determined based on the detection results of the detection area, which is conducive to improving the accuracy of image detection.

[0081] In an exemplary embodiment, through the image detection model, according to the detection task, corresponding detection processing is performed on the image to be detected, and the detection area in the image to be detected and the detection result of the detection area are determined, which specifically includes the following contents: through the image detection model, according to the detection task, the detection object of the detection task and the detection behavior of the detection task are determined; according to the detection object, corresponding detection processing is performed on the image to be detected to determine the detection area; according to the detection behavior, corresponding detection processing is performed on the detection area to determine the detection result of the detection area.

[0082] The detection object may be a detection target determined according to the detection task, such as a logistics worker.

[0083] The detection behavior may be the type of operation performed by the detection object, such as illegal operation behavior.

[0084] The detection area may be a specific image area of ​​the detection object in the image to be detected.

[0085] The detection result of the detection area may be the result of identification processing of the detection area according to the detection behavior, such as whether an illegal operation (illegal operation behavior / throwing of goods behavior) is identified.

[0086] Optionally, the terminal uses an image detection model to determine the detection object and detection behavior of the detection task according to the detection task (for example, the detection task is to identify illegal operations of logistics staff); according to the detection object, the terminal performs corresponding detection processing on the image to be detected to determine the detection area (for example, the detection area is the area of ​​logistics staff); according to the detection behavior, the terminal performs corresponding detection processing on the detection area (for example, detection processing such as identification of illegal operations) to determine the detection result of the detection area (for example, the detection result of the detection area is used to indicate whether there are illegal operations in the personnel area).

[0087] The technical solution provided in this embodiment determines the detection object and detection behavior through an image detection model, then locates the detection area according to the detection object and performs detection and processing on the content of the detection area, and finally determines the detection result of the detection area, which is conducive to more efficient and accurate determination of the detection result of the detection area, thereby helping to improve the efficiency and accuracy of image detection.

[0088] In an exemplary embodiment, the detection area is subjected to corresponding detection processing according to the detection behavior to determine the detection result of the detection area, which specifically includes the following contents: according to the detection behavior, the detection area is subjected to corresponding detection processing to determine the matching value of the detection behavior corresponding to the detection area; according to the preset matching threshold and the matching value of the detection behavior, the detection result of the detection area is determined; the detection result of the detection area is used to indicate whether there is a detection behavior in the detection area.

[0089] Among them, the matching value can be a numerical value (such as a score) obtained after the detection area content is detected and processed by the image detection model, which is used to measure the degree of matching between the detection area content and the detection behavior. The higher the value, the higher the matching degree.

[0090] The matching threshold may be a preset numerical threshold (such as a score threshold) used to determine whether the matching value is sufficient to determine that a detection behavior exists in the detection area.

[0091] Optionally, the terminal performs corresponding detection processing on the detection area based on the detection behavior, and determines the matching value of the detection behavior corresponding to the detection area, wherein the detection processing will analyze the content of the detection area and give a matching value, which indicates the degree of matching between the content of the detection area and the detection behavior. Here, the detection behavior is to identify a certain illegal operation of the logistics staff, then the matching value can use a numerical value between 0 and 1 to indicate the degree to which the detection area contains the illegal operation. The higher the numerical value, the higher the matching degree. At the same time, there is a preset matching threshold as a judgment standard; according to the preset matching threshold and the matching value of the detection behavior, the detection result of the detection area is determined, wherein, if the matching value is greater than or equal to the matching threshold, it indicates that the content of the detection area highly matches the detection behavior, and the detection result is positive, indicating that the detection area contains the detection behavior. Conversely, if the matching value is less than the matching threshold, it indicates that the matching degree is insufficient, and the detection result is negative, indicating that the detection area does not contain the detection behavior.

[0092] The technical solution provided in this embodiment is conducive to efficiently and accurately determining the detection results of the detection area by comparing the preset matching threshold with the matching value of the detection behavior, thereby helping to improve the efficiency and accuracy of image detection.

[0093] In an exemplary embodiment, the detection result of the image to be detected is determined based on the detection area and the detection result of the detection area, specifically including the following contents: when the detection result of the detection area indicates that there is a detection behavior in the detection area, the detection result of the image to be detected is determined to be that there is abnormal behavior in the image to be detected; when the detection result of the detection area indicates that there is no detection behavior in the detection area, the detection result of the image to be detected is determined to be that there is no abnormal behavior in the image to be detected.

[0094] The abnormal behavior corresponds to the detected behavior, for example, the abnormal behavior is a violation (such as throwing goods).

[0095] Optionally, the terminal determines whether the detection result of the detection area indicates that there is a detection behavior in the detection area. If the detection result of the detection area indicates that there is a detection behavior in the detection area, the detection result of the image to be detected is determined to be that there is an abnormal behavior in the image to be detected (such as the behavior of throwing goods); if the detection result of the detection area indicates that there is no detection behavior in the detection area, the detection result of the image to be detected is determined to be that there is no abnormal behavior in the image to be detected.

[0096] The technical solution provided in this embodiment determines whether there is a detection behavior in the detection area based on the detection results of the detection area, thereby determining whether there is abnormal behavior in the image to be detected, which is conducive to efficiently and accurately determining the detection results of the image to be detected, thereby helping to improve the efficiency and accuracy of image detection.

[0097] In an exemplary embodiment, in step S102, the detection task description information is input into the semantic recognition model to obtain the detection task corresponding to the detection task description information, which specifically includes the following contents: the detection task description information is semantically recognized through the semantic recognition model to obtain the semantic feature information of the detection task description information; the detection task is determined based on the semantic feature information.

[0098] The semantic recognition processing may be a process in which a semantic recognition model analyzes the detection task description information and extracts its semantic features.

[0099] The semantic feature information may be a feature representation of the detection task description information obtained through semantic recognition at the semantic level, such as a large language feature vector.

[0100] Optionally, the terminal inputs the detection task description information into a pre-trained semantic recognition model, performs semantic recognition processing on the detection task description information through the semantic recognition model, extracts the semantic features of the detection task description information, and obtains the semantic feature information of the detection task description information; determines the detection task based on the semantic feature information, for example, matches the semantic feature information output by the semantic recognition model to determine which type of detection task the user is describing, such as matching with the task of "monitoring throwing behavior", and based on the matching results, determines that the user input describes the detection task of "monitoring throwing behavior".

[0101] The technical solution provided in this embodiment inputs the detection requirements described by the user in natural language into a semantic recognition model, uses the semantic recognition model to perform semantic analysis on the description information, extracts its semantic features, and thus determines which type of specific detection task the user is describing. The semantic recognition model is used to understand the semantic intention of the user's description, which is conducive to efficiently and accurately determining the detection task, thereby helping to improve the efficiency and accuracy of image detection.

[0102] In an exemplary embodiment, after performing corresponding detection processing on the image to be detected according to the detection task through the image detection model and obtaining the detection result for the image to be detected, it also includes the following content: when the detection result of the image to be detected is that there is abnormal behavior in the image to be detected, generating corresponding warning information according to the detection result of the image to be detected; and performing corresponding warning processing on the detection object of the detection task according to the warning information.

[0103] The warning information may be a corresponding prompt information generated when the detection result determines that abnormal behavior exists.

[0104] Among them, warning processing can be an operation of prompting or informing the object of the detection task (such as logistics staff) according to the warning information.

[0105] Optionally, when the detection result of the image to be detected is that there is abnormal behavior in the image to be detected, the terminal outputs the detection result, including the detection area (such as the detection box) and the detection result of the detection area (such as the detection category), and generates corresponding warning information (for example, an alarm information that a throwing behavior is detected) based on the detection result of the image to be detected; based on the warning information, corresponding warning processing or reminder processing is performed on the detection object of the detection task.

[0106] The technical solution provided in this embodiment generates warning information and warns the detected object after abnormal behavior is discovered, which is conducive to improving the timeliness of feedback of image detection and stopping abnormal behavior in a timely manner.

[0107] The following is an example of an image detection method provided by the present application. This example uses the method applied to a terminal as an example. The main steps include:

[0108] In the first step, the terminal obtains the detection task description information for the image to be detected.

[0109] In the second step, the terminal performs semantic recognition processing on the detection task description information through the semantic recognition model to obtain semantic feature information of the detection task description information; and determines the detection task based on the semantic feature information.

[0110] In the third step, the terminal uses the image detection model to determine the detection object and detection behavior of the detection task according to the detection task; according to the detection object, the terminal performs corresponding detection processing on the image to be detected and determines the detection area.

[0111] In the fourth step, the terminal performs corresponding detection processing on the detection area based on the detection behavior, and determines the matching value of the detection behavior corresponding to the detection area; based on the preset matching threshold and the matching value of the detection behavior, the detection result of the detection area is determined; the detection result of the detection area is used to indicate whether there is a detection behavior in the detection area.

[0112] In the fifth step, when the detection result of the detection area indicates that there is a detection behavior in the detection area, the terminal determines that the detection result of the image to be detected is that there is abnormal behavior in the image to be detected; when the detection result of the detection area indicates that there is no detection behavior in the detection area, the terminal determines that the detection result of the image to be detected is that there is no abnormal behavior in the image to be detected.

[0113] In the sixth step, when the detection result of the image to be detected is that there is abnormal behavior in the image to be detected, the terminal generates corresponding warning information according to the detection result of the image to be detected; and performs corresponding warning processing on the detection object of the detection task according to the warning information.

[0114] The technical solution provided in this embodiment obtains the detection task description information of the image to be detected, inputs the detection task description information into the semantic recognition model to obtain the corresponding detection task, and performs detection processing on the image to be detected according to the detection task through the image detection model to obtain the detection result; the semantic recognition model automatically determines the detection task corresponding to the detection task description information, and the image detection model automatically determines the detection result of the image to be detected according to the detection task, which is conducive to improving the efficiency and accuracy of image detection.

[0115] The following is an application example to illustrate the image detection method provided by this application. This application example uses the method applied to a terminal as an example. Referring to FIG3 , the main steps include:

[0116] In the first step, the terminal trains a large language model, which uses the latest large language model technology, such as the general pre-trained language model, conversational language model and other frameworks. The model has the ability to understand semantics.

[0117] Among them, the model structure can adopt an open source model or a brand new structural model built by default, and its main method is still based on the transformers (feature extraction / feature conversion / conversion model) structure.

[0118] In the second step, the terminal trains a large detection model using the latest transformer technologies, such as SWIN transformers (a visual detection / segmentation model). By improving these technologies, a large detection model is obtained. This model can perform unified detection and analysis on both videos and images.

[0119] Among them, deeper network layers are used for swin-transformers (for example, the number of network layers of swin-transformers is increased, such as from the original 12 layers to 24 layers, etc., which can extract deeper feature information), and a larger multi-head mechanism is used for attention (for example, the number of heads of the multi-head attention mechanism of swin-transformers is increased. Multi-head attention can extract features from different angles. Increasing the number of heads can more comprehensively capture the spatial structure information in the image). The training here uses mixed input. The input includes the output of a large language model + pictures or videos. The network output includes two branches: detection box output and detection classification. Detection is used to detect people related, and classification is used to judge human behavior. For example, the detection classification can be the category label predicted by the detection model for each detection box (such as the detection result of the detection area).

[0120] In the third step, the terminal combines the language model and the detection model.

[0121] 4 , the language model and the detection model are combined through a fully connected layer. Specifically, the image is input into the main structure of the detection model, and the image feature information is output. Through the semantic feature information (large language feature vector) and image feature information, the detection area (such as the detection box) and the detection results of the detection area (such as the detection classification) are output.

[0122] In the fourth step, the terminal monitors the task according to the user input. For example, the user can describe the task arbitrarily, such as: help me monitor the throwing behavior (help me monitor whether anyone throws the goods), help me monitor the throwing and random stacking of express parcels.

[0123] Optionally, the terminal inputs the task into the large language model, analyzes the task through the large language model, outputs the task content, and inputs the task content into the detection large model. The detection large model detects the image (such as video, picture) in real time, outputs the real-time detection results, and determines whether there is a violation. If not, real-time detection is performed. If so, a warning is issued.

[0124] For example, after model training, the terminal will obtain detection-related thresholds through large amounts of data. For example, after the input task and a video pass through the detection model, a detection classification will be obtained. This classification will get a score. This score can be used to determine whether the user-given task exists in the video. If there is a corresponding detection box, the task is detected. Assuming that the task is illegal throwing, the detected location will be the area where people throw goods in a video. Referring to Figure 5, the output result can be an output video of people throwing goods (composed of frames of pictures, including the detection area, and some pictures are omitted).

[0125] The technical solution provided by this application example can switch at any time or perform multiple task detections simultaneously; reduce manual monitoring and achieve fully automatic monitoring; and improve the efficiency and accuracy of image detection.

[0126] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0127] Based on the same inventive concept, the present application also provides an image detection device for implementing the aforementioned image detection method. The solution provided by this device is similar to the solution described in the aforementioned method. Therefore, the specific limitations in one or more of the following embodiments of the image detection device can be found in the above-mentioned limitations on the image detection method and will not be further elaborated here.

[0128] In an exemplary embodiment, as shown in FIG6 , an image detection device is provided. The device 600 may include:

[0129] The information acquisition module 601 is used to obtain the detection task description information for the image to be detected;

[0130] An information input module 602 is used to input the detection task description information into the semantic recognition model to obtain the detection task corresponding to the detection task description information;

[0131] The image processing module 603 is used to perform corresponding detection processing on the image to be detected according to the detection task through the image detection model to obtain the detection result for the image to be detected.

[0132] In an exemplary embodiment, the image processing module 603 is also used to perform corresponding detection processing on the image to be detected according to the detection task through the image detection model, determine the detection area and the detection result of the detection area in the image to be detected; and determine the detection result of the image to be detected based on the detection area and the detection result of the detection area.

[0133] In an exemplary embodiment, the image processing module 603 is also used to determine the detection object and detection behavior of the detection task according to the detection task through the image detection model; perform corresponding detection processing on the image to be detected according to the detection object to determine the detection area; perform corresponding detection processing on the detection area according to the detection behavior to determine the detection result of the detection area.

[0134] In an exemplary embodiment, the image processing module 603 is also used to perform corresponding detection processing on the detection area based on the detection behavior, and determine the matching value of the detection behavior corresponding to the detection area; determine the detection result of the detection area based on the preset matching threshold and the matching value of the detection behavior; the detection result of the detection area is used to indicate whether there is a detection behavior in the detection area.

[0135] In an exemplary embodiment, the image processing module 603 is also used to determine that the detection result of the image to be detected is that there is abnormal behavior in the image to be detected when the detection result of the detection area indicates that there is detection behavior in the detection area; and to determine that the detection result of the image to be detected is that there is no abnormal behavior in the image to be detected when the detection result of the detection area indicates that there is no detection behavior in the detection area.

[0136] In an exemplary embodiment, the information input module 602 is further configured to perform semantic recognition processing on the detection task description information through a semantic recognition model to obtain semantic feature information of the detection task description information; and determine the detection task based on the semantic feature information.

[0137] In an exemplary embodiment, the device 600 also includes: an information generation module, which is used to generate corresponding warning information based on the detection result of the image to be detected when the detection result of the image to be detected is that abnormal behavior exists in the image to be detected; and perform corresponding warning processing on the detection object of the detection task based on the warning information.

[0138] Each module in the above-mentioned image detection device can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of one or more processors in a computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that one or more processors can call and execute the corresponding operations of each module.

[0139] In an exemplary embodiment, a computer device is provided, which may be a terminal. A diagram of its internal structure may be shown in FIG7 . The computer device includes one or more processors, memory, an input / output interface, a communication interface, a display unit, and an input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are connected to the system bus via the input / output interface. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer-readable instructions. The internal memory provides an environment for the operation of the operating system and computer-readable instructions in the non-volatile storage medium. The input / output interface of the computer device is configured to exchange information between the processor and an external device. The communication interface of the computer device is configured to communicate with an external terminal via wired or wireless communication, where the wireless communication may be achieved via Wi-Fi, a mobile cellular network, NFC (near-field communication), or other technologies. When executed by the processor, the computer-readable instructions implement an image detection method. The display unit of the computer device is configured to produce a visually visible image and may be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device casing, or an external keyboard, touchpad or mouse.

[0140] Those skilled in the art will understand that the structure shown in FIG7 is merely a block diagram of a portion of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different arrangement of components.

[0141] In an exemplary embodiment, a computer device is also provided, including a memory and one or more processors, wherein the memory stores computer-readable instructions, and the one or more processors implement the steps in the above-mentioned method embodiments when executing the computer-readable instructions.

[0142] In an exemplary embodiment, one or more computer-readable storage media are provided, on which computer-readable instructions are stored. When the computer-readable instructions are executed by one or more processors, the steps in the above-mentioned method embodiments are implemented.

[0143] In an exemplary embodiment, a computer program product is provided, including computer-readable instructions, which implement the steps of the above-mentioned method embodiments when executed by one or more processors.

[0144] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through computer-readable instructions. The computer-readable instructions can be stored in a non-volatile computer-readable storage medium. When the computer-readable instructions are executed, they can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processor involved in the various embodiments provided herein may be, but are not limited to, a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic unit, a data processing logic unit based on quantum computing, and the like.

[0145] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0146] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. An image detection method, characterized in that, Applied to a terminal; the method includes: Obtain detection task description information for an image to be detected; Input the detection task description information into a semantic recognition model to obtain a detection task corresponding to the detection task description information; Through an image detection model, perform corresponding detection processing on the image to be detected according to the detection task to obtain a detection result for the image to be detected.

2. The method according to claim 1, wherein The step of "Through an image detection model, perform corresponding detection processing on the image to be detected according to the detection task to obtain a detection result for the image to be detected" includes: Through the image detection model, perform corresponding detection processing on the image to be detected according to the detection task, and determine a detection area in the image to be detected and a detection result of the detection area; Determine the detection result of the image to be detected according to the detection area and the detection result of the detection area.

3. The method according to claim 2, wherein The step of "Through the image detection model, perform corresponding detection processing on the image to be detected according to the detection task, and determine a detection area in the image to be detected and a detection result of the detection area" includes: Through the image detection model, determine a detection object of the detection task and a detection behavior of the detection task according to the detection task; Perform corresponding detection processing on the image to be detected according to the detection object to determine the detection area; Perform corresponding detection processing on the detection area according to the detection behavior to determine the detection result of the detection area.

4. The method according to claim 3, wherein The step of "Perform corresponding detection processing on the detection area according to the detection behavior to determine the detection result of the detection area" includes: Perform corresponding detection processing on the detection area according to the detection behavior to determine a matching value of the detection behavior corresponding to the detection area; Determine the detection result of the detection area according to a preset matching threshold and the matching value of the detection behavior; the detection result of the detection area is used to indicate whether the detection behavior exists in the detection area.

5. The method according to claim 4, wherein The step of "Determine the detection result of the image to be detected according to the detection area and the detection result of the detection area" includes: When the detection result of the detection area indicates that the detection behavior exists in the detection area, determine that the detection result of the image to be detected is that there is an abnormal behavior in the image to be detected; When the detection result of the detection area indicates that the detection behavior does not exist in the detection area, determine that the detection result of the image to be detected is that there is no abnormal behavior in the image to be detected.

6. The method according to claim 1, characterized in that The step of "Input the detection task description information into a semantic recognition model to obtain a detection task corresponding to the detection task description information" includes: Through the semantic recognition model, perform semantic recognition processing on the detection task description information to obtain semantic feature information of the detection task description information; Determine the detection task according to the semantic feature information.

7. The method according to claim 1, characterized in that, After obtaining the detection result for the image to be detected through the image detection model by performing corresponding detection processing on the image to be detected according to the detection task, the method further includes: In the case that the detection result of the image to be detected indicates the existence of abnormal behavior in the image to be detected, generate a corresponding warning message according to the detection result of the image to be detected; Perform corresponding warning processing on the detection object of the detection task according to the warning message.

8. An image detection device, characterized in that, The device includes: An information acquisition module, configured to acquire detection task description information for an image to be detected; An information input module, configured to input the detection task description information into a semantic recognition model to obtain a detection task corresponding to the detection task description information; An image processing module, configured to perform corresponding detection processing on the image to be detected according to the detection task through an image detection model to obtain a detection result for the image to be detected.

9. A computer device, comprising a memory and one or more processors, the memory storing computer-readable instructions, characterized in that, When the one or more processors execute the computer-readable instructions, the steps of the method according to any one of claims 1 to 7 are implemented.

10. One or more computer-readable storage media having computer-readable instructions stored thereon, wherein, When the computer-readable instructions are executed by one or more processors, the steps of the method according to any one of claims 1 to 7 are implemented.

11. A computer program product, comprising computer-readable instructions, characterized in that, When the computer-readable instructions are executed by one or more processors, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Image content analysis method and device, equipment and medium

    CN116824278A

  • Power image anomaly detection method and device, computer equipment and storage medium

    CN117132763A

  • Target detection method, related device, equipment, system and storage medium

    CN117197423A

  • Multi-modal target detection method and device, computer equipment and storage medium

    CN117253245A

  • Attention-based explanations for artificial intelligence behavior

    US20190370587A1

Cited By

  • Container dangerous goods identification method and system based on visual language large model

    CN121330377A

  • Container dangerous goods identification method and system based on visual language large model

    CN121330377B