Image detection method and device, computer equipment and storage medium
By acquiring image detection task description information and using semantic recognition and image detection models for automated detection, the problem of traditional image detection inefficiency is solved, and more efficient and accurate detection results are achieved.
Patent Information
- Application Number
- CN202311872333.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-30
- Publication Date
- 2025-07-01
AI Technical Summary
Traditional image detection methods require manual annotation, resulting in inefficiency.
By obtaining the detection task description information of the image to be detected, it is input into the semantic recognition model to obtain the detection task, and then using the image detection model to detect the image according to the detection task, and generate the detection result.
Automatic determination of detection tasks and image detection results is realized, and the efficiency and accuracy of image detection are improved.
Smart Images

Figure CN120236283A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and particularly to an image detection method, apparatus, computer device, storage medium, and computer program product. Background Art
[0002] With the development of computer technology, image detection has important applications in many fields. By analyzing traditional image detection methods, it can be learned that the efficiency of image detection is relatively low. Therefore, how to perform image detection efficiently has become an important research direction.
[0003] Traditional technologies usually require manual annotation of images before image detection; however, performing image detection in this way requires a lot of manual processing time, resulting in low efficiency of image detection. Summary of the Invention
[0004] Based on this, it is necessary to provide an image detection method, apparatus, computer device, computer-readable storage medium, and computer program product that can improve the efficiency of image detection for the above technical problems.
[0005] In a first aspect, this application provides an image detection method. The method includes:
[0006] Obtain detection task description information for an image to be detected;
[0007] Input the detection task description information into a semantic recognition model to obtain a detection task corresponding to the detection task description information;
[0008] Through an image detection model, perform corresponding detection processing on the image to be detected according to the detection task, and obtain a detection result for the image to be detected.
[0009] In one embodiment, the step of performing corresponding detection processing on the image to be detected according to the detection task through the image detection model to obtain a detection result for the image to be detected includes:
[0010] Through the image detection model, perform corresponding detection processing on the image to be detected according to the detection task, and determine a detection area in the image to be detected and a detection result of the detection area;
[0011] Determine the detection result of the image to be detected according to the detection area and the detection result of the detection area.
[0012] In one embodiment, through the image detection model, according to the detection task, performing corresponding detection processing on the image to be detected, and determining a detection area in the image to be detected and a detection result of the detection area, includes:
[0013] Through the image detection model, according to the detection task, determining a detection object of the detection task and a detection behavior of the detection task;
[0014] According to the detection object, performing corresponding detection processing on the image to be detected to determine the detection area;
[0015] According to the detection behavior, performing corresponding detection processing on the detection area to determine a detection result of the detection area.
[0016] In one embodiment, the performing corresponding detection processing on the detection area according to the detection behavior to determine a detection result of the detection area includes:
[0017] According to the detection behavior, performing corresponding detection processing on the detection area to determine a matching value of the detection behavior corresponding to the detection area;
[0018] According to a preset matching threshold and the matching value of the detection behavior, determining a detection result of the detection area; the detection result of the detection area is used to indicate whether the detection behavior exists in the detection area.
[0019] In one embodiment, the determining a detection result of the image to be detected according to the detection area and the detection result of the detection area includes:
[0020] When the detection result of the detection area indicates that the detection behavior exists in the detection area, determining that the detection result of the image to be detected is that there is an abnormal behavior in the image to be detected;
[0021] When the detection result of the detection area indicates that the detection behavior does not exist in the detection area, determining that the detection result of the image to be detected is that there is no abnormal behavior in the image to be detected.
[0022] In one embodiment, the inputting the detection task description information into a semantic recognition model to obtain a detection task corresponding to the detection task description information includes:
[0023] Through the semantic recognition model, performing semantic recognition processing on the detection task description information to obtain semantic feature information of the detection task description information;
[0024] According to the semantic feature information, determining the detection task.
[0025] In one embodiment, after performing corresponding detection processing on the image to be detected according to the detection task through an image detection model and obtaining a detection result for the image to be detected, the method further includes:
[0026] When the detection result of the image to be detected indicates that there is an abnormal behavior in the image to be detected, generating a corresponding warning message according to the detection result of the image to be detected;
[0027] Performing corresponding warning processing on the detection object of the detection task according to the warning message.
[0028] In a second aspect, the present application further provides an image detection device. The device includes:
[0029] An information acquisition module, configured to acquire detection task description information for an image to be detected;
[0030] An information input module, configured to input the detection task description information into a semantic recognition model to obtain a detection task corresponding to the detection task description information;
[0031] An image processing module, configured to perform corresponding detection processing on the image to be detected according to the detection task through an image detection model to obtain a detection result for the image to be detected.
[0032] In a third aspect, the present application further provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0033] Acquiring detection task description information for an image to be detected;
[0034] Inputting the detection task description information into a semantic recognition model to obtain a detection task corresponding to the detection task description information;
[0035] Performing corresponding detection processing on the image to be detected according to the detection task through an image detection model to obtain a detection result for the image to be detected.
[0036] In a fourth aspect, the present application further provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the following steps are implemented:
[0037] Acquiring detection task description information for an image to be detected;
[0038] Input the detection task description information into a semantic recognition model to obtain the detection task corresponding to the detection task description information;
[0039] Through an image detection model, perform corresponding detection processing on the image to be detected according to the detection task to obtain a detection result for the image to be detected.
[0040] In a fifth aspect, the present application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, implements the following steps:
[0041] Obtain detection task description information for an image to be detected;
[0042] Input the detection task description information into a semantic recognition model to obtain the detection task corresponding to the detection task description information;
[0043] Through an image detection model, perform corresponding detection processing on the image to be detected according to the detection task to obtain a detection result for the image to be detected.
[0044] The above image detection method, device, computer device, storage medium, and computer program product obtain detection task description information for an image to be detected; input the detection task description information into a semantic recognition model to obtain the detection task corresponding to the detection task description information; through an image detection model, perform corresponding detection processing on the image to be detected according to the detection task to obtain a detection result for the image to be detected. This solution obtains the detection task description information of the image to be detected, inputs the detection task description information into the semantic recognition model to obtain the corresponding detection task, and through the image detection model, performs detection processing on the image to be detected according to the detection task to obtain the detection result; enabling the automatic determination of the detection task corresponding to the detection task description information through the semantic recognition model, and automatically determining the detection result of the image to be detected according to the detection task through the image detection model, thereby facilitating the improvement of the efficiency and accuracy of image detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0046] Figure 1 It is a schematic flowchart of an image detection method in an embodiment;
[0047] Figure 2Schematic flowchart of steps for determining detection results of an image to be detected in an embodiment;
[0048] Figure 3 Schematic flowchart of an image detection method in another embodiment;
[0049] Figure 4 Schematic flowchart of model processing in an embodiment;
[0050] Figure 5 Schematic diagram of output video in an embodiment;
[0051] Figure 6 Block diagram of the structure of an image detection device in an embodiment;
[0052] Figure 7 Internal structure diagram of a computer device in an embodiment. Detailed implementation manners
[0053] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0054] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with relevant regulations.
[0055] With the development of globalization, the scale of the logistics industry has been continuously expanding, and the workload of logistics staff has also been increasing. However, due to reasons such as high work pressure and complex working environment, the problem of illegal operations by logistics staff has gradually emerged. These illegal operations not only affect logistics efficiency, cause damage to express deliveries, but also pose a threat to logistics safety. Therefore, how to effectively identify and prevent the illegal operations of logistics staff has become an urgent problem to be solved in the logistics industry. For this reason, this application provides a system for identifying illegal operations of logistics staff based on a large model. This system uses big data and artificial intelligence technologies to monitor and analyze the behaviors of logistics staff in real time, so as to achieve the timely discovery and prevention of illegal operations. The advantages of this solution are as follows: 1. Establish a unified system solution for detecting behaviors such as throwing express deliveries, stacking express deliveries randomly, and dangerous operations; 2. Strengthen the detection of illegal behaviors of operators to improve work safety and optimize express delivery services. The difficulties of this solution are: 1. Language large model to improve the analysis ability of tasks; 2. Detection large model, which is used for detecting tasks; 3. Combine the language and detection large models to design a complete solution to achieve real-time identification of illegal operations of personnel. In response to the above-mentioned difficulties, this application has made the following improvements: 1. Increase the diversity of training data, or use more advanced models to improve the language understanding ability of the model; 2. Use richer data for the detection large model and design the detection network using the latest technologies; 3. Combine the large language model with the detection large model to make language understanding and detection stronger.
[0056] Based on this, this application provides an image detection method, device, computer device, storage medium, and computer program product. First, the image detection method provided by this application will be described.
[0057] In an exemplary embodiment, as Figure 1 shown, an image detection method is provided. This embodiment takes the application of this method to a terminal as an example for illustration; it can be understood that this method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. Among them, the terminal can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, etc.; the server can be implemented by an independent server or a server cluster composed of multiple servers. In this embodiment, the method includes the following steps:
[0058] Step S101, obtain detection task description information for the image to be detected.
[0059] Among them, the image to be detected can be a picture or video to be detected. For example, in a logistics scenario, the image to be detected can be a picture or video related to logistics, such as a picture or video of logistics staff.
[0060] Among them, the detection task description information can be the content of the detection task described by the user, and is used to determine the detection content to be detected.
[0061] Optionally, the terminal obtains the detection task description information described by the user in the form of natural language, and obtains the detection task description information for the image to be detected.
[0062] Step S102: Input the detection task description information into the semantic recognition model to obtain the detection task corresponding to the detection task description information.
[0063] Among them, the semantic recognition model can understand natural language input and can be used to convert the detection task description information into a detection task recognizable by a computer. For example, the semantic recognition model can be a large language model.
[0064] Among them, the detection task can be a task of what content in the image to be detected recognized by the semantic recognition model according to the detection task description information. For example, the detection task can be a task of monitoring whether a person throws goods.
[0065] Optionally, the terminal inputs the obtained detection task description information into a pre-trained semantic recognition model, and the semantic recognition model understands the detection task description information and outputs the corresponding specific detection task as the detection task corresponding to the detection task description information.
[0066] Step S103: Through the image detection model, perform corresponding detection processing on the image to be detected according to the detection task, and obtain the detection result for the image to be detected.
[0067] Among them, the image detection model can be used to locate and identify the corresponding content in the image to be detected according to the detection task. For example, detect the person throwing goods or the behavior of throwing goods from the image to be detected. For example, the image detection model can be a deep learning model.
[0068] Among them, the detection result can be the result of the image detection model processing the image to be detected according to the detection task. For example, the person throwing goods or the behavior of throwing goods marked and recognized.
[0069] Optionally, the terminal inputs the detection task output by the semantic recognition model into the image detection model, and through the image detection model, performs corresponding detection processing on the image to be detected according to the detection task, and outputs the detection result corresponding to the detection task as the detection result for the image to be detected.
[0070] In the above image detection method, detection task description information for the image to be detected is obtained; the detection task description information is input into a semantic recognition model to obtain a detection task corresponding to the detection task description information; through an image detection model, corresponding detection processing is performed on the image to be detected according to the detection task to obtain a detection result for the image to be detected. This solution obtains the detection task description information of the image to be detected, inputs the detection task description information into the semantic recognition model to obtain the corresponding detection task, and through the image detection model, performs detection processing on the image to be detected according to the detection task to obtain the detection result; enabling the automatic determination of the detection task corresponding to the detection task description information through the semantic recognition model, and automatically determining the detection result of the image to be detected according to the detection task through the image detection model, thereby facilitating the improvement of the efficiency and accuracy of image detection.
[0071] In an exemplary embodiment, as Figure 2 shown, in step S103, through the image detection model, corresponding detection processing is performed on the image to be detected according to the detection task to obtain a detection result for the image to be detected, which specifically includes the following content:
[0072] Step S201, through the image detection model, corresponding detection processing is performed on the image to be detected according to the detection task to determine the detection region in the image to be detected and the detection result of the detection region;
[0073] Step S202, according to the detection region and the detection result of the detection region, determine the detection result of the image to be detected.
[0074] Among them, the detection region may be a region of the corresponding content located by the image detection model in the image to be detected according to the detection task.
[0075] Among them, the detection result of the detection region may be the recognition result of the content in the detection region by the image detection model. For example, the detection result of the detection region may represent the result of whether a relevant behavior of the detection task (such as the behavior of throwing goods) occurs.
[0076] Optionally, the terminal, through the image detection model, performs corresponding detection processing on the image to be detected according to the detection task to determine the detection region in the image to be detected; performs corresponding detection processing on the detection region in the image to be detected according to the detection task to obtain the detection result of the detection region; and determines the detection result of the image to be detected according to all detection regions and the detection results corresponding to all detection regions.
[0077] The technical solution provided in this embodiment uses an image detection model to locate a detection area in an image according to a detection task, and gives the detection result of the content in the detection area. Finally, the final detection result of the image to be detected is determined based on the detection result of the detection area, which helps to improve the accuracy of image detection.
[0078] In an exemplary embodiment, through an image detection model, according to a detection task, corresponding detection processing is performed on the image to be detected, and the detection area and the detection result of the detection area in the image to be detected are determined. The specific content is as follows: Through the image detection model, according to the detection task, the detection object and the detection behavior of the detection task are determined; according to the detection object, corresponding detection processing is performed on the image to be detected to determine the detection area; according to the detection behavior, corresponding detection processing is performed on the detection area to determine the detection result of the detection area.
[0079] Among them, the detection object can be the detection target determined according to the detection task, such as a logistics worker.
[0080] Among them, the detection behavior can be the type of operation behavior performed by the detection object, such as a violation operation behavior.
[0081] Among them, the detection area can be the specific image area of the detection object in the image to be detected.
[0082] Among them, the detection result of the detection area can be the result after identifying and processing the detection area according to the detection behavior, such as whether a violation operation (violation operation behavior / throwing goods behavior) is identified.
[0083] Optionally, the terminal uses an image detection model to determine the detection object and the detection behavior of the detection task according to the detection task (for example, the detection task is to identify the violation operation of a logistics worker); according to the detection object, corresponding detection processing is performed on the image to be detected to determine the detection area (for example, the detection area is the area of the logistics worker); according to the detection behavior, corresponding detection processing (such as performing detection processing such as violation operation identification) is performed on the detection area to determine the detection result of the detection area (for example, the detection result of the detection area is used to indicate whether there is a violation operation in the personnel area).
[0084] The technical solution provided in this embodiment uses an image detection model to determine the detection object and the detection behavior, then locates the detection area according to the detection object and performs detection processing on the content of the detection area, and finally determines the detection result of the detection area, which helps to more efficiently and accurately determine the detection result of the detection area, thereby helping to improve the efficiency and accuracy of image detection.
[0085] In an exemplary embodiment, according to the detection behavior, corresponding detection processing is performed on the detection area to determine the detection result of the detection area, which specifically includes the following: According to the detection behavior, corresponding detection processing is performed on the detection area to determine the matching value of the detection behavior corresponding to the detection area; According to the preset matching threshold and the matching value of the detection behavior, the detection result of the detection area is determined; The detection result of the detection area is used to indicate whether there is a detection behavior in the detection area.
[0086] Among them, the matching value can be a numerical value (such as a score) obtained by detecting the content of the detection area through an image detection model, which is used to measure the matching degree between the content of the detection area and the detection behavior. The higher the numerical value, the higher the matching degree.
[0087] Among them, the matching threshold can be a preset numerical threshold (such as a score threshold), which is used to determine whether the matching value is sufficient to consider that there is a detection behavior in the detection area.
[0088] Optionally, the terminal performs corresponding detection processing on the detection area according to the detection behavior to determine the matching value of the detection behavior corresponding to the detection area. Among them, the detection processing analyzes the content of the detection area and gives a matching value, which represents the matching degree between the content of the detection area and the detection behavior. Here, the detection behavior is to identify a certain violation operation of logistics staff. Then, the matching value can be represented by a numerical value between 0 and 1 to indicate the degree to which the detection area contains this violation operation. The higher the numerical value, the higher the matching degree. At the same time, there is a preset matching threshold as the judgment standard; According to the preset matching threshold and the matching value of the detection behavior, the detection result of the detection area is determined. Among them, if the matching value is greater than or equal to the matching threshold, it indicates that the content of the detection area highly matches the detection behavior, and the detection result is affirmative, indicating that the detection area contains this detection behavior. On the contrary, if the matching value is less than the matching threshold, it indicates that the matching degree is insufficient, and the detection result is negative, indicating that the detection area does not contain this detection behavior.
[0089] The technical solution provided in this embodiment is beneficial to efficiently and accurately determine the detection result of the detection area by comparing the preset matching threshold with the matching value of the detection behavior, thereby being beneficial to improving the efficiency and accuracy of image detection.
[0090] In an exemplary embodiment, according to the detection area and the detection result of the detection area, the detection result of the image to be detected is determined, which specifically includes the following: When the detection result of the detection area indicates that there is a detection behavior in the detection area, it is determined that the detection result of the image to be detected is that there is an abnormal behavior in the image to be detected; When the detection result of the detection area indicates that there is no detection behavior in the detection area, it is determined that the detection result of the image to be detected is that there is no abnormal behavior in the image to be detected.
[0091] Among them, the abnormal behavior corresponds to the detection behavior. For example, the abnormal behavior is a violation behavior (such as the behavior of throwing goods).
[0092] Optionally, the terminal determines whether the detection result of the detection area indicates the existence of a detection behavior in the detection area. When the detection result of the detection area indicates the existence of a detection behavior in the detection area, it is determined that the detection result of the image to be detected is that there is an abnormal behavior (such as the behavior of throwing goods) in the image to be detected; when the detection result of the detection area indicates the non-existence of a detection behavior in the detection area, it is determined that the detection result of the image to be detected is that there is no abnormal behavior in the image to be detected.
[0093] The technical solution provided in this embodiment determines whether there is a detection behavior in the detection area according to the detection result of the detection area, so as to determine whether there is an abnormal behavior in the image to be detected, which is beneficial to efficiently and accurately determining the detection result of the image to be detected, and thus beneficial to improving the efficiency and accuracy of image detection.
[0094] In an exemplary embodiment, in step S102, the detection task description information is input into the semantic recognition model to obtain the detection task corresponding to the detection task description information, which specifically includes the following content: through the semantic recognition model, semantic recognition processing is performed on the detection task description information to obtain the semantic feature information of the detection task description information; according to the semantic feature information, the detection task is determined.
[0095] Among them, the semantic recognition processing may be the processing in which the semantic recognition model analyzes and extracts the semantic features of the detection task description information.
[0096] Among them, the semantic feature information may be the feature representation of the detection task description information at the semantic level obtained through semantic recognition, such as a large language feature vector.
[0097] Optionally, the terminal inputs the detection task description information into a pre-trained semantic recognition model. Through the semantic recognition model, semantic recognition processing is performed on the detection task description information to extract the semantic features of the detection task description information, and the semantic feature information of the detection task description information is obtained; according to the semantic feature information, the detection task is determined. For example, according to the semantic feature information output by the semantic recognition model, a match is made to determine which type of detection task the user describes. For example, it matches the task of "monitoring throwing behavior", and according to the matching result, it is determined that the detection task described by the user input is the "monitoring throwing behavior" detection task.
[0098] The technical solution provided in this embodiment inputs the detection requirements described by the user in natural language into the semantic recognition model, uses the semantic recognition model to perform semantic analysis on the description information, extracts its semantic features, and thus determines which specific type of detection task the user describes, enabling the use of the semantic recognition model to understand the semantic intention described by the user, which is beneficial to efficiently and accurately determining the detection task, and thus beneficial to improving the efficiency and accuracy of image detection.
[0099] In an exemplary embodiment, after performing corresponding detection processing on the image to be detected according to the detection task through the image detection model and obtaining the detection result for the image to be detected, the following content is further included: when the detection result of the image to be detected indicates that there is an abnormal behavior in the image to be detected, generating a corresponding warning message according to the detection result of the image to be detected; and performing corresponding warning processing on the detection object of the detection task according to the warning message.
[0100] Among them, the warning message can be a corresponding prompt message generated when it is determined that there is an abnormal behavior in the detection result.
[0101] Among them, the warning processing can be an operation of prompting or informing the object of the detection task (such as a logistics staff member) according to the warning message.
[0102] Optionally, when the detection result of the image to be detected indicates that there is an abnormal behavior in the image to be detected, the terminal outputs the detection result, including the detection area (such as a detection frame) and the detection result of the detection area (such as a detection category), generates a corresponding warning message according to the detection result of the image to be detected (for example, an alarm message for detecting a throwing behavior); and performs corresponding warning processing or reminder processing on the detection object of the detection task according to the warning message.
[0103] The technical solution provided in this embodiment generates a warning message and performs a warning on the detection object after discovering an abnormal behavior, which is beneficial to improving the feedback timeliness of image detection and promptly stopping the abnormal behavior.
[0104] The following uses an embodiment to illustrate the image detection method provided in this application. This embodiment takes the application of this method to a terminal as an example for illustration. The main steps include:
[0105] First step, the terminal obtains the detection task description information for the image to be detected.
[0106] Second step, the terminal performs semantic recognition processing on the detection task description information through the semantic recognition model to obtain the semantic feature information of the detection task description information; and determines the detection task according to the semantic feature information.
[0107] In the third step, the terminal uses an image detection model to determine the detection object and detection behavior of the detection task according to the detection task; according to the detection object, corresponding detection processing is performed on the image to be detected to determine the detection area.
[0108] In the fourth step, the terminal performs corresponding detection processing on the detection area according to the detection behavior to determine the matching value of the detection behavior corresponding to the detection area; according to the preset matching threshold and the matching value of the detection behavior, the detection result of the detection area is determined; the detection result of the detection area is used to indicate whether there is a detection behavior in the detection area.
[0109] In the fifth step, when the detection result of the detection area indicates that there is a detection behavior in the detection area, the terminal determines that the detection result of the image to be detected is that there is an abnormal behavior in the image to be detected; when the detection result of the detection area indicates that there is no detection behavior in the detection area, the terminal determines that the detection result of the image to be detected is that there is no abnormal behavior in the image to be detected.
[0110] In the sixth step, when the detection result of the image to be detected is that there is an abnormal behavior in the image to be detected, the terminal generates a corresponding warning message according to the detection result of the image to be detected; according to the warning message, corresponding warning processing is performed on the detection object of the detection task.
[0111] The technical solution provided in this embodiment obtains the detection task description information of the image to be detected, inputs the detection task description information into the semantic recognition model to obtain the corresponding detection task, and uses the image detection model to perform detection processing on the image to be detected according to the detection task to obtain the detection result; enabling the semantic recognition model to automatically determine the detection task corresponding to the detection task description information, and the image detection model to automatically determine the detection result of the image to be detected according to the detection task, which is beneficial to improving the efficiency and accuracy of image detection.
[0112] The following uses an application example to illustrate the image detection method provided in this application. This application example takes the application of this method to a terminal as an example. Refer to Figure 3 , and the main steps include:
[0113] In the first step, the terminal trains a large language model. This model adopts the latest large language model technology, such as frameworks like general pre-trained language models and dialogue language models. This model has the ability to understand semantics.
[0114] Among them, the model structure can adopt an open-source model or a completely new structure model set up in advance. Its main method is still mainly based on the transformers (feature extraction / feature transformation / transformation model) structure.
[0115] In the second step, the terminal trains a detection large model, which adopts the latest transformers technology, such as swin-transformers (visual detection / segmentation model) and other technologies. By improving these technologies, a large detection model is obtained. This model can perform unified detection and analysis on videos and pictures.
[0116] Among them, for swin-transformers, deeper network layers are used (for example, increasing the number of network layers of swin-transformers, such as increasing from the original 12 layers to 24 layers, etc., which can extract deeper feature information), and a larger multi-head mechanism is used for attention (for example, increasing the number of heads of the multi-head attention mechanism of swin-transformers. Multi-head attention can extract features from different angles, and increasing the number of heads can more comprehensively capture the spatial structure information in the image). Here, the training uses mixed input, and the input includes the output of the large language model + pictures or videos. The network output includes two branches: detection box output and detection classification. Detection is used to detect people-related, and classification is used to judge human behavior. For example, the detection classification can be the category label predicted by the detection model for each detection box (such as the detection result of the detection area).
[0117] In the third step, the terminal combines the language large model and the detection large model.
[0118] Among them, referring to Figure 4 , the language large model and the detection large model are combined through a fully connected layer. Specifically, the image is input into the backbone part of the detection model structure, and image feature information is output. Through semantic feature information (large language feature vector) and image feature information, the detection area (such as a detection box) and the detection result of the detection area (such as detection classification) are output.
[0119] In the fourth step, the terminal monitors the task according to the user input. For example, the user can arbitrarily describe the task, such as: Help me monitor the throwing behavior (help me monitor whether there is anyone throwing goods), Help me monitor the throwing and random stacking of express parcels.
[0120] Optionally, the terminal inputs the task into the large language model, analyzes the task through the large language model, outputs the task content, inputs the task content into the detection large model, and the detection large model performs real-time detection on images (such as videos, pictures), outputs the real-time detection result, judges whether it is illegal. If not, then perform real-time detection. If so, then give a warning.
[0121] For example, after model training, the terminal will obtain detection-related thresholds through a large amount of data. For example, after the input task and a video are passed through the detection model, a detection classification will be obtained. This classification will get a score. This score can be used to determine whether the user-given task exists in the video. If there is a corresponding detection box, the task is detected. Assuming that the task is illegal throwing, the detected location will be the area where people throw goods in a video. Figure 5 The output result can be an output video of a person throwing goods (composed of frames of pictures, including the detection area, and some pictures are omitted).
[0122] The technical solution provided by this application example can switch at any time or perform multiple task detections simultaneously; reduce manual monitoring and achieve fully automatic monitoring; and improve the efficiency and accuracy of image detection.
[0123] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.
[0124] Based on the same inventive concept, the embodiment of the present application also provides an image detection device for implementing the above-mentioned image detection method. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above-mentioned method, so the specific limitations in one or more image detection device embodiments provided below can refer to the limitations on the image detection method above, and will not be repeated here.
[0125] In an exemplary embodiment, Figure 6 As shown, an image detection device is provided, and the device 600 may include:
[0126] The information acquisition module 601 is used to acquire the detection task description information for the image to be detected;
[0127] An information input module 602 is used to input the detection task description information into the semantic recognition model to obtain the detection task corresponding to the detection task description information;
[0128] The image processing module 603 is configured to perform corresponding detection processing on the image to be detected according to the detection task through an image detection model, and obtain a detection result for the image to be detected.
[0129] In an exemplary embodiment, the image processing module 603 is further configured to perform corresponding detection processing on the image to be detected according to the detection task through an image detection model, determine a detection area in the image to be detected and a detection result of the detection area; and determine a detection result of the image to be detected according to the detection area and the detection result of the detection area.
[0130] In an exemplary embodiment, the image processing module 603 is further configured to determine a detection object and a detection behavior of the detection task according to the detection task through an image detection model; perform corresponding detection processing on the image to be detected according to the detection object to determine a detection area; and perform corresponding detection processing on the detection area according to the detection behavior to determine a detection result of the detection area.
[0131] In an exemplary embodiment, the image processing module 603 is further configured to perform corresponding detection processing on the detection area according to the detection behavior to determine a matching value of the detection behavior corresponding to the detection area; determine a detection result of the detection area according to a preset matching threshold and the matching value of the detection behavior; and the detection result of the detection area is used to indicate whether there is a detection behavior in the detection area.
[0132] In an exemplary embodiment, the image processing module 603 is further configured to, when the detection result of the detection area indicates that there is a detection behavior in the detection area, determine that the detection result of the image to be detected is that there is an abnormal behavior in the image to be detected; and when the detection result of the detection area indicates that there is no detection behavior in the detection area, determine that the detection result of the image to be detected is that there is no abnormal behavior in the image to be detected.
[0133] In an exemplary embodiment, the information input module 602 is further configured to perform semantic recognition processing on the detection task description information through a semantic recognition model to obtain semantic feature information of the detection task description information; and determine the detection task according to the semantic feature information.
[0134] In an exemplary embodiment, the apparatus 600 further includes: an information generation module, configured to, when the detection result of the image to be detected is that there is an abnormal behavior in the image to be detected, generate a corresponding warning information according to the detection result of the image to be detected; and perform corresponding warning processing on the detection object of the detection task according to the warning information.
[0135] Each module in the above image detection device can be implemented in whole or in part by software, hardware, or a combination thereof. Each of the above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or be stored in the memory of the computer device in software form so that the processor can call and execute the operations corresponding to each of the above modules.
[0136] In an exemplary embodiment, a computer device is provided. The computer device can be a terminal, and its internal structural diagram can be as Figure 7 shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner. The wireless manner can be implemented through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. The computer program, when executed by the processor, implements an image detection method. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0137] Those skilled in the art can understand that Figure 7 the structure shown in
[0138] is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different component layout.
[0139] In an exemplary embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0140] In an exemplary embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0141] Those of ordinary skill in the art can understand that all or part of the processes in the above method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tapes, floppy disks, flash memories, optical memories, high-density embedded non-volatile memories, resistive random access memories (ReRAMs), magnetoresistive random access memories (MRAMs), ferroelectric random access memories (FRAMs), phase change memories (PCMs), graphene memories, etc. Volatile memories can include random access memory (RAM) or external cache memories, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logics, data processing logics based on quantum computing, etc., without limitation.
[0142] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.
[0143] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. An image detection method, characterized in that, The method includes: Obtaining detection task description information for an image to be detected; Inputting the detection task description information into a semantic recognition model to obtain a detection task corresponding to the detection task description information; Performing corresponding detection processing on the image to be detected according to the detection task through an image detection model to obtain a detection result for the image to be detected.
2. The method according to claim 1, wherein The performing corresponding detection processing on the image to be detected according to the detection task through an image detection model to obtain a detection result for the image to be detected includes: Performing corresponding detection processing on the image to be detected according to the detection task through the image detection model to determine a detection area in the image to be detected and a detection result of the detection area; Determining the detection result of the image to be detected according to the detection area and the detection result of the detection area.
3. The method according to claim 2, wherein The performing corresponding detection processing on the image to be detected according to the detection task through the image detection model to determine a detection area in the image to be detected and a detection result of the detection area includes: Determining a detection object of the detection task and a detection behavior of the detection task according to the detection task through the image detection model; Performing corresponding detection processing on the image to be detected according to the detection object to determine the detection area; Performing corresponding detection processing on the detection area according to the detection behavior to determine the detection result of the detection area.
4. The method according to claim 3, wherein The performing corresponding detection processing on the detection area according to the detection behavior to determine the detection result of the detection area includes: Performing corresponding detection processing on the detection area according to the detection behavior to determine a matching value of the detection behavior corresponding to the detection area; Determining the detection result of the detection area according to a preset matching threshold and the matching value of the detection behavior; the detection result of the detection area is used to indicate whether the detection behavior exists in the detection area.
5. The method according to claim 4, wherein The determining the detection result of the image to be detected according to the detection area and the detection result of the detection area includes: When the detection result of the detection area indicates that the detection behavior exists in the detection area, determining that the detection result of the image to be detected is that there is an abnormal behavior in the image to be detected; When the detection result of the detection area indicates that the detection behavior does not exist in the detection area, determining that the detection result of the image to be detected is that there is no abnormal behavior in the image to be detected.
6. The method according to claim 1, wherein The inputting the detection task description information into a semantic recognition model to obtain a detection task corresponding to the detection task description information includes: Performing semantic recognition processing on the detection task description information through the semantic recognition model to obtain semantic feature information of the detection task description information; Determining the detection task according to the semantic feature information.
7. The method according to claim 1, wherein After performing corresponding detection processing on the image to be detected according to the detection task through an image detection model to obtain a detection result for the image to be detected, the method further includes: In the case that the detection result of the image to be detected indicates the existence of abnormal behavior in the image to be detected, generate a corresponding warning message according to the detection result of the image to be detected; Perform corresponding warning processing on the detection object of the detection task according to the warning message.
8. An image detection device, characterized in that, The device includes: An information acquisition module, configured to acquire detection task description information for an image to be detected; An information input module, configured to input the detection task description information into a semantic recognition model to obtain a detection task corresponding to the detection task description information; An image processing module, configured to perform corresponding detection processing on the image to be detected according to the detection task through an image detection model to obtain a detection result for the image to be detected.
9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.