Target detection processing method and device, electronic equipment, medium and program product
By adopting a deep learning detection model optimized by integrating attention mechanism and knowledge distillation strategy in self-service terminal equipment, the detection stability and response speed of self-service terminal equipment in complex environments is solved, real-time detection and security prompts of user privacy and illegal behavior are achieved, and the security and user experience of the equipment are improved.
Patent Information
- Application Number
- CN202510553257.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-08-01
AI Technical Summary
The existing self-service terminal equipment has poor facial recognition stability in complex environments, the three-dimensional structured optical hardware is costly and slow to respond, and lacks the ability to detect and alert users' privacy and illegal behaviors in real time. The existing detection technology lacks a state judgment mechanism that is linked to business processes.
A deep learning detection model optimized by combining attention mechanism and knowledge distillation strategy is adopted to process image data through the object detection model, abnormal state judgment is performed, and abnormal response is performed to build a closed-loop intelligent processing process.
It improves detection accuracy and reduces model calculation complexity, adapts to edge device deployment, realizes efficient detection and abnormal behavior judgment of various targets in self-service terminal scenarios, and improves privacy protection capabilities, security prevention and control capabilities and user interaction experience.
Smart Images

Figure CN120411541A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology and fintech technology, and more particularly to a method, apparatus, device, medium, and program product for target detection processing. Background Art
[0002] Currently, self-service terminal devices widely adopt face recognition technology based on two-dimensional images or three-dimensional structured light for identity comparison and user behavior judgment. However, two-dimensional image recognition is prone to recognition failure under conditions such as illumination changes, occlusion (such as masks, hats), or side faces, and it is difficult to operate stably in complex environments; while three-dimensional structured light, although having higher accuracy, has high hardware costs, complex integration, and is difficult to deploy in large-scale devices. Moreover, in high-concurrency scenarios, its response speed is slow and real-time performance is poor, affecting the user experience. In addition, traditional solutions mostly focus on identity recognition and ignore real-time detection and security prompts during the business handling process, posing a risk of user privacy exposure. Summary of the Invention
[0003] In view of the above problems, the present disclosure provides a method, apparatus, device, medium, and program product for target detection processing.
[0004] According to a first aspect of the present disclosure, there is provided a method for target detection processing, the method including: acquiring target image data in response to a screen trigger operation; processing the target image data based on a target detection model to obtain a target detection result, where the target detection model is a deep learning detection model optimized by fusing an attention mechanism and a knowledge distillation strategy; making an abnormal state judgment based on the target detection result; and performing abnormal response processing based on the result of the abnormal state judgment.
[0005] According to an embodiment of the present disclosure, the making an abnormal state judgment based on the target detection result specifically includes: obtaining target quantity information and target position information from the target detection result; making a first judgment based on the target quantity information to obtain a first state instruction; in response to the first state instruction being a target state instruction, making a second judgment based on the target position information to obtain a second state instruction; and making an abnormal state judgment based on the first state instruction and / or the second state instruction.
[0006] According to an embodiment of the present disclosure, the optimization process of the target detection model includes: selecting a target deep learning model as the teacher model, training the teacher model using a preset image dataset, and outputting soft labels including class probability distributions, where an attention mechanism module is embedded in the feature extraction network of the teacher model; performing knowledge transfer training on the student model based on the soft labels and temperature adjustment parameters; and performing model evaluation based on the accuracy and inference speed of the trained student model in the edge environment to obtain the target detection model.
[0007] According to an embodiment of the present disclosure, the performing knowledge transfer training on the student model based on the soft labels and temperature adjustment parameters specifically includes: in multiple training rounds, gradually reducing the value of the temperature adjustment parameter with a preset decay strategy; in each training round, calculating the distillation loss value based on the soft labels according to the temperature adjustment parameter corresponding to the current round; obtaining the detection loss value of the student model under the true labels; performing weighted summation on the distillation loss value and the detection loss value according to the dynamic weight set in the current training stage to obtain a combined loss value; and updating the parameters of the student model based on the combined loss value.
[0008] According to an embodiment of the present disclosure, the obtaining the target quantity information and target position information from the target detection result specifically includes: screening out target bounding boxes from the target detection result whose confidence exceeds a first preset threshold and belong to a predetermined category, and obtaining the target quantity information and target position information based on the quantity and position of the target bounding boxes.
[0009] According to an embodiment of the present disclosure, the target status instruction is represented as the status instruction when the target quantity information is two targets. The responding to the first status instruction being the target status instruction and making a second judgment based on the target position information to obtain a second status instruction specifically includes: calculating the areas of the target bounding boxes of the two targets, and sorting based on the areas of the target bounding boxes; and calculating the proportion of the target with the smaller area among the two targets in the area of the entire image region, and making a second judgment based on the proportion and a second preset threshold to obtain a second status instruction.
[0010] According to an embodiment of the present disclosure, the target detection result at least includes the detection result of a face image.
[0011] According to an embodiment of the present disclosure, the target image data includes consecutive image frames, and the method further includes: processing the target image data using deep learning-based multi-object tracking to obtain identity association information of the same target between different image frames; and performing temporal consistency verification on the target detection result based on the identity association information to obtain a verification result.
[0012] According to an embodiment of the present disclosure, the image data set includes training samples augmented via at least one image enhancement strategy for changing the spatial distribution characteristics and / or visual transformation characteristics of an image while maintaining the image semantics unchanged.
[0013] A second aspect of the present disclosure provides a processing device for object detection, the device including: a data acquisition module configured to: in response to a screen trigger operation, acquire target image data; an object detection module configured to: process the target image data based on an object detection model to obtain an object detection result, where the object detection model is a deep learning detection model optimized by fusing an attention mechanism and a knowledge distillation strategy; an abnormal state determination module configured to: perform an abnormal state determination based on the object detection result; and an abnormal processing module configured to: perform an abnormal response process based on the result of the abnormal state determination.
[0014] According to an embodiment of the present disclosure, the abnormal state determination module may further be configured to obtain target quantity information and target location information from the object detection result; perform a first determination based on the target quantity information to obtain a first state instruction; in response to the first state instruction being a target state instruction, perform a second determination based on the target location information to obtain a second state instruction; and perform an abnormal state determination based on the first state instruction and / or the second state instruction.
[0015] According to an embodiment of the present disclosure, the object detection module may further be configured to select a target deep learning model as a teacher model, train the teacher model using a preset image data set, and output a soft label including a class probability distribution, where an attention mechanism module is embedded in the feature extraction network of the teacher model; perform knowledge transfer training on a student model based on the soft label and a temperature adjustment parameter; and perform model evaluation based on the accuracy and inference speed of the trained student model in an edge environment to obtain the object detection model.
[0016] According to an embodiment of the present disclosure, the object detection module may further be configured to gradually reduce the value of the temperature adjustment parameter in multiple training rounds according to a preset attenuation strategy; in each training round, calculate a distillation loss value based on the soft label based on the temperature adjustment parameter corresponding to the current round; obtain a detection loss value of the student model under true labels; perform a weighted sum of the distillation loss value and the detection loss value according to a dynamic weight set in the current training stage to obtain a combined loss value; and update the parameters of the student model based on the combined loss value.
[0017] According to an embodiment of the present disclosure, the abnormal state judgment module can also be used to filter out target bounding boxes whose confidence exceeds a first preset threshold and belongs to a predetermined category from the target detection results, and obtain target quantity information and target position information based on the number and position of the target bounding boxes.
[0018] According to an embodiment of the present disclosure, the abnormal state judgment module can also be used to calculate the target bounding box areas of two targets and sort them based on the target bounding box areas; and calculate the proportion of the target with the smaller area of the two targets to the area of the entire image area, perform a second judgment based on the proportion and a second preset threshold, and obtain a second state instruction.
[0019] A third aspect of the present disclosure provides an electronic device, comprising: one or more processors; and a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the above method.
[0020] The fourth aspect of the present disclosure further provides a computer-readable storage medium having a computer program or instructions stored thereon, which implements the steps of the above method when the computer program or instructions are executed by a processor.
[0021] The fifth aspect of the present disclosure further provides a computer program product, comprising a computer program or instructions, which implement the steps of the above method when executed by a processor.
[0022] According to the embodiments of the present disclosure, by integrating an attention mechanism with a knowledge distillation strategy to optimize deep learning detection models, the computational complexity of the model is effectively reduced while ensuring detection accuracy. Furthermore, by constructing an intelligent processing flow with a closed loop of "image acquisition - detection and recognition - state judgment - response processing", it can adapt to edge device deployment requirements, thereby achieving efficient detection and abnormal behavior judgment of various targets in self-service terminal scenarios, and linking the system for real-time response, effectively improving the privacy protection capabilities, security control capabilities, and user interaction experience of terminal devices. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The above contents and other objects, features and advantages of the present disclosure will become more apparent through the following description of the embodiments of the present disclosure with reference to the accompanying drawings, in which:
[0024] Figure 1A Schematically illustrates an application scenario diagram of a method, apparatus, device, medium, and program product for processing target detection according to an embodiment of the present disclosure;
[0025] Figure 1B The structure of the target detection processing system according to the embodiment of the present disclosure is schematically shown;
[0026] Figure 2 Schematically shows a flowchart of a processing method for object detection according to an embodiment of the present disclosure;
[0027] Figure 3 Schematically shows a flowchart of a method for exception handling according to an embodiment of the present disclosure;
[0028] Figure 4 Schematically shows a flowchart of a model optimization process according to an embodiment of the present disclosure;
[0029] Figure 5 Schematically shows a flowchart of a knowledge distillation process according to an embodiment of the present disclosure;
[0030] Figure 6 Schematically shows an execution flowchart of a processing method for object detection based on an automated teller machine according to an embodiment of the present disclosure;
[0031] Figure 7 Schematically shows a structural block diagram of a processing apparatus for object detection according to an embodiment of the present disclosure; and
[0032] Figure 8 Schematically shows a block diagram of an electronic device suitable for implementing the processing method for object detection according to an embodiment of the present disclosure. Detailed implementation manners
[0033] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the following detailed description, for the sake of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, obviously, one or more embodiments can be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily obscuring the concepts of the present disclosure.
[0034] The terms used herein are merely for describing specific embodiments and are not intended to limit the present disclosure. The terms "including", "comprising", etc. used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0035] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0036] In the case of using expressions such as "at least one of A, B, and C", generally, it should be interpreted according to the meaning that those skilled in the art usually understand this expression (for example, "a system having at least one of A, B, and C" should include, but not be limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).
[0037] First, the technical terms described in this article are explained and illustrated as follows.
[0038] YOLOv5: An efficient single-stage object detection algorithm that uses Darknet-53 as the backbone network and achieves feature fusion through a multi-scale prediction network (PANet), enabling the simultaneous prediction of bounding boxes and confidence levels in a single forward pass.
[0039] Darknet-53: The backbone network of YOLOv5, used to extract deep features of images.
[0040] PANet: The multi-scale prediction network in YOLOv5, used for feature fusion to improve detection accuracy and speed.
[0041] Knowledge distillation: A model optimization technique that transfers the knowledge of a large model to a small model to improve the efficiency and performance of the model.
[0042] Attention mechanism (CBAM module): A technique used to enhance the model's ability to perceive key regions, improving the accuracy of detection.
[0043] With the widespread deployment of self-service terminal devices in fields such as finance, government affairs, and security, more and more business handling activities rely on users to complete autonomous operations in front of the terminal. To ensure user privacy and business security, self-service terminal devices are usually equipped with an image acquisition function and identify user identities or behaviors through object detection technology.
[0044] Currently, in such scenarios, identification means based on image processing are generally adopted, mainly including: a detection method based on two-dimensional images, which collects images through a camera, extracts the bounding box features of targets such as faces and objects, and conducts comparison or analysis; and an identification method based on three-dimensional structured light, which projects gratings or dot matrices to achieve three-dimensional modeling of the target for obtaining higher-precision structural information.
[0045] However, there are still many problems in the practical application of these existing technologies. The two-dimensional image detection technology has poor stability in complex environments and is easily affected by light changes and occlusions (such as covers and hand movements), resulting in misjudgments or missed detections. Especially in practical operation scenarios with large crowds and variable angles, it is difficult to ensure detection consistency. Although three-dimensional structured light can provide high recognition accuracy, its hardware cost is high and the system integration is complex, making it not suitable for wide deployment in self-service devices with cost sensitivity and space constraints. In addition, most existing systems focus on face recognition itself and lack the ability to detect and judge surrounding risk targets (such as users taking pictures with mobile phones during the process, others approaching the device, and the appearance of sensitive objects such as cameras), making it difficult to meet the higher requirements for privacy security and prevention and control of illegal behaviors.
[0046] More critically, most existing detection technologies focus on the processing of single-frame static images and lack a state judgment mechanism and system-level response strategy that are linked to the business process. For example, when multiple suspicious targets are detected in front of the terminal or abnormal behaviors occur for a long time, the terminal device cannot be timely linked to take countermeasures such as brightness adjustment, prompt warnings, and image capture, thus affecting the overall user experience and system security.
[0047] Based on this, the embodiments of the present disclosure provide a processing method for target detection. The method includes: in response to a screen trigger operation, obtaining target image data; processing the target image data based on a target detection model to obtain a target detection result, where the target detection model is a deep learning detection model optimized by fusing an attention mechanism and a knowledge distillation strategy; making an abnormal state judgment based on the target detection result; and performing processing of an abnormal response based on the result of the abnormal state judgment. The processing method for target detection provided by the present disclosure optimizes the deep learning detection model by fusing the attention mechanism and the knowledge distillation strategy. While improving the detection accuracy and lightweight performance, it constructs an intelligent processing flow with a closed loop of "image acquisition - detection and recognition - state judgment - response processing", which can adapt to the deployment requirements of edge devices, thereby realizing efficient detection of various targets and judgment of abnormal behaviors in the self-service terminal scenario, and linking the system for real-time response, effectively improving the privacy protection ability, security prevention and control ability, and user interaction experience of the terminal device.
[0048] It should be noted that the processing method, device, equipment, medium, and program product for target detection determined by the present disclosure can be used in the fields of artificial intelligence technology and fintech technology, and can also be used in a variety of fields other than the fields of artificial intelligence technology and fintech technology. The application fields of the processing method, device, equipment, medium, and program product provided by the embodiments of the present disclosure are not limited.
[0049] In the technical solutions of the present disclosure, the user information involved (including but not limited to user personal information, user image information, user device information such as location information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) are all information and data authorized by the user or fully authorized by all parties. Moreover, the processing of relevant data, such as collection, storage, use, processing, transmission, provision, disclosure, and application, all complies with relevant laws, regulations, and standards, takes necessary confidentiality measures, does not violate public order and good customs, and provides corresponding operation entrances for users to choose to authorize or refuse.
[0050] In the scenario of making automated decisions using personal information, the methods, devices, and systems provided by the embodiments of the present disclosure all provide corresponding operation entrances for users to choose to agree or refuse the results of automated decisions; if the user chooses to refuse, the expert decision-making process will be entered. The expression "automated decision" here refers to the activity of automatically analyzing and evaluating an individual's behavior habits, hobbies, or economic, health, credit status, etc. through a computer program and making decisions. The expression "expert decision" here refers to the activity of making decisions by personnel who are engaged in work in a specific field, have specialized experience, knowledge, and skills, and have reached a certain professional level.
[0051] Figure 1A Schematically shows an application scenario diagram of a method, device, equipment, medium, and program product for object detection according to an embodiment of the present disclosure.
[0052] As Figure 1A shown, the application scenario 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0053] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only for examples).
[0054] The first terminal device 101, the second terminal device 102, and the third terminal device 103 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, and the like.
[0055] The server 105 may be a server providing various services, such as a background management server (only for example) that supports the websites browsed by the user using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background management server may analyze and process data such as user requests received, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device.
[0056] It should be noted that the processing method for object detection provided by the embodiments of the present disclosure can generally be executed by the server 105. Correspondingly, the processing device for object detection provided by the embodiments of the present disclosure can generally be set in the server 105. The processing method for object detection provided by the embodiments of the present disclosure can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105. Correspondingly, the processing device for object detection provided by the embodiments of the present disclosure can also be set in a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105.
[0057] It should be understood that Figure 1A the numbers of terminal devices, networks, and servers in
[0058] Figure 1B are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers.
[0059] As Figure 1B shown, the object detection processing system provided by the embodiments of the present disclosure may include a data acquisition module, an edge computing module, a detection reminder module, and a background storage module. The modules cooperate with each other to form a closed-loop system for multi-object detection and abnormal reminder applicable to various scenarios.
[0060] Referring to Figure 1B , the data acquisition module may include a front camera, which is configured at a prominent position of the target device and is used to collect a video image stream in front of the device in real time during the user operation process, and transmit the collected video data to the edge computing module for processing.
[0061] The edge computing module can be deployed with an optimized deep learning object detection model, which can be optimized by integrating the attention mechanism and knowledge distillation strategy. It can effectively reduce the model's computational complexity while ensuring detection accuracy, and meet the deployment requirements of edge devices. Specifically, the edge computing module can analyze each frame of the image, identify the objects in the image, calculate their quantity and the proportion relative to the screen area, generate an abnormal state determination result according to the preset judgment logic, and submit the result to the detection reminder module.
[0062] The detection reminder module can control the target device to perform corresponding response operations according to the type of abnormality, specifically including screen prompts, voice warnings, and video screenshot functions, so as to timely remind users to pay attention to privacy and operation safety. For example, when applied to the scenario of bank self-service teller machines, the system can identify the situation of excessive non-user faces or abnormal closeness to the screen, preventing the leakage or peeping of user information; another example is that when applied to government self-service handling terminals, it can remind of abnormal behaviors such as sensitive items like cameras and mobile phones approaching the screen, preventing the risk of illegal shooting or information theft. When it is determined to be a serious abnormality, the system will also call the screenshot function to capture the current video screen and transmit the screenshot content to the background storage module.
[0063] The background storage module is used to save the video screenshot information generated in abnormal environments, providing data support for subsequent risk analysis, system optimization, and security event traceability.
[0064] The following will be based on Figure 1A and Figure 1B the described scenarios, and will describe in detail the processing method of object detection in public embodiments through Figures 2 to 6 ...
[0065] Figure 2 FIG. schematically shows a flowchart of a processing method for object detection according to an embodiment of the present disclosure.
[0066] As Figure 2 shown, the processing method for object detection in this embodiment includes operations S210 to S240, and this object detection method can be executed by server 105.
[0067] In operation S210, in response to a screen trigger operation, target image data is acquired.
[0068] Specifically, the screen trigger operation may include, but is not limited to, user interaction behaviors such as touch operation, click operation, swipe operation, or screen lighting event. When such an operation is detected, the system considers that the user is in an active interaction state and triggers the image acquisition module to obtain target image data from an image acquisition device (such as a camera). Compared with the continuous acquisition or timed acquisition method, the response based on the screen trigger operation can more accurately obtain the image data at the moment of the user's active interaction, thereby increasing the probability of the appearance of a face in the acquired image and avoiding invalid images such as empty frames and offline frames.
[0069] In an embodiment of the present disclosure, the target image data can be a single-frame static image or a continuous video stream image, depending on the service scenario. For example, in the scenario of a self-service teller machine, when the user clicks the "handle business" button, the camera is immediately activated to capture the image of the operation area in front of the device in real time for subsequent detection and analysis.
[0070] In operation S220, the target image data is processed based on the target detection model to obtain a target detection result, where the target detection model is a deep learning detection model optimized by fusing an attention mechanism and a knowledge distillation strategy.
[0071] In an embodiment of the present disclosure, the target detection model is based on a deep convolutional neural network. By introducing an attention mechanism, the model's perception ability of key region features is enhanced. Especially in the face detection scenario under complex backgrounds or low-light environments, it can effectively highlight the significant regions related to the face and suppress the interference of irrelevant backgrounds, thereby improving the target detection accuracy. At the same time, the target detection model is also optimized by combining the knowledge distillation strategy. This strategy includes training a teacher model with high accuracy but large computational overhead and using the "soft labels" in its prediction output as a knowledge source to guide the training of another lightweight student model, thereby improving the accuracy upper limit of the model and breaking through the performance bottleneck of lightweight models.
[0072] According to an embodiment of the present disclosure, the target detection model can identify various types of targets including faces, mobile phones, cameras, two-dimensional codes, etc., and extract their position, bounding box size, and category information in the image. For example, when the target detection model is deployed in a self-service government affairs handling terminal, it can detect whether a user attempts to use a mobile phone to photograph sensitive documents or information.
[0073] In operation S230, an abnormal state judgment is made based on the target detection result. For example, the abnormal state can be identified according to factors such as the number of targets, categories, and positions.
[0074] Exemplarily, when multiple faces that do not belong to the current operator appear in front of the screen, or when a camera or reflective object approaches the screen area, the system will determine it as a "crowding risk" or "sneak shot risk"; for another example, when the detection result shows that the same user stays in front of the device for a long time without further operation, the system can also identify it as an "abnormal stay" state. This abnormal judgment mechanism supports flexible configuration and can adapt to the requirements of different business types and risk levels.
[0075] In operation S240, perform abnormal response processing based on the result of the abnormal state judgment. For example, the abnormal processing can include screen brightness adjustment, voice prompt playback, video screenshot upload, temporary service locking, etc.
[0076] Exemplarily, when an illegal item approaches or there is a crowd watching, the system can automatically reduce the screen brightness and play a voice prompt to remind the user to pay attention to privacy security; in case of a high-risk situation, the system can also immediately turn off the screen and capture the current screen, and upload it to the background server for recording and analysis. In addition, the system can also associate and store the response record with the user operation process for post-event behavior review, dispute evidence collection, or system security policy optimization.
[0077] It should be noted that the target detection processing method of the embodiments of the present disclosure is applicable to various device scenarios that require image interaction and behavior monitoring, such as bank self-service teller machines, government service terminals, unattended vending systems, intelligent security terminals, etc. That is, the present disclosure does not limit the actual application scenarios. The present disclosure automatically enters the image acquisition and intelligent recognition process by responding to the user's operation trigger, and combines the optimized deep learning model to realize the recognition and dynamic response control of target behaviors or target objects, significantly enhancing the device's abnormal behavior recognition and dynamic response capabilities.
[0078] Figure 3 A flowchart of the method for abnormal processing according to an embodiment of the present disclosure is schematically shown.
[0079] As Figure 3 shown, the method for abnormal processing of this embodiment may include operation S310 to operation S340.
[0080] In operation S310, obtain the target quantity information and target position information from the target detection result. The target quantity information may include the total number of target instances recognized in the image, such as the number of recognized faces, the number of cameras, or the detection target quantity of specific categories such as mobile phones. The target position information may include the position of each target in the image, the bounding box area, the center point coordinates, etc.
[0081] In an embodiment of the present disclosure, target bounding boxes with a confidence level exceeding a first preset threshold and belonging to a predetermined category can be screened out from the target detection results, and target quantity information and target location information can be obtained based on the quantity and location of the target bounding boxes.
[0082] The predetermined category can be configured according to the specific application scenario. For example, in the scenario of an automated teller machine, it can be set to the "face" category to monitor users and potential onlookers; in a government affairs terminal, it can also be extended to include sensitive item categories such as "camera" and "mobile phone" for identifying illegal shooting behaviors; in scenarios such as access control, it can also be set to categories such as "hand", "identity document", and "moving object" to meet different security or interaction requirements.
[0083] After obtaining the bounding boxes that meet the conditions, the spatial attributes of each target can be further analyzed, including the position of the target bounding box (such as the center coordinates, upper, lower, left, and right boundaries), size (area), and relative layout relationship in the image. This information can be used for subsequent operations such as quantity judgment, ratio calculation, and trajectory tracking. For example, by performing clustering analysis on the center points of the bounding boxes, the system can identify whether there are multi-target aggregation areas; by comparing the areas, the distance relationship and interference degree between targets can be judged; by the change trend of the target positions in consecutive frames, it can also be evaluated whether there are suspicious behaviors such as a target continuously approaching the screen or staying for a long time.
[0084] In operation S320, a first judgment is made based on the target quantity information to obtain a first status instruction. The first judgment can be used to identify whether there is an abnormal target quantity exceeding normal usage behaviors, and is applicable to potential interference scenarios such as multiple people gathering and approaching abnormally.
[0085] Specifically, the first judgment can be made by comparing with a preset value to determine whether there are more suspicious targets than the reasonable quantity in the current image. For example, in a face recognition application, when the number of targets detected by the system is one, it is default considered to be in a normal business handling state, the system maintains the normal brightness of the screen, allows the user to continue the operation, and the first status instruction generated can be a normal state without triggering any abnormal prompt.
[0086] When the detection result shows that there are two targets in the picture, the system will enter S330 to further confirm the abnormal state. At this time, the first status instruction is set to a "target status instruction", for example, it can be a pending instruction. When three or more targets are detected, the system can directly consider it to be in a serious abnormal state, and the first status instruction generated can directly correspond to the abnormal situation and be used to trigger a stronger-level system response (such as screen off, voice warning, screenshot upload, etc.).
[0087] According to an embodiment of the present disclosure, the preliminary screening mechanism based on the target quantity information can quickly predict the usage environment on the premise of ensuring the system response efficiency, and reduce the possibility of misjudgment and missed judgment.
[0088] In operation S330, in response to the first status instruction being a target status instruction, a second judgment is made based on the target position information to obtain a second status instruction. The second judgment is used to further evaluate the actual interference risk of the second target in the case where two targets have been detected in the first judgment, avoid misjudging as abnormal due to accidental occlusion or background noise, and improve the accuracy and practicability of the system judgment.
[0089] In an embodiment of the present disclosure, the boundary box information of the two detected targets can be extracted, and the area occupied by each target in the image can be calculated. Subsequently, the two targets are sorted by area, and the target with the smaller area is identified. The system calculates the ratio of the area of the smaller target to the area of the entire image region to obtain its screen occupancy ratio in the image. If this ratio exceeds a preset threshold, such as 10%, it can be determined that the second target is not only close to the primary user but may also be in a position to view the screen, interfere with operations, or obtain sensitive information. Thus, the system generates a second status instruction indicating the existence of potential abnormal behavior. On the contrary, if this ratio is lower than the threshold, it means that the second target is far away or only accidentally enters the picture, and the system does not consider it to pose a serious risk and can maintain the normal business process.
[0090] According to an embodiment of the present disclosure, the second judgment mechanism can effectively filter out misjudgments caused by non-interference factors such as misidentification and brief approach, and further refine the system's perception ability of abnormal states. For example, during the use of a bank or government self-service teller machine, if a companion who stays briefly next to the primary user accidentally enters the picture without creating an actual risk, this mechanism can avoid unnecessary system interruptions and improve the user experience; while when a second person continuously approaches and clearly views or records the screen content, this mechanism can accurately identify and respond. This fine-grained risk assessment strategy based on screen occupancy provides higher sensitivity and robustness for the system in actual deployment.
[0091] In operation S340, an abnormal state judgment is made based on the first status instruction and / or the second status instruction. That is to say, operation S340 can serve as the decision-making link in the abnormal handling process. Based on the hierarchical information of the previous judgment, the detected scenario behavior is classified into normal, suspicious, or seriously abnormal states, and then the terminal device is driven to make different responses.
[0092] Specifically, when the first status instruction indicates a severe abnormal status (for example, the detected number of targets is three or more), the system can directly determine that the current is a high-risk scenario without relying on the second judgment result, and judge the final abnormal status as "severe abnormal". In this case, the system can trigger a series of response operations, such as immediately turning off the display screen, playing a voice prompt, terminating the current business process, and taking a screenshot or short video recording of the current screen and uploading it to the background system for subsequent security review or data analysis.
[0093] When the first status instruction is a target status instruction (that is, there are two targets), the system will further refine the risk level in combination with the judgment result of the second status instruction. If the second status instruction indicates that the screen occupation ratio exceeds the threshold, it means that the second target is already in a risk state of interference or peeping. The system can judge the abnormal status as "moderate abnormal" or "potential abnormal", and execute relatively mild response measures such as reducing the screen brightness, displaying a reminder message on the interface, and playing a mild prompt tone; conversely, if the second target occupation ratio is extremely low and does not constitute actual interference, the system can maintain the normal business process or only perform background recording without front-end prompting to avoid unnecessary interference with user operations.
[0094] Figure 4 A flowchart of the model optimization process according to an embodiment of the present disclosure is schematically shown.
[0095] As Figure 4 shown, the method for model optimization in this embodiment may include operation S410 to operation S430.
[0096] In operation S410, a target deep learning model is selected as the teacher model, and the teacher model is trained using a preset image data set to output a soft label including a class probability distribution, wherein an attention mechanism module is embedded in the feature extraction network of the teacher model.
[0097] In an embodiment of the present disclosure, the image data set can be selected according to the target application scenario, such as a positive and negative sample set for face detection, a multi-class image set for object recognition, etc. To further improve the feature extraction ability of the teacher model, an attention mechanism module is embedded in its feature extraction network structure, such as a channel attention mechanism (SE), a spatial attention mechanism (SAM), or a combined channel and spatial mechanism (CBAM). This module can enable the teacher model to more effectively focus on key regions, thereby improving the recognition accuracy of small targets, occluded targets, and targets under complex backgrounds. After training, the teacher model can not only output conventional detection results such as bounding boxes and confidence levels, but also output the probability distribution information corresponding to each class of targets, that is, the so-called "soft label", for guiding the student model to perform knowledge transfer.
[0098] In operation S420, knowledge transfer training is performed on the student model based on the soft labels and temperature adjustment parameters.
[0099] In an embodiment of the present disclosure, the student model can use the soft labels generated by the teacher model as a supervision signal and learn by fitting with the class distribution output by the teacher, so as to reproduce the target feature discrimination ability learned by the teacher model as much as possible while maintaining the simplicity of the model. During this process, a temperature adjustment parameter is introduced to control the smoothness of the softmax function, making the probability distribution output by the teacher model easier to be learned by the student model. The temperature parameter can adopt a preset attenuation strategy during the training process, gradually decreasing from a higher value, so that the student model can obtain richer class information in the initial stage of training and focus on the main classes in the later stage. In addition, to improve the training quality, the system can also introduce a joint loss function, fuse the distillation loss and the detection loss based on the true labels with weighted fusion, and cooperate with the dynamic loss weight control strategy to make the training process more stable.
[0100] In operation S430, the model is evaluated based on the accuracy and inference speed of the trained student model in the edge environment to obtain the target detection model.
[0101] The evaluation process can include accuracy metrics (such as mAP, Precision, Recall) and efficiency metrics (such as inference time, FPS, number of model parameters, etc.), and focus on the actual performance of the model in the edge computing environment. For example, when the student model is deployed in the front-end device of a self-service terminal, it must meet the requirements of low latency, high accuracy, and low power consumption to achieve real-time processing and high-frequency response. Based on the evaluation results, the system can further perform fine-tuning or pruning operations on the student model structure to obtain the final optimized target detection model that can be deployed.
[0102] Hereinafter, an optimization method of a target detection model based on YOLOv5 will be used as an example to specifically illustrate the embodiments of the present disclosure.
[0103] It should be noted that as a stage detection framework, YOLOv5 has the characteristics of fast end-to-end processing speed, flexible model structure, and easy pruning and compression, and is particularly suitable for edge scenarios with high requirements for real-time performance and resource sensitivity.
[0104] Specifically, a version with stronger performance in the YOLOv5 family (such as YOLOv5l or YOLOv5x) can be selected as the basic structure of the teacher model. In this basic model, YOLOv5 usually uses Darknet-53 as the backbone network to extract the deep semantic features of images, which has the characteristics of rich hierarchy and strong expression ability. This backbone network is composed of multiple convolutional layers and residual structures, and can effectively capture the key target information in the image.
[0105] After feature extraction, YOLOv5 also integrates PANet (Path Aggregation Network) as its multi-scale prediction network to achieve efficient information fusion between feature maps at different levels. PANet enhances the expressive power of low-level features through lateral connections and the downward path, thus improving the detection performance of small objects and taking into account both detection accuracy and speed. In the feature extraction network of this basic model, to further enhance the discriminative ability of the model, a CBAM attention mechanism module is introduced after each backbone convolutional block (such as the C3 module). CBAM includes two branches: channel attention and spatial attention, which can guide the model to more effectively focus on key regions and improve the ability to capture face features in complex backgrounds.
[0106] Subsequently, the prepared face dataset is used to train this teacher model. During the training process, the model output not only includes conventional object detection results such as bounding boxes and confidence levels, but also generates the probability distribution of each target category, forming a "soft label". This soft label is calculated through the softmax function, and a temperature adjustment parameter T is introduced in its calculation process to control the smoothness of the category probability distribution. Among them, the softmax function is defined as follows:
[0107] (1)
[0108] where q i represents the normalized probability of the i-th category, and z i represents the value of the i-th category in the input vector. T represents the temperature coefficient, which gradually decreases with the training progress and is used to adjust the smoothness of the output probability of the teacher model. A high temperature in the early stage makes the probability distribution smoother, and a low temperature in the later stage strengthens the learning of the main categories. For example, the initial T = 20, the final T = 5, and a total of 100 rounds of training are performed, with T decreasing by 0.15 in each round. The formula is as follows:
[0109] (2)
[0110] After the training of the teacher model is completed, a lightweight student model structure is constructed. YOLOv5s or a pruned custom structure can be selected, and knowledge transfer training is carried out with reference to the soft label generated by the teacher model. The student model calculates its own prediction probability through the softmax function at the same temperature T and calculates the distillation loss based on the KL divergence. Its expression is as follows:
[0111] (3)
[0112] Since calculating and is scaled by 1 / T times, calculating is scaled times, so the final result needs to be multiplied by , offset the scaling to get dis_loss, ensuring the stability and effectiveness of the training process.
[0113] When using soft labels to train and calculate the distillation loss to obtain the distillation loss value dis_loss (the gap with the teacher model), and using the original data to train to obtain the detection loss value task_loss, assign weights to the two (controlling the intensity of distillation knowledge transfer, gradually decreasing as the training progresses, increasing the attention to the true labels). The specific formula is as follows:
[0114] (4)
[0115] Figure 5 Schematically shows a flowchart of the knowledge distillation process according to an embodiment of the present disclosure.
[0116] As Figure 5 shown, the process of knowledge distillation includes a dual-path structure of a teacher model and a student model, where the teacher model is a large model with a deeper structure and stronger performance, and the student model is a lightweight model with fewer parameters and higher computational efficiency. Both start with the same input samples during the training phase and perform forward propagation separately.
[0117] In the teacher model path, the input data is first processed through multiple neural network layers to obtain a set of output vectors. Subsequently, the output is normalized through the Softmax function with a temperature adjustment coefficient to generate a smoothed class probability distribution, that is, the so-called soft label. A higher temperature value can expand the probability difference between classes, enabling the teacher model to not only output the class with the highest confidence but also retain the information of sub-optimal classes, thus transmitting richer "dark knowledge" to the student model.
[0118] The student model uses the same input as the teacher model and also outputs a set of prediction vectors. A part of it is processed through the same Softmax temperature as the teacher model to obtain soft predictions, which are used to compare with the soft labels of the teacher model to calculate the distillation loss. On the other hand, the student model also calculates the hard prediction result at the standard temperature (i.e., T = 1) and compares it with the true label to calculate the supervision loss of the student model itself.
[0119] Finally, the system weights and fuses the distillation loss and the student loss to form a joint loss function for backpropagation and parameter update, so that the student model can not only fit the true label but also capture the implicit features of the teacher model in class discrimination.
[0120] After the training is completed, a pre-deployment evaluation of the student model can be carried out. The mean average precision (mAP) and the frames per second (FPS) of inference can be used as performance metrics to ensure both detection accuracy and real-time processing capabilities.
[0121] It should be noted that the YOLOv5 model, as a deep learning detection model, is only an example. Those skilled in the art should be aware that other deep learning detection models that can achieve similar functions can also be applied to the embodiments of the present disclosure. The present disclosure does not limit the specific object detection algorithm. For example, the object detection model can select open-source object detection algorithms such as SSD (Single Shot MultiBox Detector) or Faster R-CNN (Region-based Convolutional Neural Network), and fine-tune and train them based on task-specific scenarios.
[0122] The following will take the scenario of an automated teller machine as an example to specifically illustrate the embodiments of the present disclosure.
[0123] In this embodiment, the object detection result at least includes the detection result of the face image.
[0124] Figure 6 The execution flowchart of the processing method for object detection based on an automated teller machine according to an embodiment of the present disclosure is schematically shown.
[0125] As Figure 6 shown, when the user clicks the screen to start the business process, the system first starts the front camera and continuously performs face detection throughout the business handling process. During the detection process, the system uses the object detection model deployed on the edge computing device (such as the YOLOv5 model optimized by integrating the attention mechanism and the knowledge distillation strategy) to analyze the number and position information of the faces in the image in real time.
[0126] The system first determines whether the number of faces detected currently is one. When only one face is detected in the result, the system determines it as the normal state, maintains the normal brightness of the screen, and allows the user to continue with the business process. When the detection result shows that there are two objects in the picture, the system further determines whether the screen ratio occupied by the second face in the image exceeds a set threshold (such as 10%). If the ratio does not exceed the threshold, the system still considers the current scenario to be within the acceptable range and continues with the business process; if the ratio exceeds the threshold, it indicates that the second face may be in a state of approaching or peeping at the screen, and the system then triggers a moderate anomaly reminder, such as dimming the screen and displaying a prompt message to guide the user to pay attention to the safety of the surrounding environment.
[0127] If the detection result shows that there are three or more people, the system will directly determine it as a severe abnormal state, which may involve behaviors such as group gathering, information leakage, or illegal assistance. The system will immediately execute a forced response, including turning off the screen, displaying a warning message, capturing the current screen, and uploading it to the background for evidence preservation processing.
[0128] Under normal or moderately abnormal conditions, the service can continue. The system continuously monitors and repeats the above judgment process until the user clicks the "End Processing" button. At this time, the system automatically turns off the camera to complete a complete service monitoring process.
[0129] In some embodiments, the target image data can be continuous image frames. For such data, to further improve the stability and accuracy of the detection results, the system can further process the target image data based on the deep learning-based multi-object tracking technology. Specifically, the trained multi-object tracking model can be used to process the image frames to obtain the identity association information between different image frames of the same target. The identity association information can be used to identify the image regions belonging to the same object or the same subject in adjacent frames, so as to achieve cross-frame tracking and matching.
[0130] Furthermore, the system can verify the temporal consistency of the detection results based on the above identity association information. Temporal consistency means that the detection results of the same target in consecutive frames should have a stable evolution trend in terms of spatial position, scale, category, etc. For this purpose, the system can perform consistency verification on the target detection boxes in consecutive frames, such as calculating indicators such as intersection over union, category matching degree, and confidence change range, to determine whether the detection of the target in the image sequence is stable. When the consistency index meets the preset rules, the detection result is determined to be valid; otherwise, there may be false detection or missed detection situations. The above verification results can be used as the basis for evaluating the credibility of the final detection output, and can also be used for prior optimization of subsequent target behavior analysis, image summary generation, and other tasks.
[0131] In other embodiments, to improve the generalization ability and robustness of the image detection model, the image dataset can also include training samples expanded by at least one image enhancement strategy. The image enhancement strategy can include methods such as rotation, cropping, flipping, scaling, brightness adjustment, color perturbation, blurring, and noise addition. The above enhancement methods are used to change the spatial distribution characteristics and / or visual transformation characteristics of the image while keeping the semantic information of the image unchanged, so as to enhance the model's adaptability to different shooting conditions, scene changes, or interference factors.
[0132] For example, for the human body detection task in the security monitoring scenario, the original image can be horizontally flipped and its brightness perturbed to simulate scenarios under different directions and lighting conditions; for the industrial detection task, image offsets and contrast changes that may occur during production can be simulated through affine transformation and edge enhancement. By introducing the above-mentioned enhanced samples to participate in model training, the robustness of the model under diverse image inputs can be effectively improved, and the false detection and missed detection situations in complex scenarios can be reduced.
[0133] Figure 7 A structural block diagram of a processing device for object detection according to an embodiment of the present disclosure is schematically shown.
[0134] As Figure 7 shown, the processing device 700 for object detection in this embodiment includes a data acquisition module 710, an object detection module 720, an abnormal state judgment module 730, and an abnormal processing module 740.
[0135] The data acquisition module 710 can be used to acquire target image data in response to a screen trigger operation. In one embodiment, the data acquisition module 710 can be used to perform the operation S210 described above, which will not be elaborated here.
[0136] The object detection module 720 can be used to process the target image data based on an object detection model to obtain an object detection result, where the object detection model is a deep learning detection model optimized by fusing an attention mechanism and a knowledge distillation strategy. In one embodiment, the object detection module 720 can be used to perform the operation S220 described above, which will not be elaborated here.
[0137] The abnormal state judgment module 730 can be used to judge the abnormal state based on the object detection result. In one embodiment, the abnormal state judgment module 730 can be used to perform the operation S230 described above, which will not be elaborated here.
[0138] The abnormal processing module 740 can be used to process the abnormal response based on the result of the abnormal state judgment. In one embodiment, the abnormal processing module 740 can be used to perform the operation S240 described above, which will not be elaborated here.
[0139] According to an embodiment of the present disclosure, the abnormal state judgment module 730 can also be used to obtain target quantity information and target position information from the object detection result; make a first judgment based on the target quantity information to obtain a first state instruction; in response to the first state instruction being a target state instruction, make a second judgment based on the target position information to obtain a second state instruction; and make an abnormal state judgment based on the first state instruction and / or the second state instruction.
[0140] According to an embodiment of the present disclosure, the target detection module 720 may also be used to select a target deep learning model as a teacher model, train the teacher model using a preset image data set, and output soft labels including class probability distributions, wherein an attention mechanism module is embedded in the feature extraction network of the teacher model; perform knowledge transfer training on the student model based on the soft labels and temperature adjustment parameters; and perform model evaluation based on the accuracy and inference speed of the trained student model in the edge environment to obtain the target detection model.
[0141] According to an embodiment of the present disclosure, the target detection module 720 may also be used to gradually reduce the value of the temperature adjustment parameter in multiple training rounds according to a preset attenuation strategy; in each training round, calculate the distillation loss value based on the soft labels according to the temperature adjustment parameter corresponding to the current round; obtain the detection loss value of the student model under the true labels; perform weighted summation on the distillation loss value and the detection loss value according to the dynamic weight set in the current training stage to obtain a combined loss value; and update the parameters of the student model based on the combined loss value.
[0142] According to an embodiment of the present disclosure, the abnormal state judgment module 730 may also be used to screen out target bounding boxes with a confidence level exceeding a first preset threshold and belonging to a predetermined category from the target detection results, and obtain target quantity information and target position information based on the quantity and position of the target bounding boxes.
[0143] According to an embodiment of the present disclosure, the abnormal state judgment module 730 may also be used to calculate the areas of the target bounding boxes of two targets, sort based on the areas of the target bounding boxes; and calculate the proportion of the target with the smaller area among the two targets in the area of the entire image region, and perform a second judgment based on the proportion and a second preset threshold to obtain a second state instruction.
[0144] According to an embodiment of the present disclosure, any multiple of the data acquisition module 710, the target detection module 720, the abnormal state determination module 730, and the abnormal processing module 740 may be combined and implemented in one module, or any one of them may be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules may be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present disclosure, at least one of the data acquisition module 710, the target detection module 720, the abnormal state determination module 730, and the abnormal processing module 740 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or any other reasonable manner of integrating or packaging circuits, etc., implemented by hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in any suitable combination of several of them. Alternatively, at least one of the data acquisition module 710, the target detection module 720, the abnormal state determination module 730, and the abnormal processing module 740 may be at least partially implemented as a computer program module, and when the computer program module is run, corresponding functions may be executed.
[0145] Figure 8 Schematically shows a block diagram of an electronic device suitable for implementing a processing method for target detection according to an embodiment of the present disclosure.
[0146] As Figure 8 shown, the electronic device 800 according to an embodiment of the present disclosure includes a processor 801, which may perform various appropriate actions and processes according to a program stored in a read only memory (ROM) 802 or a program loaded from a storage section 808 into a random access memory (RAM) 803. The processor 801 may include, for example, a general microprocessor (such as a CPU), an instruction set processor and / or a related chipset and / or a dedicated microprocessor (such as an application specific integrated circuit (ASIC)), etc. The processor 801 may also include on-board memory for caching purposes. The processor 801 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.
[0147] In the RAM 803, various programs and data required for the operation of the electronic device 800 are stored. The processor 801, the ROM 802, and the RAM 803 are connected to each other via the bus 804. The processor 801 performs various operations of the method flow according to the embodiments of the present disclosure by executing the programs in the ROM 802 and / or the RAM 803. It should be noted that the programs may also be stored in one or more memories other than the ROM 802 and the RAM 803. The processor 801 may also perform various operations of the method flow according to the embodiments of the present disclosure by executing the programs stored in the one or more memories.
[0148] According to an embodiment of the present disclosure, the electronic device 800 may further include an input / output (I / O) interface 805, and the input / output (I / O) interface 805 is also connected to the bus 804. The electronic device 800 may further include one or more of the following components connected to the input / output (I / O) interface 805: an input portion 806 including a keyboard, a mouse, etc.; an output portion 807 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage portion 808 including a hard disk, etc.; and a communication portion 809 including a network interface card such as a LAN card, a modem, etc. The communication portion 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the input / output (I / O) interface 805 as needed. A removable medium 811, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is mounted on the drive 810 as needed so that a computer program read therefrom is installed into the storage portion 808 as needed.
[0149] The present disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist separately without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to the embodiments of the present disclosure is implemented.
[0150] According to an embodiment of the present disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, device, or apparatus. For example, according to an embodiment of the present disclosure, the computer-readable storage medium may include the ROM 802 and / or RAM 803 described above and / or one or more memories other than the ROM 802 and RAM 803.
[0151] An embodiment of the present disclosure also includes a computer program product, which includes a computer program that contains program code for executing the method shown in the flowchart. When the computer program product runs in a computer system, the program code is used to enable the computer system to implement the object detection processing method provided by the embodiment of the present disclosure.
[0152] When the computer program is executed by the processor 801, it executes the above functions defined in the system / apparatus of the embodiment of the present disclosure. According to an embodiment of the present disclosure, the above-described systems, apparatuses, modules, units, etc. can be implemented by computer program modules.
[0153] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium, and be downloaded and installed through the communication part 809, and / or be installed from the removable medium 811. The program code contained in the computer program can be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0154] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 809, and / or be installed from the removable medium 811. When the computer program is executed by the processor 801, it executes the above functions defined in the system of the embodiment of the present disclosure. According to an embodiment of the present disclosure, the above-described systems, devices, apparatuses, modules, units, etc. can be implemented by computer program modules.
[0155] According to embodiments of the present disclosure, program code for executing the computer programs provided by the embodiments of the present disclosure may be written in any combination of one or more programming languages. Specifically, these computing programs may be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. The programming languages include, but are not limited to, such as Java, C++, Python, the "C" language, or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device may be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or alternatively, may be connected to an external computing device (e.g., by connecting through the Internet using an Internet service provider).
[0156] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a portion of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and combinations of blocks in the block diagram or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0157] Those skilled in the art can understand that the features described in the various embodiments of the present disclosure can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present disclosure. In particular, without departing from the spirit and teachings of the present disclosure, the features described in the various embodiments of the present disclosure can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present disclosure.
[0158] The above describes the embodiments of the present disclosure. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Although the embodiments are described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. Without departing from the scope of the present disclosure, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present disclosure.
Claims
1. A processing method for object detection, characterized in that, The method includes: In response to a screen trigger operation, obtaining target image data; Processing the target image data based on a target detection model to obtain a target detection result, where the target detection model is a deep learning detection model optimized by fusing an attention mechanism and a knowledge distillation strategy; Judging an abnormal state based on the target detection result; and Processing an abnormal response based on the result of the abnormal state judgment.
2. The method according to claim 1, characterized in that The judging the abnormal state based on the target detection result specifically includes: Obtaining target quantity information and target position information from the target detection result; Making a first judgment based on the target quantity information to obtain a first state instruction; In response to the first state instruction being a target state instruction, making a second judgment based on the target position information to obtain a second state instruction; and Making an abnormal state judgment based on the first state instruction and / or the second state instruction.
3. The method according to claim 1 or 2, characterized in that, The optimization process of the target detection model includes: Selecting a target deep learning model as a teacher model, training the teacher model using a preset image data set, and outputting a soft label including a class probability distribution, where an attention mechanism module is embedded in the feature extraction network of the teacher model; Performing knowledge transfer training on a student model based on the soft label and a temperature adjustment parameter; and Evaluating the model based on the accuracy and inference speed of the trained student model in an edge environment to obtain the target detection model.
4. The method according to claim 3, wherein The performing knowledge transfer training on the student model based on the soft label and the temperature adjustment parameter specifically includes: In multiple training rounds, gradually reducing the value of the temperature adjustment parameter according to a preset attenuation strategy; In each training round, calculating a distillation loss value based on the soft label according to the temperature adjustment parameter corresponding to the current round; Obtaining a detection loss value of the student model under a true label; Performing weighted summation on the distillation loss value and the detection loss value according to a dynamic weight set in the current training stage to obtain a combined loss value; and Updating the parameters of the student model based on the combined loss value.
5. The method according to claim 2, wherein The obtaining the target quantity information and the target position information from the target detection result specifically includes: Screening out target bounding boxes with a confidence level exceeding a first preset threshold and belonging to a predetermined category from the target detection result, and obtaining the target quantity information and the target position information based on the quantity and position of the target bounding boxes.
6. The method according to claim 5, wherein The target state instruction represents a state instruction when the target quantity information is two targets. The making a second judgment based on the target position information in response to the first state instruction being the target state instruction to obtain a second state instruction specifically includes: Calculating the areas of the target bounding boxes of the two targets, and sorting based on the areas of the target bounding boxes; and Calculating the proportion of the target with the smaller area among the two targets in the area of the entire image region, and making a second judgment based on the proportion and a second preset threshold to obtain a second state instruction.
7. The method according to claim 2, wherein The target detection result at least includes the detection result of a face image.
8. According to the method described in any one of claims 1 to 2, 4 to 7, wherein the target image data includes consecutive image frames, The method further includes: Process the target image data using multi-object tracking based on deep learning to obtain identity association information of the same target between different image frames; and Based on the identity association information, perform temporal consistency verification on the target detection result to obtain a verification result.
9. The method according to claim 3, characterized in that The image data set includes training samples augmented via at least one image enhancement strategy, and the image enhancement strategy is used to change the spatial distribution characteristics and / or visual transformation characteristics of the image while keeping the image semantics unchanged.
10. A processing device for object detection, characterized in that, The device includes: A data acquisition module, configured to: in response to a screen trigger operation, acquire target image data; A target detection module, configured to: process the target image data based on a target detection model to obtain a target detection result, wherein the target detection model is a deep learning detection model optimized by fusing an attention mechanism and a knowledge distillation strategy; An abnormal state judgment module, configured to: judge an abnormal state based on the target detection result; and An abnormal processing module, configured to: process an abnormal response based on the result of the abnormal state judgment.
11. An electronic device, comprising: One or more processors; A memory for storing one or more computer programs, Characterized in that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 9.
12. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, When the computer program or instruction is executed by the processor, the steps of the method according to any one of claims 1 to 9 are implemented.
13. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or instruction is executed by the processor, the steps of the method according to any one of claims 1 to 9 are implemented.