Work auditing method and device, equipment, storage medium and product
By using target text recognition and people recognition models, the system automates the processing of work inspection videos, solving the problems of low inspection efficiency and accuracy in existing technologies and achieving efficient and accurate inspection results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 中国移动通信集团江西有限公司
- Filing Date
- 2024-10-31
- Publication Date
- 2026-05-01
AI Technical Summary
Existing technologies for work inspection are inefficient and inaccurate, mainly because the number of video frames manually identified is small and easily affected by visual fatigue, leading to missed detections and low efficiency.
By using target text recognition models and target number recognition models, the system automatically identifies the time and number of people in work inspection videos, calculates the inspection duration and number of people, and conducts inspections in conjunction with preset ranges.
It has automated the work inspection process, improved the efficiency and accuracy of inspections, and reduced errors caused by human error and visual fatigue.
Smart Images

Figure CN121963010A_ABST
Abstract
Description
Work audit methods, devices, equipment, storage media and products Technical Field
[0001] This application relates to the field of data processing technology, and in particular to work auditing methods, apparatus, equipment, storage media and products. Background Technology
[0002] Work auditing refers to checking and verifying the completion of maintenance work. Currently, the common method for work auditing is to use work order attachments to prove the number of personnel participating in the task, the start time of the task, and the end time. Furthermore, the work order video must be taken and uploaded at the start of the task, and the video must include the faces of all maintenance personnel involved. Auditors manually identify the number of personnel and the duration of the task before work auditing can proceed. Obviously, this method requires manual recording of each video. However, when personnel are focused on their work, they are bound to miss some frames. Moreover, since the videos are often long, visual fatigue is inevitable during viewing, and the number of frames viewed at one time is relatively small. Therefore, the efficiency and accuracy of this method for work auditing are low.
[0003] The above content is only used to help understand the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention
[0004] The main objective of this application is to provide a work audit method, apparatus, equipment, storage medium, and product, which aims to solve the technical problem of low efficiency and accuracy in work auditing in the prior art.
[0005] To achieve the above objectives, this application proposes a work audit method, the method comprising:
[0006] Obtain on-site work inspection videos of the locations to be audited;
[0007] Based on the target text recognition model, the first working time and the second working time are determined according to the work inspection video, and the audit duration is calculated based on the first working time and the second working time.
[0008] The target number of personnel was determined based on the aforementioned work inspection video;
[0009] The work audit is conducted based on the audit duration and the target number of personnel.
[0010] In one embodiment, the step of determining the first working time and the second working time based on the work inspection video using the target text recognition model includes:
[0011] Calculate the Euclidean distance between image features of different frames in the work inspection video;
[0012] The similarity between different frames is determined based on the Euclidean distance of the image features.
[0013] Based on the work inspection video, samples to be clustered are obtained, and each sample in the samples to be clustered is classified separately;
[0014] The first working time and the second working time are determined based on the similarity between the different frames and the single-class samples.
[0015] In one embodiment, the step of determining the first working time and the second working time based on the similarity between the different frames and single-class samples includes:
[0016] The similarity between samples of a single class is determined based on the similarity between the different frames;
[0017] The hierarchical clustering algorithm is used to cluster the samples based on their similarity to obtain the hierarchical clustering results.
[0018] The cluster centers of the hierarchical clustering results are calculated based on a preset clustering algorithm;
[0019] Calculate the distance between each frame and the cluster center, and determine the key video frames based on the distance;
[0020] The first working time and the second working time are determined based on the key video frames using the target text recognition model.
[0021] In one embodiment, the step of determining the first working time and the second working time based on the key video frames using the target text recognition model includes:
[0022] The key video frames are preprocessed to obtain initial key images;
[0023] The initial key image is detected based on a preset text region detection strategy to obtain an image containing text regions;
[0024] Based on the target character recognition model, characters are extracted from the image containing the text region to generate the target text;
[0025] The target text is filtered using regular expressions to obtain a string that meets the time format.
[0026] The first working time and the second working time are determined based on the string.
[0027] In one embodiment, the step of determining the target number of people based on the work inspection video includes:
[0028] The work inspection video is decomposed to obtain individual video frames;
[0029] Each video frame is preprocessed;
[0030] Feature extraction is performed on each preprocessed video frame to obtain the feature vector of each video frame;
[0031] Determine the target number recognition model and load the target number recognition model using preset parameters;
[0032] Based on the loaded target number recognition model, the label positions and categories of employees and security equipment in each video frame are predicted according to the feature vectors of each video frame.
[0033] The target number of people is determined based on the label location and category;
[0034] Work audits are conducted based on the stated audit duration and target number of personnel.
[0035] In one embodiment, the step of determining the target number identification model includes:
[0036] Obtain an initial historical video frame sample set, and mark the user's location and the location of the security device in each video frame in the initial historical video frame sample set based on preset tags;
[0037] Assign a first category label to the labeled users, and assign a second category label to the labeled security devices;
[0038] Generate a sample set of target historical video frames based on the marked locations and category labels;
[0039] The model for identifying the number of people in a target is determined by training a model based on a set of historical video frames of the target using a target detection algorithm.
[0040] Furthermore, to achieve the above objectives, this application also proposes a work auditing device, which includes:
[0041] The acquisition module is used to acquire on-site work inspection videos to be audited.
[0042] The determination module is used to determine the first working time and the second working time based on the target text recognition model according to the work inspection video, and to calculate the audit duration based on the first working time and the second working time;
[0043] The determining module is also used to determine the target number of people based on the work inspection video;
[0044] The audit module is used to conduct work audits based on the audit duration and the target number of personnel.
[0045] In addition, to achieve the above objectives, this application also proposes a work auditing device, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the work auditing method as described above.
[0046] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and which, when executed by a processor, implements the steps of the work audit method described above.
[0047] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the work audit method described above.
[0048] One or more technical solutions proposed in this application have at least the following technical effects: acquiring a work inspection video of the site to be inspected; determining a first working time and a second working time based on the work inspection video using a target character recognition model; calculating the inspection duration based on the first working time and the second working time; determining the target number of people based on the work inspection video; and conducting work inspection based on the inspection duration and the target number of people. Through the above method, after acquiring the work inspection video of the site to be inspected, a target character recognition model is used to identify the work inspection video to determine the first working time and the second working time, and further, the inspection duration and the target number of people are automatically calculated. Then, work inspection is conducted from both the duration and number of people dimensions, thereby achieving automated work inspection and effectively improving the efficiency and accuracy of work inspection. Attached Figure Description
[0049] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0050] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0051] Figure 1 is a flowchart illustrating the working audit method of this application in Implementation Example 1.
[0052] Figure 2 is a flowchart of the second embodiment of the audit method of this application.
[0053] Figure 3 is a simplified flowchart of the work audit method provided in Embodiment 2 of this application;
[0054] Figure 4 is a schematic diagram of the module structure of the working audit device according to an embodiment of this application;
[0055] Figure 5 is a schematic diagram of the hardware operating environment involved in the work audit method in this application embodiment.
[0056] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0057] It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an electronic device or work auditing device capable of performing the above functions. The following description uses a work auditing device as an example to illustrate this embodiment and the subsequent embodiments.
[0058] Based on this, this application provides a work auditing method. Referring to Figure 1, which is a flowchart of the first embodiment of the work auditing method of this application.
[0059] In this embodiment, the work auditing method includes steps S10 to S40:
[0060] Step S10: Obtain the on-site work inspection video of the site to be audited.
[0061] It should be noted that work inspection videos refer to videos collected at the audit site through an application. Work inspection videos ensure the authenticity of on-site work and cannot be falsified. The condition for triggering the acquisition of work inspection videos at the audit site can be the issuance of a work task. The content of this work task includes, but is not limited to, the task title, task description, task specialty, task type, task content, required number of personnel, and time. Different specialties correspond to different task types; please refer to Table 1 for details.
[0062] Table 1:
[0063]
[0064] Step S20: Based on the target text recognition model, determine the first working time and the second working time according to the work inspection video, and calculate the audit duration based on the first working time and the second working time.
[0065] It is understandable that the target text recognition model refers to a model used to recognize text in an image. After acquiring the work inspection video, the target text recognition model is used to match strings that meet the time format, and all work times are extracted from the strings that meet the time format. The first work time (earliest time) and the second work time (latest time) are determined from all work times. The audit work order can correspond to one or more audit videos of the site to be audited, and each audit video has multiple key video frames. The first work time and the second work time can be determined from each key video frame in the above way, and the audit duration can be further calculated.
[0066] It should be understood that audit duration refers to the time required for conducting a work audit. This audit duration can be calculated based on the first working time and the second working time, specifically as follows:
[0067] Duration=EndTime-StartTime.
[0068] Where Duration represents the audit duration, StartTime represents the first working time, and EndTime represents the second working time.
[0069] Step S30: Determine the target number of people based on the work inspection video.
[0070] Further, step S30 includes: decomposing the work inspection video to obtain video frames; preprocessing each video frame; extracting features from each preprocessed video frame to obtain feature vectors for each video frame; determining a target number recognition model and loading the target number recognition model using preset parameters; predicting the tag positions and categories of employees and security equipment in each video frame based on the loaded target number recognition model and the feature vectors of each video frame; and determining the target number of people based on the tag positions and categories.
[0071] It is understandable that the target number refers to the actual number of people wearing safety equipment. Preprocessing operations on each video frame include, but are not limited to, resizing images, normalizing pixel values, and image enhancement. The purpose of these preprocessing operations is to ensure that the preprocessing methods used on the input video frames are consistent with those used when training the model.
[0072] It is understandable that the target number recognition model refers to a model used to predict the label positions and categories of employees and security equipment in each video frame. The input of the target number recognition model is the feature vector of the video frame. The preset parameters refer to the parameters that load the weights and configuration information of the target number recognition model. These preset parameters can be functions or classes. The device used to extract the feature vectors of each video frame can be a feature extractor. This feature extractor can be a convolutional neural network model used during training. Through forward propagation, it extracts the preprocessed feature vectors of each video frame, i.e., the feature vectors of each video frame. Then, the feature vectors of each video frame are input into the target number recognition model. The target number recognition model outputs the feature vectors of each video frame to predict the label positions and categories of employees and security equipment in each video frame. The above process can be implemented through forward propagation. Then, the output label positions and categories need to be post-processed. The post-processing operations include, but are not limited to, deduplication, applying non-maximum suppression to merge overlapping bounding boxes, etc. Then, combined with the confidence threshold, the number of people wearing security equipment is determined and counted, which is the target number.
[0073] Furthermore, in order to effectively improve the accuracy of the target number recognition model, the step of determining the target number recognition model includes: obtaining an initial historical video frame sample set; labeling the user's position and the security device's position in each video frame of the initial historical video frame sample set based on preset labels; assigning a first category label to the labeled user and a second category label to the labeled security device; generating a target historical video frame sample set based on the labeled position and the classified labels; and training the model using a target detection algorithm based on the target historical video frame sample set to determine the target number recognition model.
[0074] It is understood that the preset label can be a bounding box, the shape of which can be a rectangle or other polygons. This embodiment does not limit this. After obtaining the initial historical video frame sample set, each video frame in the initial historical video frame sample set is loaded and displayed in the standard tool. Then, the drawing tool in the annotation tool is used to annotate the position of the user and the position of the safety equipment in the image. During the annotation process, the size and position of the bounding box can be adjusted according to the actual situation to accurately surround the user and the safety equipment. Then, category labels need to be assigned. For the user, a first category label is assigned, which can be "person". For the safety equipment, a second category label is assigned, which can be "safety equipment". The safety equipment can be a safety helmet. In addition, it should be noted that in complex scenarios where there may be multiple people and safety helmets, different colors or numbers can be used to distinguish them.
[0075] It's important to note that during the annotation process, keyboard shortcuts or tools can be used to improve efficiency, such as copying and pasting annotated bounding boxes. After annotation is complete, the annotated locations and category labels can be saved in a standard data format for subsequent model training and evaluation. This standard data format includes, but is not limited to, XML, JSON, and CSV formats. Furthermore, during annotation, it's crucial to ensure that annotators possess strong visual judgment and accurate annotation skills; adhere to consistent annotation specifications and standards to ensure the consistency and reliability of the results; for complex or ambiguous scenarios, add annotations or explanations using the annotation tool for subsequent data analysis and model training; and regularly check and verify the quality of the annotation results to ensure accuracy and consistency.
[0076] It should be understood that after obtaining the labeled positions and classification labels in the standard data format, a sample set of historical video frames of the target is generated. The target detection algorithm can be the Faster R-CNN algorithm. The network structure of the model is constructed based on this target detection algorithm, including a feature extraction layer, a target detection layer, and a classification layer. The deep learning PyTorch framework can be used to build and train the model. On the other hand, a loss function also needs to be defined. This loss function can be the mean squared error function or the cross-entropy loss function; this embodiment does not impose any restrictions on this.
[0077] Understandably, after obtaining the target historical video frame sample set, the video frames in the sample set are preprocessed in the same way to obtain historical video frame feature vectors. These feature vectors are then used as input, with labeled bounding boxes and categories provided as the target output. The model parameters are optimized using the backpropagation algorithm to accurately predict the location and category of people and security equipment. During training, the Adam optimizer can be used to adjust the learning rate and optimization strategy. The model can also be evaluated during training. Current model performance evaluation metrics include, but are not limited to, accuracy, recall, and F1 score. If the current model performance evaluation metrics do not meet the preset requirements, model tuning is needed, such as adjusting hyperparameters, increasing training data, and data augmentation. Furthermore, the tuned people recognition model is tested. If it is successful, the target people recognition model is determined. At this point, the weights and configuration information of the target people recognition model can be saved for subsequent prediction and deployment.
[0078] Step S40: Conduct a work audit based on the audit duration and the target number of personnel.
[0079] It should be understood that after obtaining the target number of people, the work audit should be carried out in conjunction with the audit duration, that is, whether the audit duration is within the time limit and whether the target number of people is within the number limit.
[0080] It should be understood that if the audit duration is not within the specified range and / or the target number of people is not within the specified range, a manual review process will be added based on the abnormal situation. The manual review will determine whether the image recognition is normal. If there is an anomaly, the icon will be labeled as a sample and used as a sample for optimizing the target number recognition model, thereby further improving the model accuracy.
[0081] It should be noted that the duration and number of participants vary depending on the major; please refer to Table 2 for details.
[0082] Table 2:
[0083]
[0084] This embodiment acquires a work inspection video of the site to be audited; determines a first working time and a second working time based on the work inspection video using a target text recognition model; calculates the audit duration based on the first working time and the second working time; determines the target number of people based on the work inspection video; and performs a work audit based on the audit duration and the target number of people. Through this method, after acquiring the work inspection video of the site to be audited, a target text recognition model is used to identify the work inspection video to determine the first working time and the second working time. Furthermore, the audit duration and the target number of people are calculated automatically, and then the work audit is performed from both the duration and number of people dimensions. This enables automated work auditing, thereby effectively improving the efficiency and accuracy of work auditing.
[0085] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to that in the first embodiment described above can be referred to the above description and will not be repeated hereafter. Based on this, please refer to Figure 2; step S20 includes steps S201 to S204:
[0086] Step S201: Calculate the Euclidean distance between the image features of different frames of the work inspection video.
[0087] It should be noted that after obtaining the on-site inspection video, the Euclidean distance between different frames of the video can be calculated using a preset distance calculation strategy. Specifically:
[0088]
[0089] in, d represents the sum of squared differences in the positions of all pixels in a frame image. euc (A,B) represents the Euclidean distance between image feature A and image feature B.
[0090] Step S202: Determine the similarity between different frames based on the Euclidean distance of the image features.
[0091] It should be understood that since there is a specific correspondence between the Euclidean distance of image features between different frames and the similarity between different frames, for example, the smaller the Euclidean distance, the higher the similarity, after obtaining the Euclidean distance of image features, the similarity between different frames can be determined according to this specific correspondence.
[0092] Step S203: Obtain samples to be clustered based on the work inspection video, and classify each sample in the samples to be clustered separately.
[0093] Understandably, after obtaining the work inspection video, a keyframe extraction process based on work inspection video clustering is used to obtain the work inspection video sequence. At this point, the work inspection video sequence contains multiple clustering samples, which are the samples to be clustered. Before classifying each sample in the samples to be clustered separately, a termination condition needs to be determined, specifically, the number of extracted keyframes is K. The mean M and variance V of the feature vector are calculated; where D = {di | i = 1, 2, 3, ..., N} represents the distance vector, and the value of K is equal to the current Euclidean distance d > M + 2 * T, where T represents the number of video frames. After the above processing, each sample in the samples to be clustered is classified into a separate class.
[0094] Step S204: Determine the first working time and the second working time based on the similarity between the different frames and the single-class samples.
[0095] It should be understood that after obtaining a single-class sample, the two closest single-class samples are found based on the inter-frame similarity, and then the merging is performed multiple times until the number of clusters is K. Then, the first working time and the second working time are determined based on the hierarchical clustering results.
[0096] Furthermore, to effectively improve the accuracy of determining key video frames, step S204 includes: determining the similarity between single-class samples based on the similarity between different frames; performing clustering based on the similarity between the single-class samples using a hierarchical clustering algorithm to obtain hierarchical clustering results; calculating the cluster centers of the hierarchical clustering results based on a preset clustering algorithm; calculating the distance between each frame and the cluster centers, and determining key video frames based on the distances; and determining a first working time and a second working time based on the key video frames using a target text recognition model.
[0097] Understandably, after obtaining the similarity between different frames, the similarity between single-class samples is calculated sequentially based on the similarity between different frames. Then, the two closest single-class samples are found and grouped into one class. The similarity between the newly grouped class and other single-class samples is recalculated, and the search for the two closest single-class samples continues. The number of clusters is counted in real time. When the number of clusters reaches K, the classification ends. The type of this hierarchical clustering algorithm can be agglomerative.
[0098] It should be understood that the preset clustering algorithm can be the K-means clustering algorithm. After obtaining the hierarchical clustering result, the clustering result is used as the input of the K-means clustering algorithm. The center of each class generated by the hierarchical clustering is used as the initial cluster center of the K-means algorithm. This can avoid the influence of randomly generated initial cluster centers on the clustering result of the K-means algorithm.
[0099] It should be noted that the cluster center refers to the center of the cluster generated by hierarchical clustering, that is, the average value of each cluster object. After calculating the distance between each frame and the cluster center, the corresponding frames are reclassified according to the minimum distance criterion, and the cluster center of each class is recalculated. Then the distance is calculated again until the objects in each cluster no longer change. At this time, the cluster center of each class is calculated, and the frame closest to the cluster center is output as the key frame, that is, the key video frame is obtained. The target text recognition model is used to determine the first working time and the second working time.
[0100] Furthermore, to effectively improve the accuracy of determining the first and second working times, the step of determining the first and second working times based on the target character recognition model according to the key video frames includes: preprocessing the key video frames to obtain an initial key image; detecting the initial key image based on a preset text region detection strategy to obtain an image containing text regions; extracting characters from the image containing text regions based on the target character recognition model to generate target text; filtering the target text based on regular expressions to obtain a string that meets the time format; and determining the first and second working times based on the string.
[0101] It should be understood that in order to effectively improve the accuracy and stability of character recognition, key video frames need to be preprocessed. This preprocessing operation includes, but is not limited to, grayscale conversion, binarization, and denoising. The preset text region detection strategy refers to the strategy used to detect regions containing text in an image. Regions containing text are usually the boundaries of characters or text lines. In order to convert characters into numerical representations that can be processed by a classifier, it is necessary to extract the features of each character. These features include, but are not limited to, shape, angle, and texture. The classifier used for character classification can be a machine learning algorithm or a deep learning model. Then, the classified characters are cleaned. This cleaning operation includes, but is not limited to, correcting errors, correcting tilt, and removing redundancy. Finally, the cleaned characters are used to generate the target text, which can be a single character, a word, or a complete text.
[0102] Understandably, since the text recognition results may contain other characters, character anomalies, or missing characters, it is necessary to match and filter out strings that meet the time format using regular expressions. These strings that meet the time format are equivalent to time watermarks. Then, the time format is unified, and the earliest and latest times are extracted from them, namely the first working time and the second working time.
[0103] This embodiment calculates the Euclidean distance of image features between different frames of the work inspection video; determines the similarity between different frames based on the Euclidean distance of the image features; obtains samples to be clustered based on the work inspection video, and classifies each sample in the samples to be clustered separately; determines the first working time and the second working time based on the similarity between different frames and the single-class samples. Through the above method, the Euclidean distance of image features between different frames is calculated using a preset distance calculation strategy, and then the similarity between different frames corresponding to the current Euclidean distance is determined according to a specific correspondence. On the other hand, a hierarchical clustering algorithm is used to classify each sample in the samples to be clustered as one class. At this time, the similarity of the classified samples is calculated by combining the similarity between different frames, and the first working time and the second working time are further determined based on the video keyframes using regular expressions, thereby effectively improving the accuracy of determining the first working time and the second working time.
[0104] For example, to help understand the implementation flow of the work audit method obtained by combining this embodiment with the above embodiment one, please refer to Figure 3. Figure 3 provides a simplified flowchart of a work audit method, specifically:
[0105] When starting a work inspection, video footage of the site to be inspected is acquired. On one hand, a pre-trained target number recognition model is used to predict the location and category of tags for employees and safety equipment in each video frame; the target number of employees is determined based on the tag location and category. On the other hand, a target text recognition model is used to determine the first and second working times, and the inspection duration is calculated using the first and second working times. Finally, the work inspection is conducted using the inspection duration and the target number of employees, thereby achieving automated work inspection and effectively improving the efficiency and accuracy of work inspection.
[0106] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the working audit method of this application. Any simple modifications based on this technical concept are within the protection scope of this application.
[0107] This application also provides a work audit device, as shown in Figure 4, the work audit device comprising:
[0108] Module 10 is used to acquire work inspection videos of the site to be audited.
[0109] The determination module 20 is used to determine the first working time and the second working time based on the target text recognition model according to the work inspection video, and to calculate the audit duration based on the first working time and the second working time.
[0110] The determining module 20 is also used to determine the target number of people based on the work inspection video.
[0111] Audit module 30 is used to conduct work audits based on the audit duration and the target number of personnel.
[0112] This embodiment acquires a work inspection video of the site to be audited; determines a first working time and a second working time based on the work inspection video using a target text recognition model; calculates the audit duration based on the first working time and the second working time; determines the target number of people based on the work inspection video; and performs a work audit based on the audit duration and the target number of people. Through this method, after acquiring the work inspection video of the site to be audited, a target text recognition model is used to identify the work inspection video to determine the first working time and the second working time. Furthermore, the audit duration and the target number of people are calculated automatically, and then the work audit is performed from both the duration and number of people dimensions. This enables automated work auditing, thereby effectively improving the efficiency and accuracy of work auditing.
[0113] The work audit device provided in this application, employing the work audit method described in the above embodiments, can solve the technical problem of low efficiency and accuracy in existing work audits. Compared with the prior art, the beneficial effects of the work audit device provided in this application are the same as those of the work audit method described in the above embodiments, and other technical features in the work audit device are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.
[0114] In one embodiment, the determining module 20 is further configured to calculate the Euclidean distance of image features between different frames of the work inspection video; determine the similarity between different frames based on the Euclidean distance of the image features; obtain samples to be clustered based on the work inspection video, and classify each sample in the samples to be clustered separately; and determine the first working time and the second working time based on the similarity between different frames and the single-class samples.
[0115] In one embodiment, the determining module 20 is further configured to: determine the similarity between single-class samples based on the similarity between different frames; perform clustering based on the similarity between the single-class samples using a hierarchical clustering algorithm to obtain hierarchical clustering results; calculate the cluster centers of the hierarchical clustering results based on a preset clustering algorithm; calculate the distance between each frame and the cluster centers, and determine key video frames based on the distances; and determine a first working time and a second working time based on the key video frames using a target text recognition model.
[0116] In one embodiment, the determining module 20 is further configured to preprocess the key video frame to obtain an initial key image; detect the initial key image based on a preset text region detection strategy to obtain an image containing a text region; extract characters from the image containing the text region based on a target character recognition model to generate target text; filter the target text based on regular expressions to obtain a string that meets the time format; and determine a first working time and a second working time based on the string.
[0117] In one embodiment, the determining module 20 is further configured to decompose the work inspection video to obtain video frames; preprocess the video frames; extract features from the preprocessed video frames to obtain feature vectors for each video frame; determine a target number recognition model and load the target number recognition model using preset parameters; predict the tag positions and categories of employees and security equipment in each video frame based on the loaded target number recognition model and the feature vectors of each video frame; and determine the target number of people based on the tag positions and categories.
[0118] In one embodiment, the determining module 20 is further configured to acquire an initial historical video frame sample set, mark the location of the user and the location of the security device in each video frame of the initial historical video frame sample set based on preset labels; assign a first category label to the marked user and a second category label to the marked security device; generate a target historical video frame sample set according to the marked location and the classified labels; and determine a target number recognition model by training a model based on the target historical video frame sample set using a target detection algorithm.
[0119] This application provides a work auditing device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the work auditing method in Embodiment 1 above.
[0120] Referring to Figure 5 below, a schematic diagram of a work audit device suitable for implementing embodiments of this application is shown. The work audit device in embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. The work audit device shown in Figure 5 is merely an example and should not impose any limitations on the functionality and scope of use of embodiments of this application.
[0121] As shown in Figure 5, the auditing equipment may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for the operation of the auditing equipment. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to I / O interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. Communication device 1009 allows the audit equipment to communicate wirelessly or wiredly with other devices to exchange data. Although the figure shows audit equipment with various systems, it should be understood that it is not required to implement or possess all the systems shown. More or fewer systems may be implemented alternatively.
[0122] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0123] The work audit equipment provided in this application, employing the work audit method described in the above embodiments, can solve the technical problem of low efficiency and accuracy in work auditing in the prior art. Compared with the prior art, the beneficial effects of the work audit equipment provided in this application are the same as those of the work audit method provided in the above embodiments, and other technical features of this work audit equipment are the same as those disclosed in the previous embodiment method, and will not be repeated here.
[0124] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0125] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0126] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to perform the work auditing method described in the above embodiments.
[0127] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0128] The aforementioned computer-readable storage medium may be included in the work audit equipment; or it may exist independently and not assembled into the work audit equipment.
[0129] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0130] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0131] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0132] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., computer programs) for executing the above-described work auditing method, thereby solving the technical problem of low efficiency and accuracy in existing work auditing techniques. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the work auditing method provided in the above embodiments, and will not be repeated here.
[0133] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the work audit method described above.
[0134] The computer program product provided in this application can solve the technical problem of low efficiency and accuracy in existing work auditing techniques. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the work auditing methods provided in the above embodiments, and will not be repeated here.
[0135] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.
Claims
1. A work auditing method, characterized in that, The method includes: acquiring a work inspection video of the site to be inspected; determining a first working time and a second working time based on the work inspection video using a target character recognition model; calculating the inspection duration based on the first working time and the second working time; determining the target number of people based on the work inspection video; and conducting a work inspection based on the inspection duration and the target number of people.
2. The method as described in claim 1, characterized in that, The step of determining the first working time and the second working time based on the target text recognition model according to the work inspection video includes: calculating the Euclidean distance of image features between different frames of the work inspection video; determining the similarity between different frames according to the Euclidean distance of the image features; obtaining samples to be clustered according to the work inspection video, and classifying each sample in the samples to be clustered separately; determining the first working time and the second working time according to the similarity between different frames and the single-class samples.
3. The method as described in claim 2, characterized in that, The step of determining the first working time and the second working time based on the similarity between different frames and single-class samples includes: determining the similarity between single-class samples based on the similarity between different frames; performing clustering based on the similarity between single-class samples using a hierarchical clustering algorithm to obtain hierarchical clustering results; calculating the cluster centers of the hierarchical clustering results based on a preset clustering algorithm; calculating the distance between each frame and the cluster centers, and determining key video frames based on the distances; and determining the first working time and the second working time based on the key video frames using a target text recognition model.
4. The method as described in claim 3, characterized in that, The steps of determining the first working time and the second working time based on the key video frame using the target character recognition model include: preprocessing the key video frame to obtain an initial key image; detecting the initial key image based on a preset text region detection strategy to obtain an image containing a text region; extracting characters from the image containing the text region based on the target character recognition model to generate target text; filtering the target text based on regular expressions to obtain a string that meets the time format; and determining the first working time and the second working time based on the string.
5. The method as described in claim 1, characterized in that, The step of determining the target number of people based on the work inspection video includes: decomposing the work inspection video to obtain video frames; preprocessing each video frame; extracting features from each preprocessed video frame to obtain feature vectors for each video frame; determining a target number of people recognition model and loading the target number of people recognition model with preset parameters; predicting the tag positions and categories of employees and security equipment in each video frame based on the feature vectors of each video frame using the loaded target number of people recognition model; and determining the target number of people based on the tag positions and categories.
6. The method as described in claim 5, characterized in that, The steps for determining the target number recognition model include: acquiring an initial historical video frame sample set; marking the user's location and the security device's location in each video frame of the initial historical video frame sample set based on preset labels; assigning a first category label to the marked users and a second category label to the marked security devices; generating a target historical video frame sample set based on the marked locations and classified labels; and training the model using a target detection algorithm based on the target historical video frame sample set to determine the target number recognition model.
7. A work auditing device, characterized in that, The device includes: an acquisition module for acquiring a work inspection video of the site to be audited; a determination module for determining a first working time and a second working time based on a target character recognition model according to the work inspection video, and calculating the audit duration based on the first working time and the second working time; the determination module is also used to determine a target number of people based on the work inspection video; and an audit module for performing a work audit based on the audit duration and the target number of people.
8. A work auditing device, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the work audit method as described in any one of claims 1 to 6.
9. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the work auditing method as described in any one of claims 1 to 6.
10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the steps of the work auditing method as described in any one of claims 1 to 6.