Work ticket information extraction and work video linkage analysis method
Through OCR technology and language model processing work tickets, combined with convolutional neural network and timing convolutional network to analyze videos, the work ticket information extraction and job video linkage analysis is realized, solving the problems of information omission and real-time monitoring under traditional methods, and significantly improving job security and standardization.
Patent Information
- Application Number
- CN202411818555.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-10
- Publication Date
- 2025-05-06
AI Technical Summary
Traditional work ticket management and operation monitoring methods are prone to information omissions and recording errors, and it is difficult to grasp the progress of the work in real time, and it is impossible to detect and deal with on-site problems in a timely manner.
OCR technology is used to convert work tickets into processable text formats, and text cleaning technology is used to remove irrelevant information. Use pre-trained language models to identify named entities and extract relationships to generate structured data tables. The video is processed, compressed and enhanced in frames, and the object detection model of convolutional neural network and timing convolutional network is applied to identify people, devices and important scenes in the video, and analyze behavioral actions. By matching the structured information in the work ticket with the video detection results, establishing a timeline, comparing the task arrangement in the work ticket with the actual operations in the video, dynamic compliance verification and risk assessment are achieved.
It improves the efficiency and accuracy of information processing, realizes real-time monitoring and dynamic analysis, timely discovers and handles on-site problems, and significantly improves the safety and standardization of operations.
Smart Images

Figure CN119942568A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the technical field of power system data processing, and more specifically, to a method for extracting work ticket information and linking operation video analysis. Background Art
[0002] The power grid system is a key infrastructure for the country's energy supply, involving the production, transmission, distribution and consumption of electricity. In order to ensure the safe operation of the power grid, power companies need strict management and monitoring measures. In this process, the work ticket system has become an important means to ensure the safety of power grid operations. The work ticket records in detail the content, personnel, time, equipment and other key information of each operation task to ensure the planning and standardization of the operation. However, with the expansion of the scale of the power grid and the improvement of its intelligence, the traditional work ticket management and operation monitoring methods face huge challenges.
[0003] Traditional work ticket management mainly relies on paper documents. The specific process is as follows:
[0004] Work ticket application: The operator fills out a paper work ticket according to the work requirements, recording in detail the task content, schedule, participants, equipment used and other information.
[0005] Review and approval: The completed work ticket needs to go through multiple levels of review and is usually signed and approved by team leaders, technicians, dispatchers and other relevant personnel. It will only take effect after confirmation.
[0006] Operation implementation: During the operation, the operator needs to carry the work ticket with him and operate according to the contents on the work ticket. Any changes need to be recorded on the work ticket and re-signed for confirmation.
[0007] Retrieval and archiving: After the work is completed, the work ticket must be returned to the management department for archiving for subsequent audit and inspection.
[0008] In traditional work environments, job monitoring relies primarily on manual monitoring and reporting:
[0009] On-site monitoring: At key work sites, special supervisors are usually arranged to conduct on-site supervision to ensure that the work is carried out strictly in accordance with the requirements of the work ticket.
[0010] Inspection and record keeping: Inspectors conduct regular on-site inspections, record the progress of operations and any problems found, and ensure safe and standardized operations.
[0011] Post-operation report: After the operation is completed, the operator needs to fill out a detailed operation report summarizing the operation process, problems encountered and solutions.
[0012] The traditional method is prone to information omissions and recording errors; and manual monitoring and reporting methods make it difficult to grasp the progress of operations in real time, and it is impossible to discover and deal with on-site problems in a timely manner. Summary of the invention
[0013] The present invention provides a method for extracting work ticket information and analyzing the linkage with operation video, which is intended to solve the technical problems that traditional methods are prone to information omissions and recording errors, and manual monitoring and reporting methods are difficult to grasp the progress of operations in real time, and cannot promptly discover and deal with on-site problems.
[0014] The method for extracting work ticket information and linking operation video analysis includes the following steps:
[0015] Step 1: Use OCR technology to convert the work ticket into a processable text format, and apply text cleaning technology to remove phonetic characters and irrelevant information;
[0016] Step 2: Use the pre-trained language model to perform named entity recognition, extract key entities, and then use relationship extraction technology to identify the relationship between entities, generate a structured data table, and store the key information of the work ticket;
[0017] Step 3: Process the video by framing, extract key frames, and apply video compression and enhancement technology to ensure video quality and processing efficiency;
[0018] Step 4: Use the target detection model formed by convolutional neural network and temporal convolutional network to perform target detection, identify people, equipment and important scenes in the video, and analyze the behavior to obtain the video detection results;
[0019] Step 5: Match the structured information of the work ticket type with the detection results in the video. By establishing a timeline, compare the task arrangement in the work ticket with the actual operation in the video, and perform compliance verification and risk assessment based on the comparison.
[0020] The present invention uses OCR technology to convert the work ticket into a text format, and applies text cleaning technology to remove irrelevant information to ensure the integrity and accuracy of the data. Then, the pre-trained language model is used to perform named entity recognition and relationship extraction, generate a structured data table, and extract the key information of the work ticket. Then, the video is frame-processed and compressed to ensure video quality and processing efficiency, and the target detection model of the convolutional neural network and the temporal convolutional network is applied to identify the characters, equipment and important scenes in the video, analyze the behavior and obtain the video detection results. By matching the structured information in the work ticket with the video detection results, establishing a timeline, and comparing the task arrangement in the work ticket with the actual operation in the video, dynamic compliance verification and risk assessment are achieved. This method not only improves the efficiency and accuracy of information processing, but also realizes real-time monitoring and dynamic analysis, timely discovers and handles on-site problems, and significantly improves the safety and standardization of operations.
[0021] Preferably, step 2 comprises the following steps:
[0022] Data preprocessing: Use the BERT tokenizer to tokenize the text and convert it into a token sequence.
[0023] Model reasoning: Input the preprocessed data into the BERT model and perform forward propagation to obtain the hidden state of the output:
[0024] H = BERT(I,A,S);
[0025] Where: H represents the output hidden state of the BERT model; I represents the input ID sequence; A represents the attention mask; S represents the segment ID;
[0026] Label prediction: Classify each token through a fully connected layer and Softmax function to obtain the probability that each token belongs to each entity category, and select the label with the highest probability:
[0027] P i =Softmax(W·H i + b);
[0028] Where: P i represents the predicted label probability distribution of the i-th Token; W represents the weight matrix of the classification layer; H i represents the hidden state of the i-th Token; b represents the bias vector of the classification layer;
[0029] Relation extraction:
[0030] After obtaining each entity, construct entity pairs, predict the connection between entity pairs through the relation extraction model, and use a variant of BERT to classify each entity pair and predict the relationship type:
[0031] R ij =Softmax((W r ·[H i ;H j ]+b r );
[0032] Where: R ij represents the predicted probability distribution of the relationship between entity i and entity j; W r Represents the weight matrix of the relationship classification layer; [H i ;H j ] represents the hidden state concatenation vector of entity i and entity j; b r Represents the bias vector of the relation classification layer;
[0033] The results of named entity recognition and relationship extraction are structured to generate a structured data table.
[0034] Preferably, step 3 comprises the following steps:
[0035] Video frame processing: Use OpenCV to read video files and obtain basic properties of the video; read the video frame by frame and save each frame as an independent image file;
[0036] Key frame extraction: Convert each frame to an appropriate color space, and then calculate the color histogram. Using the HSV color space, the histogram calculation formula is as follows:
[0037] H i ={h i,j};
[0038] Where: H i represents the color histogram of the i-th frame; h i,j Represents the value of the jth bin in the histogram, indicating the number of pixels contained in the bin;
[0039] h i,j =∑ x,y δ(B(I i (x,y))-j);
[0040] Where: I i (x, y) represents the pixel value at position (x, y) of the i-th frame; δ represents the Dirac delta function, which outputs 1 when the input is 0, otherwise it outputs 0; B represents the assignment of pixel values to the corresponding bins;
[0041] In HSV space, the calculation of histogram is divided into three histograms through (H, S, V):
[0042]
[0043] Where: Respectively represent the histogram of the i-th frame in H, S, and V channels; Respectively represent the value of the jth bin in the corresponding channel;
[0044] The magnitude of change between frames is measured by calculating the difference between the color histograms of adjacent frames:
[0045]
[0046] Where: D i,i+1 represents the histogram difference between the i-th frame and the i+1-th frame; B represents the number of bins in the histogram;
[0047] Set a threshold. If the histogram difference exceeds the threshold, the frame is considered a key frame.
[0048] Video compression: Use FFmpeg to compress the video;
[0049] Video enhancement: Gaussian filtering is used to reduce noise in the video;
[0050] Contrast Enhancement: Enhances the contrast of the video using histogram equalization.
[0051] Preferably, step 4 comprises the following steps:
[0052] Data preprocessing: extract frames from the enhanced video data in step 3 at a fixed frame rate, scale and distort each extracted frame, and cache consecutive frames in chronological order to form a frame sequence;
[0053] Object detection module: The input is a single-frame image after data preprocessing; through multiple layers of convolution and residual connection, deep features are extracted, feature maps are output, and feature maps of different scales are extracted. They are fused through a path aggregation network to output multi-scale feature maps; object classification and positioning are performed at each scale of the multi-scale feature map, and multi-scale detection results are output, including bounding box coordinates and category probabilities, based on which people, equipment and important scenes in the video are obtained;
[0054] Action recognition module: The input is the output result of the target detection module; features are extracted from the target detection results, and the features of all detection boxes are combined into a feature vector to form a feature sequence; the feature sequence is then processed through multiple one-dimensional convolutional layers, and batch normalization and ReLU activation functions are connected after each convolution layer to output convolution features. Pooling is performed after each convolution layer to reduce the sequence length and output pooled features; the features after the last layer of convolution and pooling are flattened, and the probability distribution of the action category is obtained through a fully connected layer.
[0055] Preferably, the loss function of the target detection model is as follows:
[0056]
[0057] Where: Represents the positioning loss, which measures the difference between the predicted box and the real box, using CloU; represents the confidence loss; represents the classification loss; λ loc represents the weight of the positioning loss; conf represents the weight of execution loss; cls Represents the weight of classification loss; α det Represents the weight of target detection loss; α act represents the weight of action recognition loss; N represents the number of samples; C represents the number of action categories; y i,c Represents the cth class true label of the i-th sample; represents the predicted probability of the cth class of the ith sample.
[0058] The beneficial effects of the present invention include:
[0059] The present invention uses OCR technology to convert the work ticket into a text format, and applies text cleaning technology to remove irrelevant information to ensure the integrity and accuracy of the data. Then, the pre-trained language model is used to perform named entity recognition and relationship extraction, generate a structured data table, and extract the key information of the work ticket. Then, the video is frame-processed and compressed to ensure video quality and processing efficiency, and the target detection model of the convolutional neural network and the temporal convolutional network is applied to identify the characters, equipment and important scenes in the video, analyze the behavior and obtain the video detection results. By matching the structured information in the work ticket with the video detection results, establishing a timeline, and comparing the task arrangement in the work ticket with the actual operation in the video, dynamic compliance verification and risk assessment are achieved. This method not only improves the efficiency and accuracy of information processing, but also realizes real-time monitoring and dynamic analysis, timely discovers and handles on-site problems, and significantly improves the safety and standardization of operations. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0061] Figure 1 An overall step block diagram provided for an embodiment of the present invention. DETAILED DESCRIPTION
[0062] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0063] See also Figure 1 As shown, the best embodiment of the present invention is further described;
[0064] The method for extracting work ticket information and linking operation video analysis includes the following steps:
[0065] Step 1: Use OCR technology to convert the work ticket into a processable text format, and apply text cleaning technology to remove phonetic characters and irrelevant information;
[0066] OCR technology converts text
[0067] Image preprocessing:
[0068] Grayscale processing: Convert the image of the work ticket into a grayscale image to reduce the computational complexity.
[0069] Denoising: Apply denoising techniques such as median filtering and bilateral filtering to remove noise from the image and improve the accuracy of OCR recognition.
[0070] Binarization: Use adaptive threshold or Otsu threshold method to convert grayscale image into binary image to enhance the contrast of text.
[0071] Image Rotation and Correction: Detect and correct tilted text in images to ensure horizontal alignment of text lines.
[0072] OCR recognition:
[0073] Select OCR Engine: Convert the pre-processed images to text using advanced OCR engines such as Tesseract, Google Cloud Vision OCR, Microsoft Azure OCR, etc.
[0074] Text area detection: Use the OCR engine to automatically detect the text area in the image and segment it into different text blocks.
[0075] Character recognition: Perform character recognition on the detected text blocks to generate original text data.
[0076] Text cleaning technology
[0077] Initial cleaning:
[0078] Remove phonetic characters: Use regular expressions or character filtering technology to remove phonetic characters (such as noise, garbled characters, etc.) in the text.
[0079] Remove irrelevant information: Remove common irrelevant information (such as headers, footers, page numbers, etc.) to keep the core content of the work ticket.
[0080] Text Normalization:
[0081] Whitespace processing: remove redundant whitespace (such as consecutive spaces, tabs, newlines, etc.) and standardize text format.
[0082] Case Unification: Convert text to lowercase or uppercase as needed to ensure consistency in text processing.
[0083] Punctuation processing: Clean up or standardize punctuation to ensure the readability and parsability of the text.
[0084] Semantic cleaning:
[0085] Vocabulary standardization: Use a dictionary or predefined rules to standardize synonyms, abbreviations, etc. into a unified expression form.
[0086] Spelling Correction: Apply spell checking tools or algorithms (such as Hunspell, SymSpell, etc.) to automatically correct spelling errors in the text.
[0087] Noise word removal: Remove common noise words (such as "the", "is", "and", etc.) and retain key information.
[0088] Step 2: Use the pre-trained language model to perform named entity recognition, extract key entities, and then use relationship extraction technology to identify the relationship between entities, generate a structured data table, and store the key information of the work ticket;
[0089] The step 2 comprises the following steps:
[0090] Data preprocessing: Use the BERT tokenizer to tokenize the text and convert it into a token sequence.
[0091] Model reasoning: Input the preprocessed data into the BERT model and perform forward propagation to obtain the hidden state of the output:
[0092] H = BERT(I,A,S);
[0093] Where: H represents the output hidden state of the BERT model; I represents the input ID sequence; A represents the attention mask; S represents the segment ID;
[0094] Label prediction: Classify each token through a fully connected layer and Softmax function to obtain the probability that each token belongs to each entity category, and select the label with the highest probability:
[0095] P i =Softmax(W·H i + b);
[0096] Where: P i represents the predicted label probability distribution of the i-th Token; W represents the weight matrix of the classification layer; H i represents the hidden state of the i-th Token; b represents the bias vector of the classification layer;
[0097] Relation extraction:
[0098] After obtaining each entity, construct entity pairs, predict the connection between entity pairs through the relation extraction model, and use a variant of BERT to classify each entity pair and predict the relationship type:
[0099] R ij =Softmax((W r ·[H i ;H j ]+b r );
[0100] Where: R ij represents the predicted probability distribution of the relationship between entity i and entity j; W r Represents the weight matrix of the relationship classification layer; [H i ;H i ] represents the hidden state concatenation vector of entity i and entity j; b r Represents the bias vector of the relation classification layer;
[0101] The results of named entity recognition and relationship extraction are structured to generate a structured data table.
[0102] Step 3: Process the video by framing, extract key frames, and apply video compression and enhancement technology to ensure video quality and processing efficiency;
[0103] The step 3 comprises the following steps:
[0104] Video frame processing: Use OpenCV to read video files and obtain basic properties of the video; read the video frame by frame and save each frame as an independent image file;
[0105] Key frame extraction: Convert each frame to an appropriate color space, and then calculate the color histogram. Using the HSV color space, the histogram calculation formula is as follows:
[0106] H i ={h i,j};
[0107] Where: H i represents the color histogram of the i-th frame; h i,j Represents the value of the jth bin in the histogram, indicating the number of pixels contained in the bin;
[0108] h i,j =∑ x,y δ(B(I i (x,y))-j);
[0109] Where: I i (x, y) represents the pixel value at position (x, y) of the i-th frame; δ represents the Dirac delta function, which outputs 1 when the input is 0, otherwise it outputs 0; B represents the assignment of pixel values to the corresponding bins;
[0110] In HSV space, the calculation of histogram is divided into three histograms through (H, S, V):
[0111]
[0112] Where: Respectively represent the histogram of the i-th frame in H, S, and V channels; Respectively represent the value of the jth bin in the corresponding channel;
[0113] The magnitude of change between frames is measured by calculating the difference between the color histograms of adjacent frames:
[0114]
[0115] Where: D i,i+1 represents the histogram difference between the i-th frame and the i+1-th frame; B represents the number of bins in the histogram;
[0116] Set a threshold. If the histogram difference exceeds the threshold, the frame is considered a key frame.
[0117] Video compression: Use FFmpeg to compress the video;
[0118] Video enhancement: Gaussian filtering is used to reduce noise in the video;
[0119] Contrast enhancement: Contrast enhancement can be achieved through histogram equalization. The purpose of histogram equalization is to redistribute pixel values so that the contrast of the image is enhanced. The specific formula is as follows:
[0120] Compute the cumulative distribution function:
[0121]
[0122] Where: CDF(v) represents the cumulative distribution function of pixel value v; h(u) represents the frequency of pixel value u; N represents the total number of pixels in the image;
[0123] Calculate the equalized pixel value:
[0124]
[0125] Where: Indicates rounding down; I′ i (x, y) represents the pixel value at position (x, y) of the i-th frame after equalization; i (x, y) represents the pixel value at position (x, y) of the i-th frame before equalization; L represents the maximum value of the pixel value (for example, for an 8-bit image, L = 256);
[0126] Through the above steps, histogram equalization can effectively enhance the image contrast, making the image details more obvious and recognizable.
[0127] Step 4: Use the target detection model formed by convolutional neural network and temporal convolutional network to perform target detection, identify people, equipment and important scenes in the video, and analyze the behavior to obtain the video detection results;
[0128] The step 4 comprises the following steps:
[0129] Data preprocessing: extract frames from the enhanced video data in step 3 at a fixed frame rate, scale and distort each extracted frame, and cache consecutive frames in chronological order to form a frame sequence;
[0130] Object detection module: The input is a single-frame image after data preprocessing; through multiple layers of convolution and residual connection, deep features are extracted, feature maps are output, and feature maps of different scales are extracted. They are fused through a path aggregation network to output multi-scale feature maps; object classification and positioning are performed at each scale of the multi-scale feature map, and multi-scale detection results are output, including bounding box coordinates and category probabilities, based on which people, equipment and important scenes in the video are obtained;
[0131] Action recognition module: The input is the output result of the target detection module; features are extracted from the target detection results, and the features of all detection boxes are combined into a feature vector to form a feature sequence; the feature sequence is then processed through multiple one-dimensional convolutional layers, and batch normalization and ReLU activation functions are connected after each convolution layer to output convolution features. Pooling is performed after each convolution layer to reduce the sequence length and output pooled features; the features after the last layer of convolution and pooling are flattened, and the probability distribution of the action category is obtained through a fully connected layer.
[0132] The loss function of the target detection model is as follows:
[0133]
[0134] Where: Represents the positioning loss, which measures the difference between the predicted box and the real box, using CloU; represents the confidence loss; represents the classification loss; λ loc represents the weight of the positioning loss; conf represents the weight of execution loss; cls Represents the weight of classification loss; α det Represents the weight of target detection loss; α act represents the weight of action recognition loss; N represents the number of samples; C represents the number of action categories; y i,c Represents the cth class true label of the i-th sample; represents the predicted probability of the cth class of the ith sample.
[0135] The specific structure is as follows:
[0136] Object Detection Module
[0137] Input layer: input a single frame image after preprocessing.
[0138] Convolutional layers and residual connections:
[0139] Multiple convolutional layers extract deep features of the image.
[0140] Use residual connections (ResNet) to alleviate the gradient vanishing problem of deep networks.
[0141] Feature map generation:
[0142] After outputting the feature map, feature maps of different scales are extracted (such as FPN structure).
[0143] The path aggregation network (PANet) is used for feature fusion to generate multi-scale feature maps.
[0144] Target classification and positioning:
[0145] Object detection is performed on feature maps at each scale using bounding box prediction and classification.
[0146] Output multi-scale detection results, including bounding box coordinates and category probabilities.
[0147] Detection results: Identify people, equipment, and important scenes in the video.
[0148] Action recognition module:
[0149] Feature sequence extraction:
[0150] Extract the features of each detection box from the target detection results.
[0151] The features of all detection boxes are combined into a feature vector to form a feature sequence.
[0152] One-dimensional convolutional layer:
[0153] Multiple one-dimensional convolutional layers perform temporal processing on the feature sequence, and each layer is followed by batch normalization and ReLU activation function.
[0154] Pooling is performed after each convolutional layer to reduce the length of the feature sequence.
[0155] Fully connected layer:
[0156] Flatten the features after the last layer of convolution and pooling and input them into the fully connected layer.
[0157] Outputs the probability distribution of action categories.
[0158] In this embodiment, the target detection module uses multiple convolutional layers and residual connections, which can effectively extract deep feature information of the image and improve the accuracy and robustness of detection.
[0159] The feature maps of different scales are fused through the Path Aggregation Network (PANet), which enhances the model's ability to handle multi-scale objects and can more accurately detect objects of different sizes.
[0160] Target classification and localization at each scale of the multi-scale feature map can better cope with the changes of targets at different scales in the image and improve the detection accuracy and recall rate.
[0161] The action recognition module uses a one-dimensional convolutional layer to process the feature sequence, combined with batch normalization and ReLU activation function, to efficiently extract temporal features. This method not only retains the temporal information, but also reduces the computational complexity.
[0162] The pooling operation reduces the sequence length, reduces the dimension of the feature sequence, reduces the amount of calculation, and retains the key features, thus improving the efficiency of the model.
[0163] The loss function of the target detection model comprehensively considers the positioning loss, confidence loss and classification loss, and balances them through weights to ensure the performance of the model on different tasks.
[0164] The action recognition loss uses the cross entropy loss function, which can effectively improve the accuracy of action classification.
[0165] By adjusting the weights of the losses of target detection and action recognition, the importance of both can be adjusted according to actual needs, with high flexibility.
[0166] Step 5: Match the structured information of the work ticket type with the detection results in the video. By establishing a timeline, compare the task arrangement in the work ticket with the actual operation in the video, and perform compliance verification and risk assessment based on the comparison.
[0167] Timeline Creation
[0168] First, we need to establish a unified timeline for the video and the work ticket. Assume that the start time of the video is t0, the total duration of the video is T, and the timeline is: t0, t1, t2, …, t n , where t i Represents the timestamp of the i-th frame of the video;
[0169] For the tasks in the work ticket, we need to extract the start time and end time of the task. Assume that the task in the work ticket is Task i , whose start time is End time is
[0170] Matching work ticket tasks and video content:
[0171] In order to match the tasks in the work ticket with the actual operations in the video, we need to time-label the detection results in the video. Assume that we have obtained the event Event in the video through the target detection model j and its event segment
[0172] The key to matching is to find the work ticket task i and video events j In the overlapping part on the timeline, determine whether the task is completed on time;
[0173] In order to accurately match, further comparison is required based on the specific content in the detection results (such as people, equipment, and scenes). For example, the task description in the work ticket may involve the operation of a certain device, and we need to ensure that the device operation detected in the video is consistent with the task description.
[0174] Assume that the work ticket task Task i Contains the following key information:
[0175] Task Description: Operate Equipment A
[0176] Related Personnel: Person X
[0177] Video Event j The following test results are included:
[0178] Detected device: Device A
[0179] Person Detected: Person X
[0180] By comparing the key information of the two, we can further confirm the task. i Is it in the video event Event j was correctly implemented;
[0181] The above matching results are stored in a comparison table for subsequent compliance verification and risk assessment. The structure of the comparison table can be as follows:
[0182]
[0183] Compliance Verification and Risk Assessment
[0184] Comparison tables allow for compliance verification and risk assessment. For example:
[0185] Check whether the task is completed within the specified time.
[0186] Check whether the task execution meets the description of the work ticket.
[0187] Identify potential risk points, such as tasks not being completed on time or operational errors.
[0188] The present invention uses OCR technology to convert the work ticket into a text format, and applies text cleaning technology to remove irrelevant information to ensure the integrity and accuracy of the data. Then, the pre-trained language model is used to perform named entity recognition and relationship extraction, generate a structured data table, and extract the key information of the work ticket. Then, the video is frame-processed and compressed to ensure video quality and processing efficiency, and the target detection model of the convolutional neural network and the temporal convolutional network is applied to identify the characters, equipment and important scenes in the video, analyze the behavior and obtain the video detection results. By matching the structured information in the work ticket with the video detection results, establishing a timeline, and comparing the task arrangement in the work ticket with the actual operation in the video, dynamic compliance verification and risk assessment are achieved. This method not only improves the efficiency and accuracy of information processing, but also realizes real-time monitoring and dynamic analysis, timely discovers and handles on-site problems, and significantly improves the safety and standardization of operations.
[0189] The above are only preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present application should be included in the protection scope of the present application.
Claims
1. A method for extracting work ticket information and linking work video analysis, characterized in that: The following steps are involved: Step 1: Use OCR technology to convert the work ticket into a processable text format, and apply text cleaning technology to remove phonetic characters and irrelevant information; Step 2: Use the pre-trained language model to perform named entity recognition, extract key entities, and then use relationship extraction technology to identify the relationship between entities, generate a structured data table, and store the key information of the work ticket; Step 3: Process the video by framing, extract key frames, and apply video compression and enhancement technology to ensure video quality and processing efficiency; Step 4: Use the target detection model formed by convolutional neural network and temporal convolutional network to perform target detection, identify people, equipment and important scenes in the video, and analyze the behavior to obtain the video detection results; Step 5: Match the structured information of the work ticket type with the detection results in the video. By establishing a timeline, compare the task arrangement in the work ticket with the actual operation in the video, and perform compliance verification and risk assessment based on the comparison.
2. The method for extracting work ticket information and analyzing operation video linkage according to claim 1 is characterized in that: The step 2 comprises the following steps: Data preprocessing: Use the BERT tokenizer to tokenize the text and convert it into a token sequence. Model reasoning: Input the preprocessed data into the BERT model and perform forward propagation to obtain the hidden state of the output: H = BERT(I,A,S); Where: H represents the output hidden state of the BERT model; I represents the input ID sequence; A represents the attention mask; S represents the segment ID; Label prediction: Classify each token through a fully connected layer and Softmax function to obtain the probability that each token belongs to each entity category, and select the label with the highest probability: P i =softmax(W·H i +b); Where: P i represents the predicted label probability distribution of the i-th Token; W represents the weight matrix of the classification layer; H i represents the hidden state of the i-th Token; b represents the bias vector of the classification layer; Relation extraction: After obtaining each entity, construct entity pairs, predict the connection between entity pairs through the relation extraction model, and use a variant of BERT to classify each entity pair and predict the relationship type: R ij =Softmax((W r ·[H i ;H j ]+b r ); Where: R ij represents the predicted probability distribution of the relationship between entity i and entity j; W r Represents the weight matrix of the relationship classification layer; [H i ;H j ] represents the hidden state concatenation vector of entity i and entity j; b r Represents the bias vector of the relation classification layer; The results of named entity recognition and relationship extraction are structured to generate a structured data table.
3. The method for extracting work ticket information and analyzing operation video linkage according to claim 1 is characterized in that: The step 3 comprises the following steps: Video frame processing: Use OpenCV to read video files and obtain basic properties of the video; read the video frame by frame and save each frame as an independent image file; Key frame extraction: Convert each frame to an appropriate color space, and then calculate the color histogram. Using the HSV color space, the histogram calculation formula is as follows: H i ={h i,j }; Where: H i represents the color histogram of the i-th frame; h i,j Represents the value of the jth bin in the histogram, indicating the number of pixels contained in the bin; h i,j =∑ x,y δ(B(I i (x,y))-j); Where: I i (x, y) represents the pixel value at position (x, y) of the i-th frame; δ represents the Dirac delta function, which outputs 1 when the input is 0, otherwise it outputs 0; B represents assigning pixel values to corresponding bins; In HSV space, the calculation of histogram is divided into three histograms through (H, S, V): Where: Respectively represent the histogram of the i-th frame in H, S, and V channels; Respectively represent the value of the jth bin in the corresponding channel; The magnitude of change between frames is measured by calculating the difference between the color histograms of adjacent frames: Where: D i,i+1 represents the histogram difference between the i-th frame and the i+1-th frame; B represents the number of bins in the histogram; Set a threshold. If the histogram difference exceeds the threshold, the frame is considered a key frame. Video compression: Use FFmpeg to compress the video; Video enhancement: Gaussian filtering is used to reduce noise in the video; Contrast Enhancement: Enhances the contrast of the video using histogram equalization.
4. The method for extracting work ticket information and analyzing operation video linkage according to claim 1 is characterized in that: The step 4 comprises the following steps: Data preprocessing: extract frames from the enhanced video data in step 3 at a fixed frame rate, scale and distort each extracted frame, and cache consecutive frames in chronological order to form a frame sequence; Object detection module: The input is a single-frame image after data preprocessing; through multiple layers of convolution and residual connection, deep features are extracted, feature maps are output, and feature maps of different scales are extracted. They are fused through a path aggregation network to output multi-scale feature maps; object classification and positioning are performed at each scale of the multi-scale feature map, and multi-scale detection results are output, including bounding box coordinates and category probabilities, based on which people, equipment and important scenes in the video are obtained; Action recognition module: The input is the output result of the target detection module; features are extracted from the target detection results, and the features of all detection boxes are combined into a feature vector to form a feature sequence; the feature sequence is then processed through multiple one-dimensional convolutional layers, and batch normalization and ReLU activation functions are connected after each convolution layer to output convolution features. Pooling is performed after each convolution layer to reduce the sequence length and output pooled features; the features after the last layer of convolution and pooling are flattened, and the probability distribution of the action category is obtained through a fully connected layer.
5. The method for extracting work ticket information and analyzing operation video linkage according to claim 4 is characterized in that: The loss function of the target detection model is as follows: Where: Represents the positioning loss, which measures the difference between the predicted box and the real box, using CloU; represents the confidence loss; represents the classification loss; λ loc represents the weight of the positioning loss; conf represents the weight of execution loss; cls Represents the weight of classification loss; α det Represents the weight of target detection loss; α act represents the weight of action recognition loss; N represents the number of samples; C represents the number of action categories; y i,c Represents the cth class true label of the i-th sample; represents the predicted probability of the cth class of the ith sample.
Citation Information
Cited By
Display mode self-adaptive switching method based on continuous frame OCR (Optical Character Recognition) analysis
CN121092259A
Substation pressing plate state checking method, system, equipment and medium
CN121746305A