Video image data cleaning method and system for urban rail transit projects

By identifying key features through a multimodal large language model and a group voting mechanism, combined with data de-redundancy and augmentation processes, the high redundancy and imbalance problems of video image data in urban rail transit project construction are solved, and efficient data cleaning and intelligent analysis support are achieved.

CN120182636BActive Publication Date: 2025-10-17BEIJING URBAN CONSTRUCTION DESIGN & DEVELOPMENT GROUP CO LIMITED
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510665022.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-10-17
Estimated Expiration
2045-05-22

AI Technical Summary

Technical Problem

During the construction of urban rail transit projects, video image data suffers from high redundancy and imbalance. Existing technologies make it difficult to effectively remove redundant data, balance data distribution, and identify key features, which affects the accuracy and efficiency of intelligent analysis.

Method used

A multimodal large language model combined with a group voting mechanism is used to calculate the structural similarity index of adjacent video frames, identify key features, and build a data de-redundancy and augmentation process. The trained discriminator and autoencoder are used for data adaptation processing to achieve automatic cleaning of video image data.

Benefits of technology

Effectively remove redundant data, balance data distribution, accurately identify key features, improve data processing efficiency and storage optimization, enhance system generalization capabilities, and provide reliable data support for intelligent monitoring and fault warning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182636B_ABST
    Figure CN120182636B_ABST
Patent Text Reader

Abstract

The application discloses a kind of video image data cleaning methods and systems for urban rail transit engineering, method includes: calculating video adjacent frame structure similarity index SSIM, video block is divided according to threshold value and index set is constructed;Select no less than one multi-modal large language model, score is obtained by feature pre-processing, prompt word construction, model input calculation, the grouping voting mechanism is used to judge feature, obtain feature matrix and retain background image;Definition feature block, construct matrix to obtain sampling set, calculate sampling interval according to training sample requirement, realize redundancy removal by interval sampling or supplementary synthesis data;The amount of use of model synthesis data is used to classify and augment, adapt by discriminator and autoencoder, fusion adaptation data and original data complete augmentation, achieve video image data automatic cleaning.Reduce data storage and management cost, so that data resources are more reasonably configured.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of video image data processing, and more particularly relates to a video image data cleaning method and system for urban rail transit engineering. BACKGROUND

[0002] In the process of urban rail transit engineering construction, a large number of video monitoring devices are deployed on site to ensure construction safety and monitor project progress. A single construction site may have up to thousands of video monitoring channels, generating a huge amount of data daily. However, through in-depth analysis, it is found that these video image data have significant high redundancy and data imbalance problems. On the one hand, a large amount of video content is repetitive, such as in the case of night construction scenes, the pictures remain unchanged for several hours. Such data not only occupies a large amount of storage resources, but also increases the burden of subsequent data processing and analysis. On the other hand, the frequency of different construction events in video data varies greatly, such as normal construction state video data accounting for a large proportion, while the data proportion of key events such as equipment failure and safety hazards is relatively low. This data imbalance seriously affects the accuracy and effectiveness of subsequent intelligent analysis.

[0003] Currently, existing technologies mainly rely on traditional methods such as rule-based detection and manual annotation when processing urban rail transit engineering construction video image data. Rule-based detection requires pre-setting complex rule sets, which are difficult to cover all situations in the face of complex and variable construction scenes, and are prone to missed detection or false detection. Manual annotation completely relies on manual frame-by-frame analysis and labeling of video data, which not only consumes a lot of manpower and time cost, but also is greatly affected by subjective factors, with uneven labeling quality, making it difficult to meet the demand of mass data processing. In addition, traditional methods cannot adapt to the characteristics of key features in video data, such as dynamic changes in scenes, making it difficult to accurately extract key information such as construction equipment, personnel behavior, and environmental status, resulting in low data processing efficiency and inability to provide strong support for intelligent monitoring and management of engineering construction.

[0004] In recent years, with the development of Multimodal Large Language Models (MLLM), its powerful semantic understanding and cross-modal information processing capabilities have brought new possibilities for video data processing. However, there is currently no mature solution to effectively apply multimodal large language models to urban rail transit engineering construction video image data processing, and it is difficult to take advantage of its strengths to solve problems such as high redundancy of video data, data imbalance, and difficulty in identifying key features.

[0005] Therefore, there is an urgent need for an innovative technical solution that can fully leverage the technical advantages of multi-modal large language models to automatically clean urban rail transit construction video image data, effectively remove redundant data, balance data distribution, accurately identify key features, and thus provide a high-quality data foundation for subsequent video deep analysis and intelligent decision-making. SUMMARY

[0006] The purpose of the present application is to solve the problem of high redundancy and imbalance of urban rail transit construction video image data. By introducing a multi-modal large language model, a grouping voting mechanism is used to identify key features, and a data de-redundancy and augmentation process is constructed to accurately identify key information such as construction equipment and personnel, remove redundant data, and expand the amount of high-quality data, providing a high-quality data foundation for video deep analysis, intelligent monitoring, and other applications, and promoting the development of industry digitization.

[0007] In view of the above defects or improvement needs of the prior art, the present application provides a video image data cleaning method for urban rail transit engineering, comprising:

[0008] S1. Calculate the structural similarity index SSIM of adjacent frames of the video, and divide the video into blocks to construct a block index set;

[0009] S2. Select no less than one multi-modal large language model and use a grouping voting mechanism. First, preprocess the features to be identified and create an image set containing feature labels. Then, construct targeted prompt words. Then, input the images and features into the prompt words into the model to calculate the feature scores. Through grouping voting, determine whether a specific feature exists in the video based on the results of multiple models, and finally obtain a feature matrix, and retain the background image;

[0010] S3. Define the video block containing the specific feature as a feature block, arrange all video blocks to construct a matrix, calculate the feature block matrix, and then obtain a sampling set. According to the training sample requirements, calculate the sampling interval of each element block in the sampling set. When the sampling interval is greater than or equal to 1, sample in the data block according to the sampling interval. When the sampling interval is less than 1, supplement the synthesized data. Thus, data de-redundancy is achieved;

[0011] S4. Statistically determine the categories and quantities that need to be augmented, and use a multi-modal large language model to synthesize data. By training a discriminator and an autoencoder, the synthesized images are adapted, the adapted synthesized images are combined with the original data, and data augmentation is completed, finally achieving automatic cleaning of urban rail transit construction video image data.

[0012] Further, the calculation method of the structural similarity index SSIM in S1 is:

[0013] For a video , the first frame of the video is obtained and marked as x, the next frame of the video is obtained and marked as y, and the structural similarity index SSIM between x and y is calculated according to the following formula:

[0014] ,

[0015] wherein, and are the pixel mean values of images x and y, and are the pixel variances of images x and y, is the covariance of images x and y, and are stable constants.

[0016] Further, the specific method for constructing the block index set of the video in S1 is as follows:

[0017] Let the video set be , be a video of urban rail transit engineering construction, denote the total number of videos in the video set video;

[0018] According to the actual situation of the video, set the threshold value ; for the video , there is a corresponding frame sequence , denote the total number of frames in the frame sequence, and initially , the following steps are taken:

[0019] Let , calculate ;

[0020] When , then , continue to calculate ;

[0021] When , record the position of the frame , , and let , , continue to calculate ;

[0022] When , the process ends;

[0023] Finally, the index frame position set , is obtained, wherein the frame position recorded in the above process is , and the total number of index frame positions is denoted by

[0024] By collection The elements in are the segmentation points, which divide the video into m+1 blocks. All blocks constitute a block set. , record the first frame of the video as a block Index frame ,Location The frames at Index frame ,Location The frames at Index frame , and so on, building a block index collection ;

[0025] will be collected Perform the above operation on all elements in Stitched together to form a larger set ,right Perform the same operation and keep and The relative positions of the elements remain unchanged.

[0026] Furthermore, the method for calculating the feature score in S2 is:

[0027] The identification feature set is , Indicates the total number of features required in subsequent tasks after data cleaning is completed; manually selected Zhang, containing characteristics Frames, put a red box with four sides close to the feature edge into these images, and name these images ; for the set Perform the above operations on all elements in to obtain the image set:

[0028] ,in , For each feature The number of manually selected images;

[0029] Construct prompt words: "In the urban rail transit project construction scene, the image <figure1>The middle red box is <f>Please judge the image strictly <figure2>whether to include <f>, just answer with "yes" or "no" without explanation"

[0030] Index the block Images in Replace the prompt word <figure2>, the feature set is in the feature set replaced <f>with the elements in the image set <figure1>obtained group prompt words;

[0031] Let the number of MLMs used be , then input the above prompt words into MLLMs to obtain result sets, let be the total set of the result sets, then:

[0032] ,

[0033] wherein, for each , MLLM index number:

[0034] ,

[0035] In the formula, denotes the result set of the t-th MLLM, denotes the n groups of results obtained by inputting the n groups of prompt words into the MLLM with the corresponding index number t, if the element in the above set is "yes", the element is replaced by 1, if the element is "no", the element is replaced by 0, if the element in the set does not belong to {"yes", "no"}, the element is removed.

[0036] Further, the grouping voting mechanism in S2 is:

[0037] Set a weight for each element in to obtain a weight set , the weight is set according to the corresponding MLLM model recognition effect, the value range of the weight is 0~1, and the corresponding MLLM model score is calculated:

[0038] ,

[0039] In the formula, is element-level multiplication, that is, the elements at the same position in the two sets are multiplied and then added; is a proportionality coefficient, which is used to make the calculation result fall within the interval [0, 1], and the calculation formula is:

[0040] ;

[0041] According to the image analysis sensitivity requirement and video image analysis speed, set the discrimination value ;

[0042] Let A set of MLLM model scores is denoted as , then:

[0043] ,

[0044] Let be the proportion of the required votes that voted for the result of 1, where ; is the final labeling result; then:

[0045] ,

[0046] where, denotes the number of MLLMs, is an indicator function that is when ; and is when ;

[0047] That is, when , it is considered that the video block contains the feature , and is labeled as 1, otherwise, it is labeled as 0.

[0048] Further, the manner of obtaining the feature matrix in S2 is specifically:

[0049] For the index frame of the video block , all elements in the feature set are subjected to the above-mentioned calculation of feature scores and grouping voting mechanism to obtain the feature sequence of the index frame :

[0050] ,

[0051] The above-mentioned operation is performed on all elements in the block set and the block index set to obtain the corresponding feature sequence set , denotes the total number of elements in the index frame position set PO, that is, the total number of index frame positions, and the elements in are arranged together by column to obtain the feature matrix :

[0052] ,

[0053] In the feature matrix , there are (m+1) rows representing video blocks, and there are columns representing features of interest.

[0054] Further, the construction of the sample set S3 is as follows:

[0055] Define a video block containing features , called a feature block; Arrange all elements in the block set together to form a matrix:

[0056] , which has columns and rows, denotes the length of the block with the largest number of frames among all blocks in CM; when is smaller than , the missing position is filled with 0, indicating that there is no frame at that position; Calculate the feature block matrix: , where

[0057] is matrix multiplication, + denotes concatenating all frames of two blocks together, is zero; the feature block matrix has rows and columns; Construct the sample set =

[0058] , where denotes the th row in , that is, the sum of all row elements of the feature block matrix in the corresponding column to obtain a single-column matrix with 1 row and columns, and a set is used to represent all elements in the matrix, denotes concatenating / adding FCM in the corresponding column to form a block, and 0 represents that there is no frame at that position. Concatenating / adding with any element equals itself. Further, the specific method for synthesizing data using a multi-modal large language model in S4 is as follows:

[0059] Statistical categories and quantities that need to be augmented: the amount of data that needs to be augmented is represented by the set , and the elements

[0060] in the set are calculated using the following formula: ,

[0061] ,

[0062] where denotes the number of samples needed to process each type of feature P in the subsequent task after cleaning.​ denotes an element of the sample set FS total number of frames of the block, 0 means that no data augmentation is needed for this class, the set The maximum value of the elements in is denoted as maxAUV;

[0063] Construct the prompt word: "Replace element <f>Precise overlay to base context map <figure1>When operating, the following specifications shall be followed: (1) use the layered editing tool to place the background image in an independent group track and lock the attributes, ensuring that the coordinates and transparency parameters thereof cannot be modified; (2) new elements shall be superimposed through an independent image track; and (3) element combinations shall conform to the urban rail transit engineering construction site scene;

[0064] With the above prompt words, the categories corresponding to the elements greater than zero in the set are augmented with data, and maxAUV images are synthesized for each category.

[0065] Further, the specific method for adapting the synthesized image in S4 is:

[0066] Training the discriminator: the synthesized data is taken as a negative sample set, and the label is 0; the sample set is taken as a positive sample set, and the label is 1, and the discriminator is trained; binary cross entropy is used as the loss function:

[0067] ,

[0068] In the formula, represents the label of the sample, represents the output result of the discriminator, represents the number of training samples;

[0069] Training the autoencoder: all images in the block index set are taken as training samples, and the autoencoder is pre-trained using the reconstruction loss; next, the autoencoder is trained for image adaptation, using the reconstruction loss and the BCE loss; for a single image, the reconstruction loss includes and , and the calculation formula is as follows:

[0070] ,

[0071] ,

[0072] In the formula, represents a pixel point in the input sample, represents a pixel point in the reconstructed sample, represents the position of the pixel point in the image;

[0073] The final loss function is:

[0074] ,

[0075] In the above formula, is a weight parameter of different loss functions, which is set according to actual experience; when calculating , and respectively put into the trained discriminator by the input image and the reconstructed image respectively;

[0076] using the trained autoencoder, the domain adaptation of the synthesized image is completed;

[0077] all synthesized images are input into the autoencoder for adaptation to obtain augmented data.

[0078] As a second aspect of the present application, a video image data cleaning system for urban rail transit engineering is provided, characterized in that it comprises:

[0079] A video block and index construction unit is configured to calculate the structural similarity index (SSIM) of adjacent frames of a video, divide the video into blocks, and construct a block index set.

[0080] A key feature identification unit is configured to select no less than one multi-modal large language model and adopt a grouping voting mechanism. First, the features to be identified are preprocessed to produce an image set containing feature labels. Then, targeted prompt words are constructed. Then, the images and features are input into the model by substituting the prompt words, and the feature scores are calculated. Through grouping voting, it is determined whether a specific feature exists in the video based on the results of multiple models, and finally a feature matrix is obtained, and the background image is retained.

[0081] An adaptive data de-redundancy unit is configured to define video blocks containing specific features as feature blocks, arrange all video blocks to construct a matrix, calculate the feature block matrix, and then obtain a sampling set. According to the training sample requirements, the sampling interval of each element block in the sampling set is calculated. When the sampling interval is greater than or equal to 1, sampling is performed in the data block according to the sampling interval. When the sampling interval is less than 1, supplementary data is synthesized. Thus, data de-redundancy is achieved.

[0082] A multi-modal synthesis and adaptation data augmentation unit is configured to count the categories and quantities of data that need to be augmented, synthesize data using a multi-modal large language model, and perform adaptation processing on the synthesized images through training of discriminators and autoencoders. The adapted synthesized images are combined with the original data to complete data augmentation, and finally the automatic cleaning of urban rail transit engineering construction video image data is achieved.

[0083] Overall, compared with the prior art, the above technical solutions conceived by the present application can achieve the following beneficial effects:

[0084] 1. A video image data cleaning method for urban rail transit engineering, which introduces multi-modal large language models combined with a grouping voting mechanism to identify key features, effectively improving the comprehensiveness and accuracy of feature identification. Multiple multi-modal large language models are selected, the features to be identified are preprocessed first, a set of feature-labeled images is created, and specific prompt words are constructed and input into the model to calculate feature scores. Finally, the presence or absence of specific features in the video is determined through grouping voting. This technology fully utilizes the semantic understanding advantages of multi-modal large language models, breaks through the limitations of single feature extraction models, and can comprehensively identify multi-modal key features such as construction equipment, personnel, and equipment. The grouping voting mechanism improves the reliability of the identification results, and even in the face of new construction scenarios or complex working conditions, it can maintain a high identification accuracy, providing accurate key information for subsequent data processing.

[0085] 2. A video image data cleaning method for urban rail transit engineering, which realizes efficient utilization and storage optimization of video data by constructing a unique data redundancy reduction process. Define the video block containing specific features as a feature block, arrange all video blocks to construct a matrix, calculate the feature block matrix to obtain a sample set, and then calculate the sampling interval of each element block according to the training sample requirements. The data is sampled or supplemented according to the interval. This process realizes adaptive sampling based on the video block feature matrix, avoiding the problems of storage resource waste and low processing efficiency caused by data redundancy, ensuring efficient utilization of training samples. In the context of large amounts of video data in urban rail transit engineering, redundant data occupies less space, improves data processing speed, reduces data storage and management costs, and makes data resources more reasonably allocated.

[0086] 3. A video image data cleaning method for urban rail transit engineering, which solves the problems of data imbalance and insufficient samples and enhances the system's generalization ability through data augmentation technology based on multi-modal large language models. After counting the categories and quantities that need to be augmented, multi-modal large language models are used to synthesize data, and discriminators and autoencoders are trained to adapt the synthesized images. This technology uses multi-modal large language models to synthesize data that meets the scene specifications, combined with deep neural network training, effectively expanding the amount of high-quality data and alleviating the data imbalance situation. Optimizing data distribution allows the model to learn more rich data features during training, improving the model's ability to adapt to different scenarios and events, providing more comprehensive and reliable data support for intelligent monitoring, fault warning, and other applications in urban rail transit engineering construction, and promoting the development of the industry. BRIEF DESCRIPTION OF DRAWINGS

[0087] Figure 1 A flowchart of a video image data cleaning method for urban rail transit engineering according to an embodiment of the present invention;

[0088] Figure 2 A video block flowchart for an embodiment of the present application;

[0089] Figure 3 A structure diagram of an image adaptation method for an embodiment of the present application;

[0090] Figure 4 A structure diagram of a discriminator for an embodiment of the present application;

[0091] Figure 5 A system unit diagram for an embodiment of the present application. DETAILED DESCRIPTION

[0092] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application. In addition, the technical features involved in each embodiment of the present application described below can be combined with each other as long as they do not conflict with each other.

[0093] Embodiment 1

[0094] Please refer to Figure 1 The present embodiment 1 provides a video image data cleaning method for urban rail transit engineering, comprising:

[0095] S1. Calculate the structural similarity index SSIM of adjacent frames of the video, and divide the video into blocks to construct a block index set;

[0096] S2. Select no less than one multi-modal large language model, and use a grouping voting mechanism; first, pre-process the features to be identified, and make an image set containing feature labels; then, construct a targeted prompt word; then, input the image and the feature into the prompt word into the model to calculate the feature score; through grouping voting, it is judged whether a specific feature exists in the video according to the results of multiple models, and finally a feature matrix is obtained, and the background image is retained;

[0097] S3. Define the video block containing the specific feature as a feature block, arrange all the video blocks to construct a matrix, calculate the feature block matrix, and then obtain a sampling set; according to the training sample requirement, calculate the sampling interval of each element block in the sampling set; when the sampling interval is greater than or equal to 1, sample in the data block according to the sampling interval; when the sampling interval is less than 1, supplement the synthesized data; thereby realizing data redundancy reduction;

[0098] S4. Statistically count the categories and quantities that need to be augmented, and use a multi-modal large language model to synthesize data; through training of a discriminator and an autoencoder, the synthesized image is adapted and processed, the adapted synthesized image is combined with the original data, data augmentation is completed, and finally the automatic cleaning of the video image data of urban rail transit engineering construction is realized.

[0099] This embodiment 1 further expands the above steps.

[0100] (1) Video block

[0101] Let the video set be , a city rail transit project construction video, denote the total number of video sets video;

[0102] Please refer to Figure 2 , for a video , get the first frame of the video and mark it as x, get the next frame of the video and mark it as y, and the structural similarity index SSIM between x and y is calculated as follows:

[0103] ,

[0104] where, and are the pixel means of images x and y, and are the pixel variances of images x and y, is the covariance of images x and y, and are stability constants.

[0105] According to the actual situation of the video, set the threshold value ; for the video there is a corresponding frame sequence , denote the total number of frames in the frame sequence, initially , the following steps are taken:

[0106] Let , calculate ;

[0107] When , then , continue to calculate ;

[0108] When , record the position of the frame , is the recorded serial number, and let , , continue to calculate ;

[0109] When , the process ends;

[0110] Finally, the index frame position set , The frame position recorded in the above process, Indicates the total number of index frame positions.

[0111] The video is divided into m+1 blocks with the elements in the set as the division point, and all blocks form a block set The first frame image of the video is recorded as the index frame of block , the frame at position is recorded as the index frame of block , and the frame at position is recorded as the index frame of block , and so on, to build the block index set ;

[0112] All elements in the set are processed as described above, and each processed set is spliced together to form a larger set , and the same operation is performed on , and the relative position relationship between the elements in and must be maintained during the operation.

[0113] (2) Key feature recognition

[0114] Select multiple multimodal large language models (Multimodal Large Language Models / MLLM), such as Janus-Pro-7B, Qwen2-VL, Llama 3.2-Vision, and use a group voting mechanism to identify key features in images.

[0115] Key feature preprocessing: the identified feature set is , Indicates the total number of features needed in subsequent tasks after data cleaning; for example: is a crane, is an operator, is fire-fighting equipment, etc. Manually select n frames (adjust according to actual conditions. If the MLLM accurately identifies a certain feature in most cases, n can be set to 1) containing the feature , place a red square with four edges close to the feature edge in each of these images, and name these images ; perform the above operation on all elements in the set to obtain the image set:

[0116] , wherein , indicates the above for each feature the number of images manually selected;

[0117] Constructing prompt words: "In the construction scene of urban rail transit engineering, the image <figure1>The middle red box is <f>Please judge the image strictly <figure2>whether to include <f>, only answer with "Yes" or "No", without explanation. The prompt words can be adjusted according to different MLLMs.

[0118] Compute feature scores: Replace the image in the chunk index with the image in the prompt words. <figure2>, the feature set is in the feature set replaced <f>with the elements in the image set <figure1>Afterwards Group prompt words;

[0119] Assume that the number of MLLMs used is Then use the above prompt words to enter MLLM A result set, For this The total set of result sets, then:

[0120] ,

[0121] Among them, for each , Indicates the index number of MLLM:

[0122] ,

[0123] Where, represents the result set of the t-th MLLM, Indicates that n groups of prompt words are input into the corresponding MLLM numbered t, and n groups of results are obtained. If the above If the element in the set is "yes", replace it with 1; if the element is "no", replace it with 0. If the element in the set does not belong to {"yes", "no"}, remove it.

[0124] Group voting mechanism: Set a weight for each element in the to get the weight set , said The weight is set according to the corresponding MLLM model recognition effect. The value range of the weight is 0~1. The corresponding MLLM model score is calculated. :

[0125] ,

[0126] In the formula, * is element-wise multiplication, that is, the elements at the same position in the two sets are multiplied and then added; k is the proportional coefficient, which is used to make the calculation result fall into the interval [0,1]. The calculation formula is: ;

[0127] Set the discrimination value according to the image analysis sensitivity requirements and video image analysis speed ;

[0128] set up The MLLM model score set is ,but:

[0129] ,

[0130] Let be the ratio value of the required votes for the voting result to be 1, where ; be the final labeling result; then:

[0131] ,

[0132] where, denotes the number of MLLMs; is an indicator function, when , ; when , ;

[0133] i.e. when , it is considered that the video block contains the feature , and it is labeled as 1, otherwise it is labeled as 0.

[0134] In some preferred embodiments, when the recognition effect of one MLLM meets the expectation, there is no need to set multiple sets to perform the same operation.

[0135] Obtain the feature matrix: for the index frame of the video block , after the above operation on all elements in the feature set , the feature sequence of the index frame is obtained:

[0136] ,

[0137] After the above operation on all elements in the block set and the block index set , the corresponding feature sequence set is obtained, wherein the total number of elements in the index frame position set PO, i.e. the total number of index frame positions, is denoted by , and the elements in are arranged together by column to obtain the feature matrix :

[0138] ,

[0139] In the feature matrix , there are (m+1) rows representing video blocks, and there are columns representing features of interest.

[0140] Keep the background image: after completing the obtaining of the feature matrix, the block set which does not belong to the feature sequence set is kept.The blocks in a set are put together with a background set to represent that this part of data is used in the data augmentation of domain adaptation.

[0141] (3) Data de-redundancy

[0142] Define a video block containing features , called a feature block with .

[0143] Arrange all elements in the block set together to form a matrix: , which has columns, rows, , which represents the length of the block with the most frames among all blocks in CM; when is less than , the missing position is filled with 0, indicating that there is no frame at that position.

[0144] Calculate the feature block matrix: , where is matrix multiplication, + represents concatenating all frames of two blocks together, is zero; the feature block matrix has rows columns.

[0145] Construct the sampling set = , where represents the th row in , that is, the sum of all row elements of the feature block matrix in the corresponding column to obtain a single-column matrix with 1 row and columns, and a set is used to represent all elements in the matrix, , which represents the concatenation / sum of FCM in the corresponding column into a block, and 0 represents that there is no frame at that position. Concatenation / sum with any element is equal to itself.

[0146] According to the sample demand of each category during training to . Assuming that images are needed for each feature category in features, in general, as many training samples as possible are needed, so g is generally taken as the maximum value of the elements in .

[0147] ​The sampling interval is calculated for each element block: , wherein is the total number of frames of the i-th block, represents the integer part. When is greater than or equal to 1, the data in the data block is sampled according to this interval, and when <1, the data synthesized by the data augmentation method is put into the set in the corresponding position, ensuring that all elements in ≥1. This operation is performed on all elements in .

[0148] By constructing the sampling set , the ability of MLLM is utilized to achieve video deduplication.

[0149] (3) Data augmentation

[0150] Statistics of the categories and amounts that need to be augmented: the amount of data that needs to be augmented is represented by the set , and the elements in the set are calculated using the following formula:

[0151] ,

[0152] , wherein represents the number of samples needed to process each type of feature P in the subsequent task after cleaning up; represents the total number of frames of the element block in the sampling set FS, and 0 indicates that data augmentation is not needed for this type. The maximum value of the elements in the set is recorded as maxAUV.

[0153] Construct a prompt word: "put the element <f>Precise overlay to base context map <figure1>When operating, the following specifications shall be followed: (1) use the layered editing tool to place the background image in an independent group track and lock the attributes, ensuring that the coordinates and transparency parameters cannot be modified; (2) new elements shall be superimposed through an independent image track; and (3) element combinations shall conform to the urban rail transit engineering construction site scenario;

[0154] With the above prompt words, the categories corresponding to the elements greater than zero in the set are augmented with data, and maxAUV images are synthesized for each category.

[0155] Referring to Figure 3 , an image adapter is further constructed, and specifically as follows:

[0156] Training the discriminator: the synthesized data is taken as a negative sample set, and the label is 0; the sample set is taken as a positive sample set, and the label is 1; referring to Figure 4 , the discriminator is trained; binary cross-entropy is used as the loss function:

[0157] ,

[0158] In the formula, represents the label of the sample, represents the output result of the discriminator, represents the number of training samples;

[0159] Training the autoencoder: all images in the block index set are taken as training samples, and the autoencoder is pre-trained using the reconstruction loss; next, the autoencoder is trained for image adaptation, using the reconstruction loss and the BCE loss; for a single image, the reconstruction loss includes and , and the calculation formula is as follows:

[0160] ,

[0161] ,

[0162] In the formula, represents a pixel point in the input sample, represents a pixel point in the reconstructed sample, represents the position of the pixel point in the image;

[0163] The final loss function is:

[0164] ,

[0165] In the above formula, is a weight parameter of different loss functions, which is set according to actual experience; the calculation Time, and respectively put into the trained discriminator by the input image and the reconstructed image respectively;

[0166] Using the trained autoencoder, the domain adaptation of the synthesized image can be performed;

[0167] All synthesized images are input into the autoencoder for adaptation to obtain augmented data.

[0168] Through the above operation, automatic data cleaning is realized.

[0169] Embodiment 2

[0170] Please refer to Figure 5 The embodiment 2 provides a video image data cleaning system for urban rail transit engineering, comprising:

[0171] A video block and index construction unit is configured to calculate the structural similarity index SSIM of adjacent frames of a video, divide the video into blocks, and construct a block index set.

[0172] A key feature identification unit is configured to select no less than one multi-modal large language model and adopt a grouping voting mechanism. The unit first pre-processes the features to be identified, and then constructs an image set containing feature labels. Next, it constructs a targeted prompt word. Then, the image and the feature are input into the model by substituting the prompt word, and the feature score is calculated. Through grouping voting, it is determined whether a specific feature exists in the video according to the results of multiple models, and finally a feature matrix is obtained, and the background image is retained.

[0173] An adaptive data deduplication unit is configured to define a video block containing a specific feature as a feature block, arrange all video blocks to construct a matrix, calculate the feature block matrix, and then obtain a sampling set. According to the training sample requirements, the sampling interval of each element block in the sampling set is calculated. When the sampling interval is greater than or equal to 1, sampling is performed in the data block according to the sampling interval. When the sampling interval is less than 1, supplementary synthesized data is performed. Thus, data deduplication is realized.

[0174] A multi-modal synthesis and adaptation data augmentation unit is configured to count the categories and quantities of data that need to be augmented, and to synthesize data using a multi-modal large language model. Through training of a discriminator and an autoencoder, the synthesized image is adapted and processed. The adapted synthesized image is combined with the original data to complete data augmentation, and finally the automatic cleaning of the video image data of urban rail transit engineering construction is realized.

[0175] Embodiment 3

[0176] The embodiment 3 also provides a computer readable storage medium, and the computer readable storage medium stores a computer program.

[0177] The computer readable storage medium can include a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media capable of storing program codes. For the computer readable storage medium provided in the present application, refer to the above method embodiments, and the present application will not be repeated here.

[0178] Those skilled in the art can easily understand that the above description is only a preferred embodiment of the present application, and is not used to limit the present application, and any modification, equivalent replacement and improvement made within the spirit and principle of the present application shall be included in the protection scope of the present application. < / f> ​ ​​< / f> < / f> < / f> < / f> ​​< / f> < / f> < / f>

Claims

1. A video image data cleaning method for urban rail transit projects, characterized in that: include: S1. Calculate the structural similarity index SSIM of adjacent frames of the video and divide the video into blocks to construct a block index set; S2. Select at least one large multimodal language model and employ a group voting mechanism. Preprocess the features to be identified to create a feature-labeled image set. Then, construct targeted prompt words. Substitute the images and features into the prompt words and input the model to calculate the feature score. Through group voting, the presence of specific features in the video is determined based on the results of multiple models, and the feature matrix is ​​finally obtained, while the background image is retained; S3. Define video blocks containing specific features as feature blocks, arrange all video blocks to construct a matrix, calculate the feature block matrix, and then obtain a sampling set; calculate the sampling interval of each element block in the sampling set according to the training sample requirements; when the sampling interval is greater than or equal to 1, sample the data block according to the sampling interval; When the sampling interval is less than 1, supplementary synthetic data is performed to achieve data redundancy removal; S4. Count the categories and usage that need to be augmented, and use a multimodal large language model to synthesize data. Adapt the synthesized images by training the discriminator and autoencoder, and combine the adapted synthesized images with the original data to complete data augmentation, ultimately achieving automatic cleaning of video image data from urban rail transit project construction.

2. The video image data cleaning method for urban rail transit engineering according to claim 1 is characterized in that: The calculation method of the structural similarity index SSIM in S1 is: For a video , get the first frame of the video and mark it as x, get the next frame of the video and mark it as y, and the structural similarity index SSIM calculation formula between x and y is: , in, and is the pixel mean of image x and y, and is the pixel variance of image x and y, is the covariance of images x and y, and is a stability constant.

3. The video image data cleaning method for urban rail transit engineering according to claim 2 is characterized in that: The specific method of constructing the block index set of the video partition block in S1 is: Set video collection , This is a video about the construction of an urban rail transit project. Indicates the total number of videos in the video collection; Set the threshold according to the actual situation of the video ; For videos There is a corresponding frame sequence , Indicates the total number of frames in the frame sequence, initially , proceed as follows: make ,calculate ; when When , continue to calculate ; when When recording frames Location , is the serial number of the record, and , , continue to calculate ; when When , the process ends; Finally, we get the index frame position set , is the frame position recorded in the above process, Indicates the total number of index frame positions; By collection The elements in are the segmentation points, which divide the video into m+1 blocks. All blocks constitute a block set. , record the first frame of the video as a block Index frame ,Location The frames at Index frame ,Location The frames at Index frame , and so on, building a block index collection ; will be collected Perform the above operation on all elements in Stitched together to form a larger set ,right Perform the same operation and keep and The relative positions of the elements remain unchanged.

4. The video image data cleaning method for urban rail transit engineering according to claim 1 is characterized in that: The method for calculating the feature score in S2 is: The identification feature set is , Indicates the total number of features required in subsequent tasks after data cleaning is completed; manually selected Zhang, containing characteristics Frames, put a red box with four sides close to the feature edge into these images, and name these images ; for the set Perform the above operations on all elements in to obtain the image set: ,in , For each feature The number of manually selected images; Construction prompt words: "In the urban rail transit project construction scene, the image <figure1>The red box is <f>, please strictly judge the image <figure2>Whether to include <f> , just answer with "yes" or "no" without explanation"< / f> < / f> Index the block Images in Replace the prompt word <figure2>, the feature set is in replace <f>, using image collection in Replace the elements in <figure1>Afterwards Group prompt words; < / f> Assume that the number of MLLMs used is Then use the above prompt words to enter MLLM A result set, For this The total set of result sets, then: , Among them, for each , Indicates the index number of MLLM: , Where, represents the result set of the t-th MLLM, Indicates that n groups of prompt words are input into the corresponding MLLM numbered t, and n groups of results are obtained. If the above If the element in the set is "yes", replace it with 1; if the element is "no", replace it with 0. If the element in the set does not belong to {"yes", "no"}, remove it.

5. The video image data cleaning method for urban rail transit engineering according to claim 4 is characterized in that: The group voting mechanism in S2 is: Will Set a weight for each element in the to get the weight set , said The weight is set according to the corresponding MLLM model recognition effect. The value range of the weight is 0~1. The corresponding MLLM model score is calculated. : , In the formula, It is element-wise multiplication, that is, the elements in the same position in the two sets are multiplied separately and then added; is the proportional coefficient, which is used to make the calculation result fall into the interval [0,1]. The calculation formula is: ; Set the discrimination value according to the image analysis sensitivity requirements and video image analysis speed ; set up The MLLM model score set is ,but: , set up is the proportion of votes required to determine the voting result as 1, where ; is the final marking result; then: , in, represents the number of MLLMs; is the characteristic function, when hour, ;when hour, ; That is When the video block Features exist in , mark it as 1, otherwise mark it as 0.

6. The video image data cleaning method for urban rail transit engineering according to claim 5 is characterized in that: The characteristic matrix in S2 is obtained in the following manner: For video blocks Index frame , the feature set After all elements in the above are subjected to the feature score calculation and the group voting mechanism, the index frame is obtained. The characteristic sequence of: , The block collection and block index collections The corresponding feature sequence set of all elements obtained by performing the above operations , Represents the total number of elements in the index frame position set PO, that is, the total number of index frame positions. The elements in are arranged in columns to obtain the feature matrix : , In the feature matrix There are (m+1) rows representing video blocks, total Columns represent the focus Features.

7. The video image data cleaning method for urban rail transit engineering according to claim 6 is characterized in that: The sampling set in S3 is constructed as follows: Definition contains features Video Block , called Feature blocks; The block collection Arrange all the elements together to form a matrix: , the matrix has List, OK, Indicates the length of the block with the largest number of frames among all the blocks in CM; when The length is less than When , the missing position is filled with 0, indicating that there is no frame at that location; Calculate the feature block matrix: , where is matrix multiplication, + Indicates splicing all frames of the two blocks together. is zero; the characteristic block matrix Total OK List; Constructing a sampling set = , where express The Row, that is, the feature block matrix All row elements are added together in the corresponding columns to get a 1-row A single-column matrix with columns and a set representing all elements in the matrix. Indicates that FCM is concatenated / added to form a block on the corresponding column. 0 indicates that there is no frame at that location. Concatenating / adding any element is equal to itself.

8. The video image data cleaning method for urban rail transit engineering according to claim 1 is characterized in that: The specific method of synthesizing data using the multimodal large language model in S4 is: Statistics need to expand the categories and usage: the amount of data that needs to be expanded, use the collection Indicates that the elements in the set Calculated using the following formula: , Where, Indicates the number of samples required to process each type of feature P in the subsequent tasks after cleaning is completed; Represents the elements of the sampling set FS The total number of frames in the block, 0 means that this class does not need data augmentation, and the collection The maximum value of the elements in is recorded as maxAUV; Build prompt word: "Element <f>Accurately superimposed on the basic background image <figure1> , the following specifications should be followed during operation: (1) Use the layered editing tool to place the background image in an independent group track and lock the properties to ensure that its coordinates and transparency parameters cannot be modified; (2) New elements should be superimposed through independent image tracks; (3) The combination of elements must conform to the construction scenario of the urban rail transit project. < / f> Using the above prompt words, in the collection The categories corresponding to the elements greater than zero are augmented, and maxAUV images are synthesized for each category.

9. The video image data cleaning method for urban rail transit engineering according to claim 3 is characterized in that: The specific method of performing adaptation processing on the composite image in S4 is: Training the discriminator: Use the synthetic data as a negative sample set with a label of 0; Sampling set The data is used as a positive sample set with a label of 1 to train the discriminator; Use binary cross entropy as the loss function: , Where, represents the label of the sample, Represents the result of the discriminator output, Indicates the number of training samples; Training the Autoencoder: Block Index Set All images in are used as training samples, and the autoencoder is pre-trained using reconstruction loss; then the autoencoder is trained for image adaptation using reconstruction loss and BCE loss; for a single image, the reconstruction loss includes and , the calculation formula is as follows: , , Where, represents the pixel point in the input sample, Represents the pixel points in the reconstructed sample, Indicates the position of the pixel in the image; The final loss function is: , In the above formula, is the weight parameter of different loss functions, which is set according to actual experience; calculate hour, and The input image and the reconstructed image are respectively put into the trained discriminator; Use the trained autoencoder to complete domain adaptation of the synthetic image; All synthesized images are input into the autoencoder for adaptation to obtain augmented data.

10. A video image data cleaning system for urban rail transit projects, characterized in that: include: A video segmentation and index building unit is used to calculate the structural similarity index (SSIM) of adjacent frames of a video and to build a block index set by dividing the video into blocks; The key feature recognition unit is used to select at least one multimodal large language model using a group voting mechanism. It first pre-processes the features to be recognized and creates a set of images with feature labels. It then constructs targeted prompt words. The images and features are then substituted into the prompt word input model to calculate the feature score. Through group voting, the presence of specific features in the video is determined based on the results of multiple models, and the feature matrix is ​​finally obtained, while the background image is retained; The adaptive data de-redundancy unit is used to define video blocks containing specific features as feature blocks, arrange all video blocks to construct a matrix, calculate the feature block matrix, and then obtain a sampling set; calculate the sampling interval of each element block in the sampling set according to the training sample requirements; when the sampling interval is greater than or equal to 1, sample the data block according to the sampling interval; When the sampling interval is less than 1, supplementary synthetic data is performed to achieve data redundancy removal; The multimodal synthesis and adaptation data augmentation unit is used to count the categories and usage that need to be augmented, and use the multimodal large language model to synthesize data; by training the discriminator and autoencoder, the synthesized image is adapted and combined with the original data to complete the data augmentation, and ultimately realize the automatic cleaning of urban rail transit project construction video image data.

Citation Information

Patent Citations

  • Method and system for cleaning video redundant data

    CN117456149A

  • Data cleaning method

    CN119884611A