Metro protection area illegal operation mechanical identification method based on large model scene understanding
Through large-model scenario understanding technology, combined with multimodal large language model and deep learning algorithm, the identification accuracy and adaptability of the subway protection area monitoring system in complex environments is solved, and efficient and reliable identification and monitoring of illegal operation machinery is achieved, ensuring the safety of subway facilities and the stability of the environment.
Patent Information
- Application Number
- CN202510667223.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-08-15
AI Technical Summary
When identifying illegal operating machinery, the existing subway protection area monitoring system has low identification accuracy, serious missed detection and missed inspections, and it is difficult to adapt to complex and changeable construction environments. It lacks flexibility and adaptability, which increases labor costs and reduces the system's response speed and real-time performance.
Using a method based on large-model scenario understanding, combining multimodal large language model and deep learning algorithm, image features are extracted through convolutional neural networks, and a large language model is used for comprehensive analysis, a target discrimination system is constructed, and environmental information analysis is analyzed in combination with support vector machines to realize automatic identification and monitoring of illegal operation machinery in the subway protection area.
It improves identification accuracy, reduces labor costs, enhances system response capabilities, realizes equipment-environment-personnel three-in-one supervision, adapts to stability and accuracy in different scenarios, reduces missed detection rates and identification time, and improves monitoring coverage and response speed for handling violations.
Smart Images

Figure CN120495994A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of subway intelligent technology, and specifically to a method for identifying machinery operating illegally in a subway protection zone based on large-scale model scene understanding. Background Art
[0002] Identifying and monitoring illegally operating machinery has always been a significant technical challenge in subway protection zone management. Traditional recognition methods primarily rely on single visual recognition algorithms. While effective in certain simple scenarios, they often exhibit low recognition rates and significant false positives and missed detections in the complex and ever-changing subway construction environment. These methods often fail to fully utilize multimodal data, particularly when processing both visual and environmental data simultaneously. Furthermore, traditional methods rely on human oversight, which increases labor costs and reduces the system's responsiveness and real-time performance.
[0003] Existing monitoring systems for subway protection zones face challenges such as low recognition accuracy, high false alarm rates, and difficulty adapting to complex environments. Single visual recognition algorithms often overlook rich environmental information and operational characteristics, limiting the system's recognition capabilities when the environment changes dramatically or construction conditions are complex. In some cases, environmental factors such as surface changes and vegetation damage are not effectively integrated, preventing the system from making timely and accurate judgments. Furthermore, traditional methods lack flexibility and adaptability, making it difficult to maintain stable recognition performance across diverse scenarios.
[0004] To address these issues, technologies combining large-scale scene understanding with large models have become a new trend. By introducing large multimodal language models and deep learning algorithms, it is possible to understand images at a deep level and perform more accurate identification and early warning. By leveraging the contextual understanding capabilities of large language models, combined with classifiers such as support vector machines (SVMs), image data and environmental features can be vectorized to construct a target discrimination system, thereby improving target recognition accuracy. Large-scale pre-trained models can perform logical reasoning within a multi-task learning framework, replacing manual secondary review, reducing subjective judgment, and improving operational efficiency. This approach not only reduces labor costs but also enhances the system's ability to respond to illegal operations within subway protection zones, ensuring the safety of subway facilities and the stability of the surrounding environment. Summary of the Invention
[0005] The purpose of the present invention is to provide a technology for identifying illegal mechanical operations in subway protection zones based on large-scale model scene understanding, which solves the problems of limited recognition rate, false detection and missed detection in the existing technology, and has the advantages of improving recognition accuracy, reducing labor costs, and enhancing system responsiveness. The existing management of subway protection zones faces severe challenges, especially in complex environments, where a single visual recognition algorithm is difficult to cope with. By combining a multimodal large language model and a deep learning algorithm, the present invention can understand images at a deep level and perform more accurate recognition and early warning.
[0006] The present invention provides a method for identifying illegal operating machinery in a subway protection zone based on large-scale model scene understanding, comprising the following steps:
[0007] S1. Use a deep learning algorithm to perform preliminary recognition of images within the subway protection zone. A convolutional neural network model is used to extract image features, determine whether there is suspected construction equipment, and calculate the target recognition confidence level.
[0008] S2. For images with low recognition confidence, they are input into multiple large language models for processing, and comprehensive analysis and judgment are performed using natural language processing technology and logical reasoning capabilities. In the large language models, a target discrimination system is constructed based on the environmental information and work behavior characteristics in the image;
[0009] S3. Analyze environmental information indicators using artificial intelligence algorithms; the environmental information indicators include surface changes, soil exposure, vegetation destruction, and garbage loading indicators;
[0010] S4. Automatically identify and monitor illegal machinery operating within the subway protection zone; characteristics of illegal machinery include detection of construction protection facilities, placement of construction materials, and transport vehicles;
[0011] S5. Automatically identify and monitor illegal operations within the subway protection zone; the target identification system also includes supporting personnel indicators to determine whether there are construction workers wearing safety helmets and engineering uniforms.
[0012] Further preferably, in step S1, image features are extracted through a convolutional neural network model and then classified, and the model parameters are optimized using a cross-entropy loss function. The calculation formula for target recognition confidence is C=p(y|x), where p is the recognition probability, y is the target category, and x is the input image.
[0013] Further preferably, the step S2 specifically includes the following steps:
[0014] S21. For suspected images with a confidence level lower than 50%, they are fed into multiple large language models. Each large language model is responsible for a specific task. The model uses natural language processing technology to parse the image description information and conducts detailed analysis based on specific task requirements.
[0015] S22. In the large language model, we use the environmental information and work behavior characteristics in the image to build a target discrimination system, vectorize the environmental features in the image, and combine it with the support vector machine classifier to perform discriminant analysis to improve the accuracy of target recognition;
[0016] S23. Utilize the logical reasoning capabilities of a large model to conduct comprehensive analysis and judgment, replacing manual secondary review. The logical reasoning is achieved through a large-scale pre-trained model. The large-scale pre-trained model has contextual understanding capabilities and performs reasoning under a multi-task learning framework.
[0017] Further preferably, the step S3 is specifically as follows: constructing a deep learning model, using artificial intelligence algorithms of computer vision and image processing technology to analyze environmental information indicators, and analyzing and identifying illegal operations within the subway protection zone in real time; the environmental information includes surface changes, soil exposure, vegetation destruction and garbage loading indicators.
[0018] Further preferably, the surface change indicators include traces of surface landfill or excavation and surface wet traces; in the deep learning model, by collecting high-definition image data and combining it with geographic information system data, the surface changes are dynamically monitored, and the traces of surface landfill or excavation and surface wet trace change indicators are identified through preprocessing, feature extraction and classification of image data; in the feature extraction, a convolutional neural network is used to extract features from the image, and a long short-term memory network is combined to perform time series analysis to enhance the understanding and prediction capabilities of surface changes; in the identification of traces of surface landfill or excavation and surface wet traces, a multi-level indicator system is constructed to map surface change indicators to potential illegal operations, and by comparing and analyzing the spatiotemporal evolution characteristics of the change indicators, it is determined whether there are activities of illegal operating machinery; geospatial statistical analysis methods are introduced to quantify the spatial distribution pattern and change intensity of surface changes.
[0019] Further preferably, in step S4, the mechanical behavior characteristics include detection of construction work protective facilities, construction material placement, and transport vehicles;
[0020] The detection of construction protection facilities includes determining whether there are fences and warning signs, using the deep learning capabilities of large models to fully understand the subway protection zone scene, collecting and analyzing on-site data in multiple dimensions, applying convolutional neural networks to extract features of the construction environment in the image, distinguishing between construction areas and non-construction areas through image segmentation technology, and on this basis, identifying the presence of fences and warning signs.
[0021] The detection of construction work protective facilities is comprehensively judged through multiple characteristic parameters such as color, shape and location, and further combined with machine learning algorithms, including but not limited to support vector machines or random forest algorithms, to optimize and verify the detection results;
[0022] Verify the standardization of detected fences and warning signs, and provide real-time reminders and reports on non-compliance with standards;
[0023] The calculation model calculates the detection accuracy using the following formula:
[0024] Among them, TP represents the positive example of correctly identifying protective facilities, TN represents the negative example of correctly identifying no protective facilities, FP represents the positive example of incorrectly identifying, and FN represents the positive example of not identifying.
[0025] Further preferably, the detection of the placement of construction materials is based on a target detection algorithm, which analyzes image data of the construction site to identify the type and placement of materials, and compares them with preset safety standards to determine whether there are any violations;
[0026] The detection of transport vehicles uses motion target detection technology combined with vehicle recognition algorithms to track and identify vehicles entering and leaving the subway protection zone in real time, and determines whether the vehicle complies with relevant operating specifications by analyzing the vehicle's movement trajectory and behavioral characteristics.
[0027] Further preferably, in step S5, a deep neural network model is introduced, including but not limited to a convolutional neural network and a long short-term memory network, and is trained in combination with actual scene data of a subway construction site to form a recognition model with high accuracy and high robustness; a multi-layer feature extraction technology is used to parse the target objects in the scene layer by layer, and then refine the features such as the color, shape and position of the construction workers' helmets and engineering clothes to improve the accuracy of recognition; a feature map is extracted by performing a convolution operation on the pixels of the input image, and classification and judgment are performed through a fully connected layer, and finally the supporting indicators of the construction workers are output.
[0028] Further preferably, the detection of the supporting personnel is based on the analysis of the clothing features of the personnel in the video image. By comprehensively applying computer vision technology and deep learning algorithms, the appearance features of the personnel, including but not limited to color, shape and material information, are extracted from the video stream in real time, and compared with the standard supporting personnel clothing features pre-stored in the database. Feature matching algorithms, including but not limited to Euclidean distance or cosine similarity calculation, are used to determine whether the personnel in the current scene are wearing construction clothes.
[0029] Compared with the prior art, the present invention has the following beneficial effects:
[0030] (1) Improved recognition accuracy
[0031] The present invention uses a hybrid model architecture of CNN (such as ResNet50) + LSTM, combines image features with spatiotemporal context information, and improves the accuracy of mechanical recognition. It adopts a multi-level confidence judgment mechanism (direct output / large model review / manual intervention) to perform semantic enhancement analysis on low-confidence samples to reduce the missed detection rate. The combination of environmental feature vectorization + SVM classifier improves the accuracy of surface change recognition.
[0032] (2) Efficiency optimization
[0033] The large-scale model logical reasoning of the present invention replaces manual review work, reducing the time consumption of single recognition; the adaptive learning module shortens the model update cycle and reduces maintenance costs.
[0034] (3) Multi-dimensional regulatory capabilities
[0035] It realizes the three-in-one supervision of "equipment-environment-personnel". Mechanical identification covers 8 major categories of construction equipment, and environmental monitoring includes 4 types of surface indicators (fill and excavation traces / wet traces, etc.); personnel detection supports 5 safety features (safety helmets / engineering uniforms, etc.), and the moving target tracking algorithm improves the accuracy of vehicle trajectory restoration.
[0036] (4) System adaptability and management benefits
[0037] In order to improve the robustness and applicability of the system, the present invention also introduces an adaptive learning module. By continuously updating and optimizing model parameters, the recognition system can adapt to different environmental changes and construction conditions, ensuring its stability and accuracy in various complex scenarios. The overall architecture design of the system is mainly modular, which is easy to expand and maintain. It is suitable for promotion and application in different scenarios of multiple subway protection zones. It supports robust recognition of 12 complex scenarios such as lighting / weather / occlusion, and the dynamic optimization of model parameters shortens the adaptation cycle of new scenarios. Compared with traditional manual inspection methods, the monitoring coverage rate is improved, the labor cost is reduced, and the response speed of handling violations is significantly improved.
[0038] The present invention can provide an efficient and reliable means of identifying and monitoring machinery operating in violation of regulations within a subway protection zone, thereby ensuring the safety of subway facilities and the stability of the environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 is a flow chart of the method of the present invention;
[0040] Figure 2 Schematic diagram of the structure initially identified by deep learning;
[0041] Figure 3 This is the operating logic diagram of the large language model. DETAILED DESCRIPTION
[0042] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0043] The preferred embodiments of the present invention are described in further detail below with reference to the accompanying drawings.
[0044] Example 1:
[0045] A method for identifying illegal machinery in subway protection areas based on large-scale model scene understanding, such as Figure 1 The method of the present invention is shown in the flow chart, comprising the following steps:
[0046] S1. Use a deep learning algorithm to perform preliminary recognition of images within the subway protection zone, extract image features through a convolutional neural network model, determine whether there is suspected construction equipment, and calculate the target recognition confidence.
[0047] In step S1, image features are extracted through a convolutional neural network model and then classified. The model parameters are optimized using the cross-entropy loss function. The target recognition confidence is calculated as C = p(y|x), where p is the recognition probability, y is the target category, and x is the input image.
[0048] S2. For images with low recognition confidence, they are fed into multiple large language models for processing, utilizing natural language processing technology and logical reasoning capabilities for comprehensive analysis and judgment. Within these large language models, a target recognition system is constructed based on the environmental information and work behavior characteristics in the image.
[0049] Step S2 specifically includes the following steps:
[0050] S21. For suspected images with a confidence level lower than 50%, they are input into multiple large language models. Each large language model is responsible for a type of task. The image description information is parsed through natural language processing technology, and detailed analysis is performed based on specific task requirements.
[0051] S22. In the large language model, the environmental information and work behavior characteristics in the image are used to build a target discrimination system, the environmental features in the image are vectorized, and the support vector machine classifier is combined to perform discriminant analysis to improve the accuracy of target recognition.
[0052] S23. Utilize the logical reasoning capabilities of a large model to conduct comprehensive analysis and judgment, replacing manual secondary review. The logical reasoning is achieved through a large-scale pre-trained model. The large-scale pre-trained model has contextual understanding capabilities and performs reasoning under a multi-task learning framework.
[0053] S3. Use artificial intelligence algorithms to analyze environmental information indicators; the environmental information indicators include surface changes, soil exposure, vegetation destruction and garbage loading indicators.
[0054] Step S3 specifically involves constructing a deep learning model and using artificial intelligence algorithms based on computer vision and image processing technologies to analyze environmental information indicators, enabling real-time analysis and identification of illegal operations within subway protection zones. The environmental information includes indicators of surface change, exposed soil, vegetation destruction, and garbage loading. Surface change indicators include traces of landfill or excavation and surface wetness. The deep learning model dynamically monitors surface changes by collecting high-definition image data and combining it with geographic information system data. These traces of landfill or excavation and surface wetness are identified through image data preprocessing, feature extraction, and classification. In feature extraction, a convolutional neural network is used to extract features from the images, combined with a long-short-term memory network for time series analysis, enhancing understanding and prediction of surface changes. In identifying traces of landfill or excavation and surface wetness, a multi-level indicator system is constructed to map surface change indicators to potential illegal operations. By comparing and analyzing the spatiotemporal evolution of these change indicators, the presence of illegal machinery is determined. Geospatial statistical analysis methods are introduced to quantify the spatial distribution patterns and intensity of surface changes.
[0055] S4. Automatically identify and monitor illegal operating machinery within the subway protection zone; the characteristics of illegal operating machinery include detection of construction work protection facilities, construction material placement, and transport vehicles.
[0056] In step S4, the mechanical behavior characteristics include detection of construction work protective facilities, construction material placement, and transport vehicles;
[0057] The detection of construction protection facilities includes determining whether there are fences and warning signs, using the deep learning capabilities of large models to fully understand the subway protection zone scene, collecting and analyzing on-site data in multiple dimensions, applying convolutional neural networks to extract features of the construction environment in the image, distinguishing between construction areas and non-construction areas through image segmentation technology, and on this basis, identifying the presence of fences and warning signs.
[0058] The detection of construction work protective facilities is comprehensively judged through multiple characteristic parameters such as color, shape and location, and further combined with machine learning algorithms, including but not limited to support vector machines or random forest algorithms, to optimize and verify the detection results;
[0059] Verify the standardization of detected fences and warning signs, and provide real-time reminders and reports on non-compliance with standards;
[0060] The calculation model calculates the detection accuracy using the following formula:
[0061] Among them, TP represents the positive example of correctly identifying protective facilities, TN represents the negative example of correctly identifying no protective facilities, FP represents the positive example of incorrectly identifying, and FN represents the positive example of not identifying.
[0062] The detection of construction material placement is based on a target detection algorithm. By analyzing image data from the construction site, the type and placement of materials are identified, and by comparing them with preset safety standards, it is determined whether there are any violations.
[0063] The detection of transport vehicles uses motion target detection technology combined with vehicle recognition algorithms to track and identify vehicles entering and leaving the subway protection area in real time. By analyzing the vehicle's movement trajectory and behavioral characteristics, it is determined whether it complies with relevant operating specifications.
[0064] S5. Automatically identify and monitor illegal operations within the subway protection zone; the target identification system also includes supporting personnel indicators to determine whether there are construction workers wearing safety helmets and engineering uniforms.
[0065] In step S5, a deep neural network model is introduced, including but not limited to a convolutional neural network and a long short-term memory network, and is trained in combination with actual scene data from the subway construction site to form a recognition model with high accuracy and high robustness. Multi-layer feature extraction technology is used to analyze the target objects in the scene layer by layer, and then refine the features such as the color, shape, and position of the construction workers' helmets and engineering uniforms to improve the accuracy of recognition. By performing a convolution operation on the pixels of the input image, a feature map is extracted, and classification and judgment are performed through the fully connected layer, and finally the supporting indicators of the construction workers are output.
[0066] The detection of supporting personnel indicators is based on the analysis of the clothing features of personnel in video images. Through the comprehensive application of computer vision technology and deep learning algorithms, the appearance features of personnel, including but not limited to color, shape and material information, are extracted from the video stream in real time, and compared with the standard supporting personnel clothing features stored in the database in advance. Feature matching algorithms, including but not limited to Euclidean distance or cosine similarity calculation, are used to determine whether the personnel in the current scene are wearing construction clothing.
[0067] Compared with the prior art, the present invention has the following beneficial effects:
[0068] (1) Improved recognition accuracy
[0069] The present invention uses a hybrid model architecture of CNN (such as ResNet50) + LSTM, combines image features with spatiotemporal context information, and improves the accuracy of mechanical recognition. It adopts a multi-level confidence judgment mechanism (direct output / large model review / manual intervention) to perform semantic enhancement analysis on low-confidence samples to reduce the missed detection rate. The combination of environmental feature vectorization + SVM classifier improves the accuracy of surface change recognition.
[0070] (2) Efficiency optimization
[0071] The large-scale model logical reasoning of the present invention replaces manual review work, reducing the time consumption of single recognition; the adaptive learning module shortens the model update cycle and reduces maintenance costs.
[0072] (3) Multi-dimensional regulatory capabilities
[0073] It realizes the three-in-one supervision of "equipment-environment-personnel". Mechanical identification covers 8 major categories of construction equipment, and environmental monitoring includes 4 types of surface indicators (fill and excavation traces / wet traces, etc.); personnel detection supports 5 safety features (safety helmets / engineering uniforms, etc.), and the moving target tracking algorithm improves the accuracy of vehicle trajectory restoration.
[0074] (4) System adaptability and management benefits
[0075] In order to improve the robustness and applicability of the system, the present invention also introduces an adaptive learning module. By continuously updating and optimizing model parameters, the recognition system can adapt to different environmental changes and construction conditions, ensuring its stability and accuracy in various complex scenarios. The overall architecture design of the system is mainly modular, which is easy to expand and maintain. It is suitable for promotion and application in different scenarios of multiple subway protection zones. It supports robust recognition of 12 complex scenarios such as lighting / weather / occlusion, and the dynamic optimization of model parameters shortens the adaptation cycle of new scenarios. Compared with traditional manual inspection methods, the monitoring coverage rate is improved, the labor cost is reduced, and the response speed of handling violations is significantly improved.
[0076] The present invention can provide an efficient and reliable means of identifying and monitoring machinery operating in violation of regulations within a subway protection zone, thereby ensuring the safety of subway facilities and the stability of the environment.
[0077] Example 2:
[0078] This embodiment provides a method for identifying illegal operating machinery in subway protection zones based on large-scale model scene understanding. It uses a deep learning algorithm to perform preliminary identification of images in subway protection zones, determines whether there are suspected construction equipment, and calculates the target recognition confidence. This step extracts image features and classifies them by training a convolutional neural network (CNN) model, and uses a cross-entropy loss function to optimize model parameters. The target recognition confidence is calculated as C=p(y|x), where p is the recognition probability, y is the target category, and x is the input image. For suspected images with a confidence level lower than 50%, they are input into multiple large language models, each of which is responsible for a type of task, and natural language processing technology is used to process the image. The technology is used to analyze the image description information and conduct detailed analysis in combination with specific task requirements; the logical reasoning ability of the large model is used to conduct comprehensive research and judgment, replacing manual secondary review, reducing subjective judgment, and improving work efficiency. Logical reasoning is achieved through a large-scale pre-trained model. The model has the ability to understand the context and can perform reasoning under the multi-task learning framework; in the large language model, the environmental information, work behavior and other features in the image are fully utilized to build a target discrimination system, conduct a comprehensive analysis of the on-site environment, and improve the accuracy of target recognition. This process vectorizes the environmental features in the image and combines classifiers such as support vector machines (SVM) for discriminant analysis to improve recognition accuracy.
[0079] This embodiment utilizes advanced artificial intelligence algorithms and analysis of environmental information indicators to efficiently identify illegal mechanical operations within subway protection zones. Based on a deep learning model, this technology, combined with computer vision and image processing techniques, enables real-time analysis and identification of illegal operations within subway protection zones. The environmental information includes indicators such as surface changes, exposed soil, vegetation destruction, and garbage loading.
[0080] Surface change indicators include traces of landfill or excavation and surface moisture. To effectively identify illegal machinery operating within subway protection zones, this embodiment utilizes a deep learning algorithm to accurately analyze surface changes and identify possible illegal operations. By collecting high-definition image data and integrating it with geographic information system (GIS) data, surface changes are dynamically monitored. Image data preprocessing, feature extraction, and classification are used to identify indicators of change, such as traces of landfill or excavation and surface moisture. A convolutional neural network (CNN) is used to extract features from the images, combined with a long short-term memory (LSTM) network for time series analysis, to enhance understanding and prediction of surface changes. During the identification process, a multi-level indicator system is constructed to map surface change indicators to potential illegal operations. By comparing and analyzing the spatiotemporal evolution of these change indicators, the presence of illegal machinery is determined. Furthermore, by introducing geospatial statistical analysis methods, the spatial distribution patterns and intensity of surface changes are quantified, enabling precise location and timely warning of illegal operations within subway protection zones. Through the above steps, the present invention can provide efficient and reliable means of identifying and monitoring illegal machinery within subway protection zones, ensuring the safety of subway facilities and environmental stability. Soil exposure and vegetation damage are identified using a support vector machine (SVM) classification model, and garbage piles are located and identified using target detection algorithms such as YOLO or Faster R-CNN. The formula for detecting surface changes is:
[0081] Among them, St1 and St2 represent the surface conditions at time t1 and t2 respectively.
[0082] This embodiment realizes the automatic identification and monitoring of various illegal operating machines in the subway protection zone through the combination of deep learning algorithms and computer vision technology. The operating machine characteristics include the detection of construction work protective facilities, construction material placement and transport vehicles. Among them, the detection of construction work protective facilities is achieved by combining the convolutional neural network (CNN) model and image semantic segmentation technology to accurately identify the protective facilities in the protection zone, ensuring the integrity and correctness of the facilities. The detection of construction material placement is based on the target detection algorithm. By analyzing the image data of the construction site, the type and placement of the material are identified, and by comparing with the preset safety standards, it is determined whether there are any violations. The detection of transport vehicles uses moving target detection technology, combined with the vehicle recognition algorithm, to track and identify vehicles entering and leaving the subway protection zone in real time. By analyzing the vehicle's motion trajectory and behavioral characteristics, it is determined whether it complies with the relevant operating specifications. Through deep training of big data analysis and machine learning, the present invention can effectively improve the safety monitoring efficiency of the subway protection zone and reduce the pressure and error rate of human supervision.
[0083] In practical applications, the system first extracts continuous frames from surveillance video and uses a preprocessing algorithm to remove noise and interference. These preprocessed images are then fed into a trained large model for feature extraction and recognition. During feature extraction, a deep network structure based on convolutional and pooling layers is employed to ensure that detailed image information is preserved while extracting features. The recognition component embeds a multi-task learning mechanism within the model, enabling simultaneous detection of different types of violations, improving both accuracy and real-time performance. Through real-time analysis and feedback of detection results, the system can promptly issue alerts and automatically record image and video data of violations for subsequent review and processing.
[0084] To enhance the system's robustness and applicability, the present invention also incorporates an adaptive learning module. By continuously updating and optimizing model parameters, the recognition system can adapt to varying environmental changes and construction conditions, ensuring stability and accuracy in a variety of complex scenarios. The system's overall modular architecture facilitates expansion and maintenance, making it suitable for application in diverse scenarios across multiple subway protection zones.
[0085] The detection of construction protection facilities includes determining whether there are fences and warning signs. In this technology, the deep learning capabilities of the large model are first used to fully understand the subway protection zone scene. The system collects and analyzes the on-site data in multiple dimensions, applies convolutional neural networks (CNN) to extract features of the construction environment in the image, and uses image segmentation technology to distinguish between construction areas and non-construction areas. On this basis, the presence or absence of fences and warning signs is identified. The detection of protective facilities can be comprehensively judged through multiple feature parameters such as color, shape, and position, and further combined with machine learning algorithms such as support vector machines (SVM) or random forests (Random Forest) to optimize and verify the detection results. The system verifies the standardization of the detected fences and warning signs, and provides real-time reminders and reports for non-compliance with the standards. The calculation model calculates the detection accuracy using the following formula:
[0086] TP represents a positive example where protective equipment is correctly identified, TN represents a negative example where no protective equipment is correctly identified, FP represents a positive example where it is incorrectly identified, and FN represents a positive example where it is not identified. This method effectively improves the accuracy of identifying illegal machinery within subway protection zones, ensuring construction safety.
[0087] This embodiment also uses advanced image recognition algorithms and big data analysis technologies to automatically identify and monitor illegal operations within subway protection zones. The target identification system also includes supporting personnel indicators to determine whether construction workers are wearing helmets and construction uniforms. Specifically, this technology incorporates deep neural network models, such as convolutional neural networks (CNNs) and long short-term memory networks (LSTMs), and trains them with real-world scene data from subway construction sites to form a highly accurate and robust recognition model. To improve recognition accuracy, the present invention utilizes multi-layer feature extraction technology to analyze target objects in the scene layer by layer, further refining the identification of features such as the color, shape, and position of construction workers' helmets and construction uniforms. By performing convolution operations on the pixels of the input image, a feature map is extracted, which is then classified and judged through a fully connected layer, ultimately outputting supporting indicators for the construction workers. To ensure real-time and high efficiency, the present invention also utilizes optimized algorithmic processes and hardware acceleration technologies, enabling rapid and accurate identification of illegal operations in complex subway construction environments. In this way, the present invention can provide strong technical support for subway construction safety management, ensuring construction safety and the stability of subway operations. The detection of supporting personnel indicators is based on the analysis of personnel clothing features in video images. By comprehensively applying computer vision technology and deep learning algorithms, the appearance features of personnel, including but not limited to color, shape and material information, are extracted from the video stream in real time. These features are then compared and analyzed with the standard supporting personnel clothing features stored in the database in advance. Feature matching algorithms, including but not limited to Euclidean distance or cosine similarity calculations, are used to determine whether the personnel in the current scene are wearing construction clothing. The use of convolutional neural network (CNN) models further improves the ability to accurately detect personnel clothing features. The model parameters are trained and optimized with a large amount of sample data to enhance the robustness to changes in personnel clothing in complex scenes. The system integrates an anomaly detection mechanism, which performs a probability assessment on the identified clothing features. Once an anomaly is detected, an alarm mechanism is immediately triggered to prompt relevant management personnel to intervene. The entire detection process can be analyzed in time series on multiple frames to avoid false alarms caused by misjudgment of a single frame.
[0088] In this embodiment, to effectively identify illegal machinery operating within subway protection zones, the present invention introduces scene understanding technology based on large-scale pre-trained models and uses deep learning algorithms to accurately analyze surface changes to identify possible illegal operations. By collecting high-definition image data and combining it with geographic information system (GIS) data, dynamic monitoring of surface changes is performed. Convolutional neural networks (CNNs) are used to extract features from the images, and long short-term memory networks (LSTMs) are used for time series analysis to enhance understanding and prediction of surface changes.
[0089] Operational behavior characteristics include the detection of construction protection facilities, construction material placement, and transport vehicles. The detection of construction protection facilities combines a convolutional neural network (CNN) model with image semantic segmentation technology to accurately identify the protective facilities within the protection zone, ensuring the integrity and correctness of the facilities. The detection of construction material placement is based on a target detection algorithm. By analyzing image data from the construction site, the type and placement of the materials are identified, and by comparing them with preset safety standards, it is determined whether any construction violations have occurred. The detection of transport vehicles uses motion target detection technology, combined with a vehicle recognition algorithm, to track and identify vehicles entering and leaving the subway protection zone in real time. By analyzing the vehicle's motion trajectory and behavioral characteristics, it is determined whether the vehicle is engaged in construction operations.
[0090] In this example, through deep training using big data analysis and machine learning, the efficiency of security monitoring in subway protection zones can be effectively improved, reducing the burden of human oversight and the rate of false positives. In practice, the system extracts continuous frames from surveillance video, uses a preprocessing algorithm to remove noise and interference, and then feeds these preprocessed images into a trained large model for feature extraction and recognition.
[0091] To enhance the robustness and applicability of the system, this invention incorporates an adaptive learning module. By continuously updating and optimizing model parameters, the recognition system can adapt to varying environmental changes and construction conditions, ensuring stability and accuracy in a variety of complex scenarios. The system's overall modular architecture facilitates expansion and maintenance, making it suitable for application in diverse scenarios across multiple subway protection zones.
[0092] The technical solution of this embodiment also involves supporting personnel indicators to determine whether there are construction workers wearing helmets and engineering uniforms. Specifically, by introducing deep neural network models such as convolutional neural networks (CNNs) and long short-term memory networks (LSTMs), combined with actual scene data from subway construction sites for training, a recognition model with high accuracy and robustness is formed. To improve recognition accuracy, multi-layer feature extraction technology is used to analyze the target objects in the scene layer by layer, which can be refined to identify features such as the color, shape, and position of the construction workers' helmets and engineering uniforms.
[0093] By performing convolution operations on the pixels of the input image, a feature map is extracted, which is then classified and judged through a fully connected layer, ultimately outputting supporting indicators for construction workers. To ensure real-time and high efficiency, an optimized algorithm process and hardware acceleration technology are used, enabling the rapid and accurate identification of illegal work behaviors in complex subway construction environments. In this way, the present invention can provide strong technical support for subway construction safety management, ensuring construction safety and the stability of subway operations.
[0094] The present invention improves the performance of target recognition by combining deep learning with a large language model. Figure 2 The schematic diagram of the structure of deep learning preliminary recognition is shown, in which the input layer inputs multi-dimensional data into the convolutional neural network (CNN) to extract features. After feature extraction, a large language model is used for comprehensive analysis, such as Figure 3 In this model, the input feature vector is converted into a semantic vector, and the information is filtered and weighted through the self-attention mechanism, thereby improving the accuracy and robustness of recognition.
[0095] In complex environments, environmental change indicators affect recognition results. By monitoring environmental changes in real time, the system can dynamically adjust recognition parameters. Specifically, the system calculates the impact coefficient Ce of environmental changes on recognition results, and its formula is:
[0096] Among them, xi and yi are the characteristic values of the target in different environments, and u and v are their means.
[0097] Furthermore, by analyzing operator behavior patterns, the system can predict potential errors and make adjustments. The large language model reduces recognition differences between different operators through semantic analysis, improving the system's generalization capabilities.
[0098] During operation, the system records and analyzes human experience data to further optimize model parameters and reduce the need for manual review. Ultimately, the timeliness of target recognition is guaranteed and labor costs are significantly reduced.
[0099] Through the above steps, the present invention demonstrates how to achieve efficient target recognition in complex environments and ensure the interpretability and generalization ability of the system.
[0100] The above is a further detailed description of the present invention in conjunction with specific preferred embodiments, and the specific implementation of the present invention should not be considered to be limited to these descriptions. For those skilled in the art of the present invention, without departing from the concept of the present invention, several simple deductions or substitutions can be made, which should be considered to fall within the scope of protection of the present invention.
Claims
1. A method for identifying illegal machinery in subway protection areas based on large-scale model scene understanding, characterized by: The following steps are involved: S1. Use a deep learning algorithm to perform preliminary recognition of images within the subway protection zone. A convolutional neural network model is used to extract image features, determine whether there is suspected construction equipment, and calculate the target recognition confidence level. S2. For images with low recognition confidence, they are input into multiple large language models for processing, and comprehensive analysis and judgment are performed using natural language processing technology and logical reasoning capabilities. In the large language models, a target discrimination system is constructed based on the environmental information and work behavior characteristics in the image; S3. Analyze environmental information indicators using artificial intelligence algorithms; the environmental information indicators include surface changes, soil exposure, vegetation destruction, and garbage loading indicators; S4. Automatically identify and monitor illegal machinery operating within the subway protection zone; characteristics of illegal machinery include detection of construction protection facilities, placement of construction materials, and transport vehicles; S5. Automatically identify and monitor illegal operations within the subway protection zone; the target identification system also includes supporting personnel indicators to determine whether there are construction workers wearing safety helmets and engineering uniforms.
2. The method for identifying illegal operating machinery in subway protection areas based on large model scene understanding according to claim 1 is characterized in that: In step S1, image features are extracted through a convolutional neural network model and then classified. The model parameters are optimized using a cross-entropy loss function. The target recognition confidence is calculated using the formula C=p(y|x), where p is the recognition probability, y is the target category, and x is the input image.
3. The method for identifying illegal operating machinery in subway protection areas based on large-scale model scene understanding according to claim 1 is characterized in that: The step S2 specifically includes the following steps: S21. For suspected images with a confidence level lower than 50%, they are fed into multiple large language models. Each large language model is responsible for a specific task. The model uses natural language processing technology to parse the image description information and conducts detailed analysis based on specific task requirements. S22. In the large language model, we use the environmental information and work behavior characteristics in the image to build a target discrimination system, vectorize the environmental features in the image, and combine it with the support vector machine classifier to perform discriminant analysis to improve the accuracy of target recognition; S23. Utilize the logical reasoning capabilities of a large model to conduct comprehensive analysis and judgment, replacing manual secondary review. The logical reasoning is achieved through a large-scale pre-trained model. The large-scale pre-trained model has contextual understanding capabilities and performs reasoning under a multi-task learning framework.
4. The method for identifying illegal operating machinery in subway protection areas based on large-scale model scene understanding according to claim 3 is characterized in that: The step S3 specifically includes: building a deep learning model, using artificial intelligence algorithms based on computer vision and image processing technology to analyze environmental information indicators, and real-time analysis and identification of illegal operations within the subway protection zone; the environmental information includes surface changes, soil exposure, vegetation destruction, and garbage loading indicators.
5. The method for identifying illegal operating machinery in subway protection areas based on large-scale model scene understanding according to claim 4 is characterized in that: The surface change indicators include traces of surface landfill or excavation and surface wetness; in the deep learning model, by collecting high-definition image data and combining it with geographic information system data, surface changes are dynamically monitored, and through preprocessing, feature extraction and classification of image data, surface landfill or excavation traces and surface wetness change indicators are identified; in the feature extraction, convolutional neural networks are used to extract features from images, and long short-term memory networks are combined to perform time series analysis to enhance the understanding and prediction capabilities of surface changes; in the identification of surface landfill or excavation traces and surface wetness, a multi-level indicator system is constructed to map surface change indicators to potential illegal operations, and by comparing and analyzing the spatiotemporal evolution characteristics of the change indicators, it is determined whether there are activities of illegal operating machinery; geospatial statistical analysis methods are introduced to quantify the spatial distribution pattern and change intensity of surface changes.
6. The method for identifying illegal operating machinery in subway protection areas based on large-scale model scene understanding according to claim 1 is characterized in that: In step S4, the mechanical behavior characteristics include detection of construction work protective facilities, construction material placement, and transport vehicles; The detection of construction protection facilities includes determining whether there are fences and warning signs, using the deep learning capabilities of large models to fully understand the subway protection zone scene, collecting and analyzing on-site data in multiple dimensions, applying convolutional neural networks to extract features of the construction environment in the image, distinguishing between construction areas and non-construction areas through image segmentation technology, and on this basis, identifying the presence of fences and warning signs. The detection of construction work protective facilities is comprehensively judged through multiple characteristic parameters such as color, shape and location, and further combined with machine learning algorithms, including but not limited to support vector machines or random forest algorithms, to optimize and verify the detection results; Verify the standardization of detected fences and warning signs, and provide real-time reminders and reports on non-compliance with standards; The calculation model calculates the detection accuracy using the following formula: Among them, TP represents the positive example of correctly identifying protective facilities, TN represents the negative example of correctly identifying no protective facilities, FP represents the positive example of incorrectly identifying, and FN represents the positive example of not identifying.
7. The method for identifying illegal operating machinery in subway protection areas based on large-scale model scene understanding according to claim 6 is characterized in that: The detection of construction material placement is based on a target detection algorithm that analyzes image data of the construction site to identify the type and placement of materials, and compares them with pre-set safety standards to determine whether there are any violations. The detection of transport vehicles uses motion target detection technology combined with vehicle recognition algorithms to track and identify vehicles entering and leaving the subway protection zone in real time, and determines whether the vehicle complies with relevant operating specifications by analyzing the vehicle's movement trajectory and behavioral characteristics.
8. The method for identifying illegal operating machinery in subway protection areas based on large-scale model scene understanding according to claim 1 is characterized in that: In step S5, a deep neural network model is introduced, including but not limited to a convolutional neural network and a long short-term memory network, and is trained in combination with actual scene data from the subway construction site to form a recognition model with high accuracy and high robustness. Multi-layer feature extraction technology is used to analyze the target objects in the scene layer by layer, and then refine the features such as the color, shape, and position of the construction workers' helmets and engineering uniforms to improve the accuracy of recognition. By performing a convolution operation on the pixels of the input image, a feature map is extracted, and classification and judgment are performed through a fully connected layer, and finally the supporting indicators of the construction workers are output.
9. The method for identifying illegal operating machinery in subway protection areas based on large-scale model scene understanding according to claim 8 is characterized in that: The detection of the supporting personnel indicators is based on the analysis of the clothing features of the personnel in the video image. By comprehensively applying computer vision technology and deep learning algorithms, the appearance features of the personnel, including but not limited to color, shape and material information, are extracted from the video stream in real time, and compared and analyzed with the standard supporting personnel clothing features pre-stored in the database. Feature matching algorithms, including but not limited to Euclidean distance or cosine similarity calculation, are used to determine whether the personnel in the current scene are wearing construction clothing.
Citation Information
Patent Citations
Rotary tillage operation intelligent control system based on surface topography feature information
CN116158215A
Image description method and device, equipment, storage medium and product
CN118467776A
Monitoring data processing method and device and computer storage medium
CN119169538A
Intelligent construction site safety monitoring method and system fused with multi-modal large model
CN119399702A