5s intelligent monitoring method and system based on multi-task image recognition
Patent Information
- Application Number
- CN202610877858.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-17
- Publication Date
- 2026-09-11
AI Technical Summary
[0005]本申请公开了一种基于多任务图像识别的5S智能监控方法及系统,旨在解决在存在视觉干扰源的复杂工业生产环境中,传统多任务图像识别系统在火灾检测、人员监控和通道物体识别等任务中存在的性能冲突、误报、漏报以及识别精度下降等技术问题
Smart Images

Figure CN122737818A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image recognition technology, and more specifically, to a 5S intelligent monitoring method and system based on multi-task image recognition. Background Technology
[0002] In modern industrial production, strict adherence to the 5S management principles is crucial for ensuring safety and improving efficiency. Traditional safety monitoring systems typically rely on multiple independent visual recognition methods, each focused on a specific monitoring task. This fragmented deployment approach leads to enormous hardware resource consumption, and the need for independent data processing workflows and computing units increases processing latency, making it difficult to meet the stringent real-time early warning requirements of industrial scenarios.
[0003] To overcome these limitations, an intelligent monitoring system based on multi-task image recognition has emerged, aiming to process multiple monitoring tasks simultaneously from a single video stream. This system employs an integrated deep learning approach, but when deployed in complex industrial production environments such as welding and grinding, where there is localized intense light and a large amount of sparks, this approach reveals serious performance conflicts.
[0004] Specifically, the image features of the electric arc and metal sparks generated during welding operations are highly similar to early signs of a fire (such as small flames or electrical sparks), causing the system to frequently misjudge normal production phenomena as fires, generating a large number of false fire alarms. To solve this problem, technicians have used retraining methods to force the method to ignore the "bright, fast-moving small particles" features of welding sparks, thereby reducing false fire alarms. However, since the feature extraction part within the multi-task method is shared by all recognition tasks, this targeted inhibitory learning also affects the method's sensitivity to features of other things. As a result, when a real fire occurs, the system may fail to identify the true fire precursor in time due to ignoring sparks, leading to dangerous missed alarms. Summary of the Invention
[0005] This application discloses a 5S intelligent monitoring method and system based on multi-task image recognition, which aims to solve the technical problems of performance conflicts, false alarms, missed alarms and decreased recognition accuracy of traditional multi-task image recognition systems in complex industrial production environments with visual interference sources in tasks such as fire detection, personnel monitoring and channel object recognition.
[0006] The technical solution of this application is as follows:
[0007] In a first aspect, this application discloses a 5S intelligent monitoring method based on multi-task image recognition, comprising the following steps:
[0008] Acquire video and sensor information of the work area where visual interference sources exist; these visual interference sources include image interference sources generated by welding and grinding processes.
[0009] The status information of the work area is determined based on the video information and sensor information; the status information refers to the parameters corresponding to the work mode determined by the visual interference image of the work area.
[0010] Based on this status information, the response sensitivity of the strong light point, the flame pattern and duration information are adjusted to obtain the adjustment results of fire detection; among them, the flame pattern refers to the change in the area of the flame area, the change in the flame color and the smoke pattern.
[0011] Based on this status information, and combined with the extracted time series, the human body contour information and corresponding color areas are identified, and the adjustment results of personnel monitoring are obtained.
[0012] Based on this status information, the object's edge, shape, and preset area are identified to obtain the adjustment result of the channel object.
[0013] Alarm information is generated based on the adjustment results of fire detection, personnel monitoring, and passageway objects.
[0014] Secondly, this application also discloses a 5S intelligent monitoring system based on multi-task image recognition, the system comprising:
[0015] The information acquisition module is used to acquire video information and sensor information of the work area where there are visual interference sources; the visual interference sources include image interference sources generated by welding and grinding processes.
[0016] The status information recognition module is used to determine the status information of the work area based on the video information and sensor information; the status information refers to the parameters corresponding to the work mode determined by the visual interference image of the work area.
[0017] The fire detection adjustment module is used to adjust the response sensitivity of the strong light point, the flame pattern and duration information according to the status information to obtain the adjustment result of fire detection; wherein, the flame pattern refers to the change in the area of the flame area, the change in the flame color and the smoke pattern.
[0018] The personnel monitoring adjustment module is used to identify human body contour information and corresponding color areas based on the status information and the extracted time series, so as to obtain the adjustment result of personnel monitoring.
[0019] The channel object adjustment module is used to identify the object's edge, shape, and preset area based on the status information, and obtain the adjustment result of the channel object.
[0020] The alarm generation module is used to generate alarm information based on the adjustment results of fire detection, personnel monitoring, and passageway objects.
[0021] Beneficial effects
[0022] This application provides a 5S intelligent monitoring method based on multi-task image recognition. By acquiring video and sensor information of work areas with visual interference sources, and determining the status information of the work area based on this information, it can specifically adjust the response sensitivity of strong light points, the shape and duration of firelight for fire detection, combine time-series recognition of human contour information and corresponding color regions for personnel monitoring, and identify object edges, shapes, and preset area regions for channel object monitoring. Finally, alarm information is generated based on these adjustments. This method effectively solves the problems of false fire alarms, difficulties in personnel identification, and decreased accuracy in channel debris identification in existing multi-task image recognition systems under strong visual interference environments such as welding and grinding. By introducing work area status information for adaptive adjustment, this application can distinguish between normal production phenomena and real dangers, avoiding false alarms, while ensuring accurate identification of personnel and objects under complex lighting conditions. This significantly improves the reliability, accuracy, and real-time early warning capability of the monitoring system, overcoming the performance bottleneck caused by inter-task constraints in existing technologies. Attached Figure Description
[0023] Figure 1 A schematic diagram of a 5S intelligent monitoring method based on multi-task image recognition provided in this application.
[0024] Figure 2 This is a schematic diagram of a 5S intelligent monitoring system based on multi-task image recognition provided for this application. Detailed Implementation
[0025] The technical solutions of this application will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this application, and not all embodiments. The components of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0026] Reference Figure 1 The diagram illustrates an embodiment of a 5S intelligent monitoring method based on multi-task image recognition according to the present invention, which may specifically include the following steps:
[0027] S101, acquire video information and sensor information of the work area where visual interference sources exist; the visual interference sources include image interference sources generated by the welding process and the grinding process;
[0028] S102, determine the status information of the work area based on the video information and sensor information; the status information refers to the parameters corresponding to the work mode determined by the visual interference image of the work area;
[0029] S103, Based on the state information, adjust the response sensitivity of the strong light point, the flame pattern and duration information to obtain the adjustment result of fire detection; wherein, the flame pattern refers to the change in the area of the flame area, the change in the flame color and the smoke pattern.
[0030] S104. Based on the status information and combined with the extracted time series, human body contour information and corresponding color areas are identified to obtain the adjustment result of personnel monitoring.
[0031] S105, Based on the state information, identify the object edge, object shape, and preset area of the object to obtain the adjustment result of the channel object;
[0032] S106 generates alarm information based on the adjustment results of fire detection, personnel monitoring, and passageway objects.
[0033] This application acquires video and sensor information from the work area and determines the area's status based on this information. This allows for targeted adjustment of parameters for fire detection, personnel monitoring, and object recognition in passageways, effectively addressing challenges posed by visual interference sources. Consequently, this application significantly improves the accuracy and reliability of multi-task image recognition in complex industrial environments, reduces false alarms and missed alarms, and ensures the effective operation of the 5S intelligent monitoring system.
[0034] To better understand the technical solution of this application, some key terms involved will be explained first.
[0035] "Visual interference sources" refer to factors within the work area that may interfere with image recognition, such as arc light and metal sparks generated during welding processes, and sparks and fumes generated during grinding processes. The image characteristics of these interference sources may resemble signs of a fire, or cause localized overexposure and reduced contrast in the image, thereby affecting the recognition accuracy of the monitoring system.
[0036] "Video information" refers to image or video data that reflects the real-time situation of the work area, obtained through visual sensors such as cameras.
[0037] "Sensor information" refers to data acquired through other types of sensors besides video information, such as acoustic information acquired by acoustic sensors, thermal image information acquired by thermal imagers, and chemical composition information acquired by chemical sensors. This information can help determine the true condition of the work area.
[0038] "Status information" refers to the current working mode or environmental characteristics of the work area determined by a combination of video and sensor information. For example, when welding or grinding operations are detected, the system will identify the corresponding working mode and determine the parameters related to that mode, such as light intensity and spark density.
[0039] "Response sensitivity to strong light points" refers to the fire detection module's sensitivity in identifying areas of strong light in an image. In work areas with visual interference sources, this sensitivity needs to be adjusted based on status information to distinguish between strong light generated during normal operations and signs of fire.
[0040] "Flame morphology" refers to the visual characteristics of a fire or similar fire phenomenon, including changes in the area of the flame, changes in the color of the flame, and the shape of the smoke. These characteristics are important criteria for identifying a fire.
[0041] A "time series" refers to a series of data points arranged in chronological order. In personnel monitoring, time series can be used to track the movement trajectory of human body outlines, thereby more accurately identifying and judging personnel activities.
[0042] In personnel monitoring, "color zones" specifically refer to specific color areas related to personnel safety, such as the color of safety helmets, work clothes, and name tags.
[0043] "Object edge", "object shape" and "preset area of object" are key features for channel object recognition, used to determine whether there are unwanted objects in the channel.
[0044] The implementation environment of this application is typically an industrial production workshop, which is equipped with multiple surveillance cameras, various sensors (such as acoustic sensors, thermal imaging sensors, chemical composition sensors, etc.), and computing units for data processing and alarm generation. These devices work together to acquire and process various information from the work area in real time.
[0045] This application provides a 5S intelligent monitoring method based on multi-task image recognition. Its core lies in the intelligent perception and utilization of the status information of the work area to achieve adaptive adjustment of multiple tasks such as fire detection, personnel monitoring and channel object recognition, thereby effectively coping with visual interference in complex industrial environments.
[0046] There are several ways to acquire video and sensor information from work areas with visual interference sources. For example, video information can be acquired using industrial-grade high-definition cameras deployed within the work area, which can capture images or video streams of the work area in real time. Sensor information can be acquired using various sensors integrated into the monitoring system; for example, acoustic sensors can be used to detect abnormal noise, thermal imaging sensors can be used to monitor temperature changes, and chemical composition sensors can be used to detect the concentration of specific harmful substances in the air. These sensors can be deployed independently or integrated into a multi-functional sensor unit. In one implementation, video recording can be manually initiated by the operator, and sensor information can be manually read and input into the system by the operator.
[0047] To determine the status information of the work area based on the video and sensor information, the following methods can be used. The system can preset multiple work modes, each corresponding to a specific type and intensity of visual interference sources. For example, when a large amount of arc light and metal sparks appear in the video information, and sensor information (such as acoustic information) shows high-frequency welding noise, the system can determine that the current work area is in the "welding process" mode. When a large amount of smoke and sparks appear in the video information, and sensor information (such as thermal imaging information) shows a local temperature increase, the system can determine that the current work area is in the "grinding process" mode. These mode determinations can be achieved through matching preset rules or by training a machine learning model. In one implementation, the status information of the work area can be manually input into the system by on-site management personnel according to the actual situation.
[0048] In terms of adjusting the response sensitivity of the strong light point, the flame pattern, and the duration information based on the aforementioned state information to obtain the adjusted fire detection results, the parameters of the fire detection algorithm can be dynamically adjusted according to the determined work area state information. For example, when the system identifies that it is currently in a "welding process" mode, since welding sparks are similar to the initial signs of a fire, the system can reduce the response sensitivity to the feature of "bright, fast-moving small particles," while increasing the weight of more discriminative features such as changes in the area of the flame, changes in the color of the flame, and the shape of the smoke. In addition, the judgment threshold for the duration of the flame can be extended according to the characteristics of the welding operation to avoid misjudging short-lived welding sparks as a fire. In one implementation, the adjusted fire detection results can be searched and applied based on a preset fixed parameter table, which is preset according to different work modes.
[0049] In terms of identifying human contour information and corresponding color regions based on the aforementioned state information and combined with the extracted time series data to obtain the adjustment results of personnel monitoring, the system can optimize the personnel monitoring algorithm based on the state information of the work area. For example, in environments with strong light interference, the system can enhance its edge detection capability for human contours and use time series information to track human contours, maintaining continuous identification of personnel even under partial occlusion or insufficient light. Simultaneously, by combining the time series data, specific color regions on the human contour can be identified more accurately, such as the color of the safety helmet and the color of work clothes and name tags, thereby determining whether personnel comply with safety regulations. In one implementation, the adjustment results of personnel monitoring may rely solely on the identification of human contours and color regions in a single frame image, without considering time series information.
[0050] In terms of identifying object edges, shapes, and preset area regions based on the aforementioned state information to obtain the adjustment results for channel objects, the system can adjust the parameters for channel object identification based on the state information of the work area. For example, in areas with complex lighting or dust, the system can enhance the edge detection algorithm and improve the ability to recognize blurred images. Simultaneously, by identifying the shape and preset area region of objects, it can determine whether there are unwanted objects in the channel. For example, the shapes and sizes of tools or materials that are allowed to be temporarily placed can be preset; when an identified object does not conform to these preset characteristics, it is judged as unwanted material. In one implementation, the adjustment results for channel objects can be obtained by identifying objects simply through image thresholding without performing complex edge, shape, and area analysis.
[0051] Regarding the generation of alarm information based on adjustments made to fire detection, personnel monitoring, and passageway object detection, the system will immediately generate alarm information when the fire detection module determines that there is a fire risk, the personnel monitoring module detects personnel violations (such as not wearing a safety helmet, not wearing work clothes or name tags), or the passageway object detection module detects that the passageway is blocked by debris. Alarm information can include various forms such as audible alarms, visual alarms (such as flashing lights), and notifications to relevant management personnel to ensure timely response and handling.
[0052] The overall working principle of this application lies in the introduction of the core concept of "state information," which enables adaptive adjustment of the multi-task image recognition system. Traditional multi-task systems, when facing work environments with strong visual interference such as welding and grinding, suffer from performance degradation in other tasks (such as personnel monitoring and object recognition) due to their shared feature extraction network treating all tasks equally. This results in optimizations made to address one task (e.g., false fire alarms) weakening the performance of others. This application, by acquiring video and sensor information from the work area, can accurately determine the specific work mode of the current work area, such as whether it is a welding or grinding process. Based on this state information, the system no longer blindly performs uniform processing but can selectively adjust the recognition parameters of each sub-task.
[0053] Specifically, when the system identifies a welding process in the work area, it adjusts the sensitivity of the fire detection module's strong light response, the flame pattern, and its duration based on this status information. This means the system will reduce its sensitivity to normal production phenomena such as welding sparks, while paying more attention to more fire-indicating features such as changes in the area and color of the flame zone and the smoke pattern, and extending the threshold for judging the flame duration, thereby effectively reducing false fire alarms.
[0054] Meanwhile, the accuracy of personnel monitoring tasks can be affected by strong light interference. This application can adjust the parameters of the personnel monitoring module through status information, and enhance the ability to recognize human body contours by combining the extracted time series. Even in cases of local overexposure or low contrast, it can more accurately identify human body contour information and corresponding color areas, such as safety helmets and work clothes and name tags, through continuous frame motion trajectory information, ensuring effective monitoring of personnel safety and compliance.
[0055] Furthermore, for channel object recognition tasks, object edges and shapes are easily blurred in complex lighting and dusty environments. This application also utilizes state information to adjust the parameters of the channel object recognition module, optimize the recognition algorithms for object edges, object shapes, and the preset area of objects, improve recognition accuracy under adverse lighting conditions, and thus accurately determine whether the channel is blocked by debris.
[0056] Ultimately, the adjustments made to fire detection, personnel monitoring, and passageway object detection are aggregated into the alarm generation module. Based on these adjusted and more accurate identification results, corresponding alarm information is generated. This collaborative approach allows each task to independently and intelligently optimize its parameters when facing specific interference sources, avoiding negative impacts between tasks and significantly improving the reliability and accuracy of the entire 5S intelligent monitoring system in complex industrial environments.
[0057] The core innovation of this application lies in the introduction of the concept of "state information" and the implementation of adaptive adjustment of the multi-task image recognition system based on this. Compared with the closest existing technology, although the existing technology also adopts a multi-task image recognition method, when facing complex industrial production environments such as welding and grinding where there is local strong light and a large number of sparks, the shared feature extraction network treats all tasks equally. This leads to optimizations made to solve one task (such as false fire alarms) actually weakening the performance of other tasks (such as personnel monitoring and channel object recognition), resulting in false alarms and missed alarms.
[0058] This application acquires video and sensor information from the work area and determines the "status information" of the work area based on this information, such as whether the current process is welding or grinding. This intelligent perception of the work mode allows the system to adjust the recognition parameters of each sub-task accordingly. For example, during a welding process, the fire detection module reduces its sensitivity to welding sparks and focuses more on key features such as the shape and duration of the flame, thereby effectively reducing false fire alarms. Existing technologies, in order to reduce false alarms, may use retraining to force the method to ignore features like "bright, fast-moving small particles" such as welding sparks. This inhibitory learning affects the method's sensitivity to features of other objects, potentially leading to missed detections when a real fire occurs. This application, however, uses intelligent adjustment to ensure the accuracy of fire detection while avoiding negative impacts on other tasks.
[0059] Furthermore, this application adjusts the parameters for personnel monitoring and channel object recognition based on status information. Under strong light interference, existing technologies struggle to accurately determine whether welders are wearing safety helmets or identify their work badges; the outlines and boundaries of necessary tools and workpieces within the channel also become blurred, leading to decreased recognition accuracy. This application significantly improves recognition accuracy in complex lighting and dusty environments by combining time-series recognition of human contour information and corresponding color regions, as well as recognizing object edges, shapes, and preset area regions, and adjusting these parameters based on status information.
[0060] Therefore, this application overcomes the problem of performance degradation caused by mutual interference between tasks in complex scenarios in the existing multi-task methods, and realizes the improvement of the reliability of personnel monitoring and channel object recognition while ensuring the accuracy of fire detection, providing a more reliable and efficient solution for safety monitoring of complex workstations.
[0061] In some of the embodiments described above in this application, traditional 5S intelligent monitoring methods based on single video information may suffer from inaccurate judgment of the work area's condition in areas with visual interference sources such as welding and grinding processes. This interference from factors like strong light and smoke can affect the reliability of fire detection, leading to false alarms or missed alarms. If these problems are not addressed, safety hazards may go undetected, or the monitoring system's reliability may be reduced due to frequent false alarms. Therefore, this application further proposes to improve the accuracy and robustness of monitoring in complex industrial environments by introducing multimodal sensor information and refining the fire detection process.
[0062] Specifically, the sensor information includes acoustic information, thermal imaging information, and chemical composition information; determining the status information of the work area based on the video information and sensor information includes:
[0063] The status information of the work area is determined based on the video information, the acoustic information, the thermal imaging information, and the chemical composition information.
[0064] The step of adjusting the response sensitivity of the strong light point, the flame pattern, and the duration information based on the state information to obtain the adjustment result of fire detection includes:
[0065] Preset rule information is obtained based on the state information;
[0066] Based on the preset rule information, the video information, the acoustic information, the thermal imaging information, and the chemical composition information are matched to obtain fire conclusion information; the fire conclusion information refers to the conclusion of whether a fire has occurred.
[0067] Based on the fire conclusion information, the response sensitivity of the strong light point, the flame pattern and duration information are adjusted to obtain the adjusted fire detection results.
[0068] Sensor information can be understood as non-visual data, in addition to video information, used to assist in judging the status and potential risks of the work area. Specifically, acoustic information refers to environmental sound data collected through devices such as microphones, such as specific noise patterns generated during welding or grinding; thermal imaging information refers to temperature distribution data obtained through infrared thermal imagers, which can reflect abnormal heating of objects or areas; chemical composition information refers to the concentration data of specific chemical substances in the air detected by devices such as gas sensors, such as carbon monoxide, carbon dioxide, or other combustion products in smoke. This multimodal sensor information can provide richer and more comprehensive environmental perception capabilities than single video information, helping to overcome the limitations caused by visual interference.
[0069] Furthermore, determining the status information of the work area based on video and sensor data is accomplished through comprehensive analysis of video, acoustic, thermal imaging, and chemical composition information. For example, in a welding process, video data may show strong light and smoke, acoustic data may detect arc sounds, thermal imaging may show localized high temperatures, and chemical composition information may detect welding fumes. The system uses this multi-source data, combined with preset work mode parameters, to more accurately determine whether the current work area is in a normal welding mode, an abnormal overheating mode, or a potential fire mode.
[0070] In practical applications, the process of adjusting fire detection results based on status information involves several steps. First, based on the determined status information of the work area, the system dynamically acquires or generates a set of preset rule information. These rules are defined for fire characteristics under different work modes and environmental conditions. For example, in normal welding mode, a certain intensity of flames and smoke is allowed, but their shape and duration should conform to specific rules; while during non-work periods, any flame or smoke may be considered abnormal. Second, using this preset rule information, real-time video information, acoustic information, thermal imaging information, and chemical composition information are matched and analyzed. For example, if the video detects a flame, while the thermal image shows a sharp increase in temperature, the acoustic image detects an explosion, and the chemical composition detects a high concentration of carbon monoxide, these multi-source data highly match the fire rules, and a fire conclusion can be drawn. The fire conclusion clearly indicates whether a fire has occurred. Finally, based on this fire conclusion information, the system adjusts the response sensitivity of the strong light point, the recognition threshold of the flame shape, and the judgment criteria for the flame duration accordingly, thereby obtaining a more accurate fire detection adjustment result. For example, if a fire is detected, an alarm will be triggered immediately, and the threshold for subsequent detections may be lowered to increase sensitivity; if a false alarm is detected, parameters will be adjusted to reduce false alarms in similar situations.
[0071] This application's solution incorporates acoustic, thermal imaging, and chemical composition information as components of sensor data, and refines the process for determining state information and adjusting fire detection. This effectively solves the problem of false alarms or missed alarms when relying solely on video information for fire detection in work areas with visual interference sources (such as strong light and smoke from welding or grinding). Specifically, when visual interference exists in the work area, video information alone may not accurately distinguish between normal work phenomena and fire signs. For example, strong light and smoke from welding may be misjudged as a fire. By combining acoustic information, specific sound patterns related to fire (such as cracking sounds and burning sounds) can be identified and distinguished from normal work sounds (such as welding arc sounds and grinding friction sounds). Thermal imaging information can monitor abnormal temperature rises, which are often important indicators of fire occurrence, while localized high temperatures generated during normal work usually have a specific range and duration. Chemical composition information can detect combustion products (such as carbon monoxide and carbon dioxide in smoke), thus providing more direct evidence of a fire. The fusion of these multimodal information modalities enables the system to comprehensively assess the state of the work area from multiple dimensions, avoiding the limitations of relying solely on visual information. Furthermore, by deriving preset rule information from the state information and matching it with multi-source information to arrive at fire conclusions, this solution achieves intelligent and adaptive fire detection. The preset rule information dynamically adjusts the fire judgment criteria based on the current work mode (determined by the state information). For example, during welding operations, the tolerance for flames and smoke is higher than during non-operational periods. This rule-based matching mechanism allows the system to more accurately distinguish between normal work phenomena and genuine fire risks, thereby significantly improving the accuracy and reliability of fire detection in complex and ever-changing industrial environments.
[0072] As a specific implementation, consider a 5S intelligent monitoring system based on multi-task image recognition deployed in a factory workshop with welding and grinding processes. In addition to high-definition cameras acquiring video information, the system is equipped with acoustic sensors to collect ambient noise, an infrared thermal imager to monitor temperature distribution, and gas sensors to detect chemical components in the air such as carbon monoxide and carbon dioxide. When welding operations are underway in the workshop, the system first identifies the welding arc light and smoke through video information. Simultaneously, the acoustic sensors detect the characteristic arc sound of welding, the thermal imager shows a localized temperature increase at the welding point, but the gas sensors do not detect abnormal concentrations of combustion products. At this point, based on this multi-source information and combined with preset welding operation mode parameters, the system determines that the work area is in a "normal welding mode" state. According to the "normal welding mode" state information, the system loads a set of preset fire detection rules for the welding process. These rules allow for the presence of strong light and smoke within a certain range and set relatively lenient thresholds for the flame pattern (such as changes in area and color) and duration. Subsequently, the system continuously matches real-time video, acoustic, thermal imaging, and chemical composition information. If, at this moment, the video feed shows the flames suddenly and rapidly expanding, changing color from bright white to orange-red, accompanied by a large amount of thick smoke, while the thermal imager detects a rapid temperature spread in the surrounding area, the acoustic sensor captures abnormal popping sounds, and the gas sensor detects a sharp increase in carbon monoxide concentration, the system will match this multi-source data with preset fire rules. Because these characteristics highly match the fire rules, the system will conclude that a fire has occurred. Based on this conclusion, the system will immediately adjust the fire detection parameters, such as maximizing the sensitivity of the strong light point, setting the thresholds for judging the flame shape and duration to the most stringent level, and quickly generating an alarm to notify relevant personnel to take action. In this way, even in welding environments with severe visual interference, the system can accurately distinguish between normal operations and fire risks through the fusion of multimodal sensor information and rule matching, thereby achieving efficient and reliable intelligent monitoring.
[0073] Specifically, based on the status information and combined with the extracted time series, the human body contour information and corresponding color areas are identified to obtain the adjustment results of personnel monitoring, which can be further refined.
[0074] In some embodiments of this application, the color region includes a safety helmet image in the head region and a work uniform and name tag image in the torso region; the step of identifying human body contour information and corresponding color regions based on the status information and combined with the extracted time series to obtain the adjustment result of personnel monitoring specifically includes the following steps:
[0075] First, based on the state information, the human body region in the video information is separated into multiple independent human body contours.
[0076] Next, for each individual human silhouette, the safety helmet image in the head region and the work clothes and name tag image in the torso region are identified.
[0077] Subsequently, time series analysis is used to track each individual human body contour to extract images of safety helmets and work clothes / signs, thereby obtaining the correlation between the safety helmet images, work clothes / sign images and human body contours.
[0078] Furthermore, by using human posture information, adjacent or partially overlapping human silhouettes are distinguished, ultimately resulting in adjustments to personnel monitoring.
[0079] Specifically, the color-coded areas refer to specific visual regions used in image recognition to identify personnel identity or safety compliance. Specifically, the helmet image in the head area refers to the visual features of the helmet detected and identified in the personnel's head area using image processing technology, such as its shape, color, and texture. The work clothes and name tag images in the torso area refer to the visual features of the work clothes and name tags detected and identified in the personnel's torso area, such as their patterns, colors, text, or logos. This image information is an important basis for determining whether personnel comply with 5S management standards.
[0080] In this embodiment of the invention, contour separation of the human body region in the video information refers to using an image segmentation algorithm to accurately separate the human target in the video frame from the background, forming an independent and complete individual contour. This facilitates subsequent independent feature extraction and analysis of each individual. Tracking each independent human contour using time series analysis refers to establishing the continuity of the individual in the time dimension by analyzing the changes in the position, shape, and motion trajectory of the human contour in consecutive video frames, thereby achieving continuous tracking of a specific person.
[0081] The method of distinguishing adjacent or partially overlapping human body contours through human posture information refers to using deep learning or computer vision technology to analyze the posture characteristics of each individual, such as limb direction and joint position, when multiple workers are close to or occlude each other in the work area, in order to accurately distinguish different individuals and avoid identifying overlapping human bodies as a single target or erroneously merging information from different individuals.
[0082] This application's solution refines the adjustments made to personnel monitoring, addressing issues such as inaccurate identification, discontinuous tracking, and overlapping occlusion that can arise with traditional personnel identification methods in complex work environments. Specifically, by separating the human body regions in video information based on status information, each individual can be accurately extracted from the complex background, laying the foundation for subsequent feature recognition. Subsequently, for each individual human silhouette, the system identifies the safety helmet image of the head region and the work clothes and name tag image of the torso region, enabling the system to directly obtain crucial information on whether personnel are wearing safety equipment and work clothes, which is essential for safety compliance checks in 5S intelligent monitoring. Furthermore, by using time-series tracking of human silhouettes, the system ensures that the identity and status of personnel can be continuously and stably monitored during movement, avoiding information loss due to brief occlusion or changes in perspective. Crucially, by introducing human posture information to distinguish adjacent or partially overlapping human silhouettes, the system effectively overcomes the problem of traditional methods failing to accurately distinguish individuals when multiple people are active in the same area, significantly improving the accuracy and robustness of personnel monitoring.
[0083] In addition, this application proposes an improved method for monitoring objects in passageways, which identifies and judges long-term objects in the passageway area, thereby more effectively managing the cleanliness and safety of the work area.
[0084] Specifically, the above method includes: detecting object images appearing in the channel area based on the video information, and obtaining the initial position information and initial shape information of the object images;
[0085] If, after a preset time has elapsed, the initial position and shape information of the object image remain unchanged, then the object is determined to be a long-term object.
[0086] Analyze the shape characteristics of the long-term object;
[0087] Based on the shape characteristics of the long-term object, determine whether the shape characteristics of the long-term object match the preset shape of the items that are allowed to be temporarily placed.
[0088] If the shape of the long-term object does not match the preset shape of the items that can be temporarily placed, a clutter alarm will be triggered.
[0089] Specifically, video information refers to continuous images or video streams captured by surveillance cameras installed within the work area. A passageway area refers to a specific path or space within the work area designated for personnel passage or material transport. An object image refers to any non-background element identified in the video information using image processing techniques, such as tools, equipment, or material piles. Initial position information can be understood as the position of an object image in the image coordinate system when it is first detected, such as its centroid coordinates or the coordinates of the top-left corner of its bounding box. Initial shape information can be understood as the geometric features of an object image, such as its outline, area, and aspect ratio, when it is first detected.
[0090] The preset time refers to a time threshold set by the system to determine whether an object remains in the passage area for an extended period. This preset time can be flexibly configured according to the actual working environment, 5S management requirements, and the importance of the passage; for example, it can be set to 5 minutes, 15 minutes, or 30 minutes. If the initial position and shape information of the object image remain unchanged after the preset time, it indicates that the object is stationary within the passage area and has not undergone significant deformation or movement; in this case, the object is judged as a long-term object.
[0091] In practical applications, analyzing the shape characteristics of long-term objects can include extracting their geometric features (such as length, width, height, and volume estimation), texture features, and color features. Preset shapes for temporarily placed items refer to the shape characteristics of items explicitly permitted for temporary placement in aisle areas according to 5S management standards, such as toolboxes of certain sizes or temporary storage boxes. These preset shape characteristics can be stored in the system's database as a basis for judgment.
[0092] When the shape of a long-term object does not match the preset permitted shapes for temporary items, it indicates that the object is unauthorized clutter, and the system will trigger a clutter alarm. This alarm may include audible alerts, visual alerts (such as highlighting on the monitoring interface), SMS notifications, or email notifications to remind relevant management personnel to handle the situation promptly.
[0093] This application effectively addresses the limitations of traditional monitoring methods in identifying and managing long-term debris in passageways. By introducing the judgment of object dwell time and comparison of shape characteristics, this solution achieves accurate identification and alarm for long-term lingering and non-compliant debris in passageways. This not only improves the intelligence level of the 5S intelligent monitoring system and reduces the workload of manual inspections, but also enables the timely detection and handling of potential passageway blockages and safety hazards, thereby significantly improving the cleanliness, traffic efficiency, and overall safety of the work area, providing strong support for enterprises to achieve lean production and safety management.
[0094] In some preferred embodiments, a specific example is given below. Assume a high-definition surveillance camera is installed in a material aisle area of a production workshop. The system continuously collects video information from this aisle area. One day, a worker temporarily places a large discarded wooden crate on the edge of the aisle and leaves the area. The system detects the image of the wooden crate in the video information and records its initial position and shape information. The system has a preset time limit of 15 minutes for debris to remain in the aisle. During the next 15 minutes, the system continuously monitors the wooden crate and finds that its position and shape information remain unchanged. Therefore, the system determines that the wooden crate is a permanent object. Subsequently, the system analyzes the shape characteristics of the wooden crate, extracting its length, width, height, and overall outline. The system compares these characteristics with a preset database of allowed temporary object shapes. This database may include standard-sized turnover boxes or small tool carts, but does not include discarded wooden crates. Because the shape characteristics of the wooden crate do not match the preset allowed temporary object shapes, the system immediately triggers a debris alarm. The alarm was broadcast audibly via the workshop's public address system, and the area where the wooden crate was located was highlighted on the monitoring center's screen. Simultaneously, a text message notification was sent to the workshop management personnel's mobile phones. Upon receiving the alarm, the management personnel quickly went to the scene and removed the abandoned wooden crate, thus preventing passageway blockage and potential safety risks.
[0095] This application further proposes the aforementioned 5S intelligent monitoring method based on multi-task image recognition, wherein the sensor information includes thermal imaging information and auxiliary production control system data; the determination of the status information of the work area based on the video information and sensor information includes:
[0096] The status information of the work area is determined based on the video information, the thermal imaging information, and the data from the auxiliary production control system.
[0097] The step of adjusting the response sensitivity of the strong light point, the flame pattern, and the duration information based on the state information to obtain the adjustment result of fire detection includes:
[0098] Based on the state information, preset rule information for material heat treatment of welding and grinding processes is obtained;
[0099] Based on the preset rules for material heat treatment in the welding and grinding processes, the video information, the thermal imaging information, and the auxiliary production control system data are matched to obtain fire conclusion information.
[0100] Based on the fire conclusion information, the response sensitivity of the strong light point, the flame pattern and duration information are adjusted to obtain the adjusted fire detection results.
[0101] Specifically, thermal imaging information refers to temperature distribution data of the work area acquired through infrared thermal imagers. It reflects the thermal radiation of an object's surface, thus providing temperature characteristics that are difficult to capture visually. Auxiliary production control system data refers to real-time data related to the production process, such as the operating status of welding equipment, the power of grinding equipment, the type of material being processed, the preset heat treatment parameters of the material, and production plans. This data provides important contextual information for determining the normal heat generation in the work area.
[0102] Determining the status information of the work area based on video information, thermal imaging information, and data from the auxiliary production control system can be understood as comprehensively utilizing multi-source heterogeneous data to conduct a comprehensive assessment of the current operating conditions, ambient temperature, and equipment operating status of the work area, thereby forming a more accurate and instructive status description. For example, when the auxiliary production control system data shows that high-temperature welding is in progress, the system will identify that it is currently in a high-heat-generating working mode.
[0103] In practical applications, the preset rules for material heat treatment in welding and grinding processes specifically refer to pre-defined judgment criteria based on the normal heat, sparks, and fumes that different materials may generate during welding or grinding, as well as their unique thermal and visual characteristics under abnormal conditions (such as overheating or ignition). These rules may include temperature thresholds, temperature change rates, specific spectral responses, and matching patterns between flame duration and morphology. For example, for welding a specific metal, its normal operating temperature range and spark characteristics are known; exceeding this range or exhibiting atypical flame patterns may trigger an alarm.
[0104] Furthermore, matching video information, thermal imaging information, and auxiliary production control system data to obtain fire conclusion information involves comparing and analyzing real-time acquired video, thermal imaging, and production control data against preset rules. For example, if thermal imaging information shows that the temperature in a certain area has risen sharply and exceeded the preset safety threshold for material heat treatment, and video information captures abnormal flames or smoke patterns, and the auxiliary production control system data does not indicate any normal production activity that can explain this phenomenon, then fire conclusion information can be derived, i.e., determining whether a fire has occurred.
[0105] This application significantly improves the accuracy and reliability of fire detection in work areas with visual interference sources such as welding and grinding processes. Specifically, the introduction of thermal imaging information enables the system to make more accurate judgments about fires from a temperature perspective, effectively distinguishing between normal production heat and abnormal fire heat. The integration with auxiliary production control system data provides crucial production context for fire detection, allowing the system to dynamically adjust detection strategies based on actual working conditions, thereby significantly reducing false alarms caused by production activities (such as welding sparks and high temperatures from grinding). Furthermore, by pre-setting rules for material heat treatment in welding and grinding processes, the logic of fire detection is more refined and intelligent, enabling earlier and more accurate identification of potential fire risks, buying valuable time for timely countermeasures, and effectively ensuring the safety of the work area.
[0106] In some preferred embodiments, a specific example is given below. Suppose that in a metal processing workshop, stainless steel welding is being carried out in a certain area. The monitoring system first captures the intense arc light and sparks generated by the welding through video information. Simultaneously, a thermal imager acquires real-time temperature distribution data for the area, showing that the local temperature at the welding point is as high as several hundred degrees Celsius, while the temperature of the surrounding area is normal. In addition, data from the auxiliary production control system indicates that the welding equipment is currently operating normally and is processing stainless steel, and the corresponding preset rules for material heat treatment are activated.
[0107] At this point, based on video information, thermal imaging information, and data from the auxiliary production control system, the system determines that the work area is in a "stainless steel welding operation" state. Based on this state, the system invokes preset material heat treatment rules for the stainless steel welding process. These rules may include: the welding point temperature is normal within a specific range, but if the surrounding area temperature rises rapidly within a short period and exceeds a certain threshold, or if atypical flame patterns appear (e.g., the flame area expands abnormally, and the color changes from the characteristic bluish-white of the welding arc to orange-red), it may indicate a fire.
[0108] The system matches real-time video information, thermal imaging information, and data from the auxiliary production control system with these preset rules. If the matching results show that although the temperature at the welding point is high, the temperature in the surrounding area is stable, and the flame pattern conforms to the normal characteristics of stainless steel welding, the fire conclusion information is "no fire". At this time, the system will adjust the response sensitivity of the strong light point to a lower level based on this conclusion and ignore the normal welding flame pattern and duration to avoid false alarms.
[0109] However, if thermal imaging shows a sudden, abnormal temperature rise in a piece of scrap near the welding point, accompanied by a continuous, spreading orange-red flame and thick smoke in the video, and the auxiliary production control system data does not show any other normal production activities to explain this phenomenon, the system will conclude that a fire has occurred. Based on this fire conclusion, the system will immediately adjust the response sensitivity of the high-intensity light point to a high level and pay close attention to the fire's shape (such as area and color) and duration, thereby quickly generating an alarm message to notify relevant personnel to handle the situation.
[0110] In practical applications, this application further proposes a 5S intelligent monitoring method based on multi-task image recognition, which also includes monitoring and judging the activity status of personnel within a preset safe area to enhance the safety management capability of the work area.
[0111] In this embodiment of the invention, the method further includes: setting a preset safety zone within the work area based on video information;
[0112] Monitor the location information of personnel within the preset safety area;
[0113] When the time a person stays in the preset safe area reaches a preset threshold, and the location information of the person is still detected within the preset safe area, the activity status of the person is determined.
[0114] If the activity status of the person is not accessible, an alarm message is generated.
[0115] Specifically, the preset safety zone refers to a specific spatial area within the work area designated according to safety management or 5S management requirements. This could be a hazardous materials storage area, an emergency exit, a restricted equipment operation area, or a temporary storage area. This area can be configured by virtually drawing lines on a video surveillance screen or by importing CAD drawings.
[0116] The location information of the monitoring personnel within the preset safe area can be obtained by analyzing video information in real time, using target detection and tracking algorithms to identify the personnel, and continuously acquiring their coordinate positions in video frames, thereby determining whether the personnel have entered or remained within the preset safe area.
[0117] Furthermore, when the time a person spends within the preset safe area reaches a preset threshold, the system will detect whether the person's location information is still within the area. The dwell time refers to the length of time a person remains within the preset safe area after entering it. The preset threshold can be flexibly configured according to the actual working environment and safety regulations. For example, for emergency exits, the threshold may be set to a very short time, while for certain areas that allow short stays, the threshold can be appropriately extended.
[0118] Based on this, if the detected location information of the person is still within the preset safe area, it is necessary to determine the person's activity status. The activity status can be analyzed based on information such as the person's movement trajectory and speed changes over a period of time. For example, if the person remains stationary or only makes slight movements within the preset safe area for a long time, it can be determined as a non-passage state. The non-passage state refers to a state where the person is not in a normal passage, operation, or authorized stay state, such as prolonged lingering, loitering, or remaining stationary.
[0119] Therefore, if the activity status of the personnel is determined to be an unauthorized access status, the system will generate an alarm message. This alarm message can take various forms, such as visual alarms, audible and visual alarms, and notifications to management personnel, to remind relevant personnel to take timely action.
[0120] Through the aforementioned technical solution, this application enables refined management and monitoring of personnel activities in specific safety areas within the work area. This solution not only effectively identifies and warns of prolonged lingering or non-accessible states of personnel in critical safety areas, significantly reducing the risk of safety accidents caused by improper personnel behavior, but also helps improve the "Sort" and "Sweep" effects in 5S management, ensuring the standardization and safety of the work environment. Compared to basic solutions, this application, by introducing preset safety areas and activity status judgment mechanisms, greatly enhances the intelligent monitoring system's capabilities in personnel behavior management, providing more comprehensive and proactive safety assurance for the work area.
[0121] In some preferred embodiments, it is assumed that within a large production workshop, there exists an area for storing flammable and explosive chemicals, which is explicitly designated as a "hazardous materials storage safety area." Simultaneously, the workshop also has multiple emergency evacuation routes, which are designated as "emergency exit safety areas."
[0122] The system first sets the aforementioned "hazardous materials storage safety area" and "emergency passage safety area" as preset safety areas based on the workshop floor plan or manually on the monitoring screen.
[0123] When an unauthorized employee enters the "hazardous materials storage safe area," the system continuously monitors the employee's location via video analytics. If the employee remains in the area for more than 5 minutes (a preset threshold) and their location remains almost unchanged during those 5 minutes (determined as an unauthorized access point), the system will immediately trigger an alarm. This alarm may manifest as an audible and visual alarm sounding in the area, while simultaneously sending an alert notification containing the employee's image and location information to the mobile terminal of the workshop management personnel.
[0124] For example, if an employee remains stationary within the "emergency exit safety zone" for an extended period, such as more than 30 seconds (a preset threshold), the system will determine that they are not in a passable state and immediately generate an alarm. This helps to promptly detect passage blockages or potential emergencies, thereby ensuring unobstructed emergency evacuation routes.
[0125] In this way, the solution proposed in this application can provide customized, real-time monitoring and early warning for different types of safety areas and potential risks, greatly improving the safety management level of the work area.
[0126] In some of the embodiments described above in this application, a method for monitoring personnel location information within a preset safe area is proposed. However, in its implementation, when personnel overlap or are obstructed, and when the preset safe area is located in a multi-story building, relying solely on a single two-dimensional location information may lead to inaccurate personnel location judgment, thereby affecting the reliability of monitoring and the timeliness of alarms.
[0127] In response, this application further proposes that the aforementioned location information includes second location information and third location information; the location information of the monitoring personnel within the aforementioned preset safety area includes: when there is overlap or obstruction between personnel, the location information of the personnel is corrected by utilizing the characteristics of the unobstructed personnel and the motion trajectory information in continuous time frames to obtain second location information; when the aforementioned preset safety area is located in a multi-story building, the location information of the personnel is mapped and fused spatially by utilizing the personnel location information from different levels of perspective to determine the third location information of the personnel in three-dimensional space.
[0128] Specifically, the second location information refers to more accurate two-dimensional location data obtained by comprehensively analyzing the visible features and historical movement trajectories of a person when they are partially occluded or overlap with other people. The unoccluded features of the person can be understood as the visible and identifiable parts of their body from the current viewpoint, such as the head, shoulders, arms, or work clothes of a specific color. The motion trajectory information in continuous time frames refers to the system's record of the person's movement path over a period of time. By analyzing their movement trends and speed, their possible location when occluded can be predicted. Correcting the person's location information involves refining the initial identified or predicted location by combining the unoccluded features and motion trajectory information to improve the accuracy of the location.
[0129] Third-party location information refers to the precise coordinates of a person in three-dimensional space by integrating perspective data from multiple cameras at different heights or angles in a multi-story building environment. Specifically, the location information of a person at different levels refers to the two-dimensional or three-dimensional location data of the same person captured by monitoring equipment installed on different floors or at different heights. Spatial location mapping and fusion refers to unifying these local location information from different perspectives into a global three-dimensional coordinate system through geometric transformations, coordinate transformations, and other techniques, and then integrating the data to eliminate errors and determine the precise location of the person in three-dimensional space.
[0130] This application's solution effectively addresses the problem of inaccurate personnel location information acquisition in complex monitoring scenarios by introducing second and third location information. Specifically, when personnel overlap or are occluded, traditional methods based on single image frames or simple contour recognition struggle to accurately distinguish and locate individuals. This application utilizes the features of unoccluded personnel, such as visible body parts or clothing color, combined with their motion trajectory information across consecutive time frames, to reasonably infer and correct the position of occluded personnel. For example, even if a person's main body is occluded, but their head or feet are still visible, and their motion trajectory shows they are moving in a certain direction, the system can correct their overall position accordingly.
[0131] In some preferred embodiments, this application is implemented as follows:
[0132] For example, in a welding workshop, when two workers approach each other and partially overlap, traditional image recognition systems may struggle to distinguish between the two individuals. The system in this application identifies the unobstructed features of each worker, such as one worker's safety helmet and the other's tool bag. Simultaneously, the system analyzes their movement trajectories over the preceding seconds; for example, one worker is moving left while the other is moving right. Based on these unobstructed features and movement trajectory information, the system can accurately correct the individual positions of the two workers, thereby obtaining their respective second position information, ensuring accurate tracking of each individual even when they are overlapping.
[0133] In a preferred embodiment of this application, a second position information is further proposed to correct the position information of personnel by utilizing the features of unobstructed personnel and motion trajectory information in continuous time frames when personnel overlap or are occluded, including:
[0134] Extract the motion trajectory information of the personnel in consecutive time frames;
[0135] Based on the movement trajectory information of the person, identify the person's current movement state; the movement state includes uniform movement, accelerated movement, or decelerated movement.
[0136] Based on the current movement state of the person, the parameters for predicting the movement trajectory are adjusted to obtain the adjusted movement trajectory prediction parameters;
[0137] Based on the adjusted motion trajectory prediction parameters, the position information of the person in the time frame is predicted to obtain the predicted position information;
[0138] By using the characteristics of the person not being obscured and the predicted location information, the location information of the person is corrected to obtain the second location information.
[0139] Specifically, extracting a person's motion trajectory information across consecutive time frames refers to using video analytics to continuously record and track a person's position at different points in time, thereby forming their movement path over a period of time. The motion trajectory information can be understood as a set of coordinate points of a person in space that change over time.
[0140] Furthermore, based on the movement trajectory information of the person, the current movement state of the person can be identified, with the aim of gaining a more refined understanding of the person's movement pattern. Specifically, the movement state can include uniform motion, accelerating motion, or decelerating motion. For example, by calculating the speed and acceleration of the person between consecutive frames, it can be determined whether they are moving at a constant speed, gradually increasing speed, or gradually decreasing speed.
[0141] The purpose of adjusting the parameters for predicting the motion trajectory based on the person's current motion state is to improve the accuracy of position prediction. For example, a linear prediction model can be used for a person moving at a constant speed; while a nonlinear model can be used for a person moving at an accelerating or decelerating speed, or the parameters of the linear model can be dynamically adjusted to better fit their motion trend.
[0142] In practical applications, predicting a person's position within a time frame based on adjusted motion trajectory prediction parameters refers to using an optimized prediction model to estimate the person's likely location at a future point in time. For example, Kalman filtering, particle filtering, or other motion model-based prediction algorithms can be used, combining historical trajectories and current motion states, to infer the person's next position.
[0143] Finally, the location information of the person is corrected by using the features of the person who is not occluded and the predicted location information. The purpose is to comprehensively utilize multi-source information to overcome the uncertainty caused by occlusion. Specifically, when a person is partially occluded, their visible features (such as head, hands, or feet) can provide some location clues. Combined with the predicted location information, the initial estimated location can be corrected to obtain more accurate second location information.
[0144] This application significantly improves the accuracy and reliability of personnel position information correction in complex visual environments such as overlapping or occluded personnel. Compared to methods that rely solely on general trajectory information for correction, this solution dynamically identifies the movement state of personnel and adjusts prediction parameters, making the position prediction more closely match the actual movement patterns of personnel, thereby effectively reducing position estimation errors caused by changes in movement patterns. Therefore, in a 5S intelligent monitoring system, more accurate secondary personnel position information can be obtained, which is of great significance for the timely detection of abnormal personnel stays, intrusion into dangerous areas, or unsafe operations, further enhancing the safety management level of the work area.
[0145] In some preferred embodiments, assuming a busy production workshop, the monitoring system needs to accurately track a person performing equipment maintenance. While this person is moving, parts of their body may be briefly obscured by large equipment or another passing colleague.
[0146] First, the system continuously extracts the person's motion trajectory information in consecutive video frames. When occlusion is detected, the system analyzes the person's most recent motion trajectory. For example, if the person was moving in a certain direction at a constant speed before the occlusion, the system will identify their current motion state as uniform motion.
[0147] Next, the system will adjust its internal motion trajectory prediction parameters based on the identified uniform motion state. For example, it will select a prediction model suitable for uniform motion and optimize its smoothing factor or prediction step size.
[0148] Then, using these adjusted prediction parameters, the system predicts the possible location of the person during the occlusion period. For example, by using a Kalman filter, combined with historical trajectories and a uniform motion model, it predicts the person's possible coordinates in the next frame.
[0149] Finally, when an unobstructed feature of the person (e.g., an exposed safety helmet or part of their work clothes) is detected again, the system fuses and corrects the location information of this visible feature with the previously predicted location information. For example, a weighted average of the predicted location and the visible feature location can be used, or optimization can be performed centered on the predicted location under the constraints of the visible feature, thus obtaining a highly accurate second location information. Even during periods of occlusion, this effectively avoids location drift or misjudgment due to missing information.
[0150] This application further proposes an optimization scheme that enhances personnel features through multimodal information fusion, thereby improving the accuracy of personnel position information correction in complex scenes. The method also includes: acquiring visual information, depth information, and thermal imaging information corresponding to the features of the unoccluded parts of the personnel;
[0151] The visual information, depth information, and thermal imaging information are fused to obtain enhanced personnel feature information;
[0152] Based on the enhanced personnel feature information and the personnel's motion trajectory information in continuous time frames, the personnel's position information is corrected to obtain second position information.
[0153] Specifically, acquiring visual, depth, and thermal information corresponding to the features of the unobstructed parts of a person involves simultaneously collecting image data, distance data, and temperature distribution data of the unobstructed areas of the person using multiple sensors deployed in the work area, such as visible light cameras, depth sensors (e.g., ToF cameras or structured light sensors), and thermal imagers. Visual information provides two-dimensional features such as the person's appearance, color, and texture; depth information provides distance information between the person and the sensors, helping to construct the person's three-dimensional structure and distinguish occlusion relationships; and thermal information reflects the person's body temperature distribution, enabling effective identification of people even in low-light or camouflaged conditions.
[0154] Furthermore, fusing the visual information, depth information, and thermal imaging information can be understood as employing a multimodal data fusion algorithm to effectively integrate heterogeneous data from different sensors. For example, feature-level fusion, decision-level fusion, or hybrid fusion strategies can be used. In feature-level fusion, feature vectors from each modality can be extracted, and then these feature vectors can be concatenated or learned through a neural network to generate a more discriminative joint feature representation. The aim is to comprehensively utilize the advantages of different modal information to compensate for the shortcomings of single-modal information, thereby obtaining a more comprehensive and robust description of the person.
[0155] Therefore, the enhanced personnel feature information obtained refers to the comprehensive feature representation that incorporates multi-dimensional information such as visual, depth, and thermal imaging after fusion processing. Compared with single-modal features, this enhanced feature information has higher robustness and discriminative power, and can more accurately represent the identity and status of personnel, especially in complex environments or under partial occlusion.
[0156] Finally, based on the enhanced personnel feature information and the personnel's motion trajectory information in consecutive time frames, the personnel's position information is corrected to obtain second position information. This means that in the personnel position correction process, it no longer relies solely on the original visual features, but utilizes features enhanced by multimodal fusion, combined with the personnel's motion patterns in the time series, and uses more precise matching and tracking algorithms to finely adjust the personnel's actual position, thereby obtaining more accurate second position information.
[0157] Secondly, referring to Figure 2This application further proposes a 5S intelligent monitoring system based on multi-task image recognition, the system comprising:
[0158] Information acquisition module 201 is used to acquire video information and sensor information of the work area where there are visual interference sources; the visual interference sources include image interference sources generated by welding and grinding processes;
[0159] The status information recognition module 202 is used to determine the status information of the work area based on the video information and sensor information; the status information refers to the parameters corresponding to the work mode determined by the visual interference image of the work area.
[0160] The fire detection adjustment module 203 is used to adjust the response sensitivity of the strong light point, the flame pattern and duration information according to the status information to obtain the adjustment result of fire detection; wherein, the flame pattern refers to the change in the area of the flame area, the change in the flame color and the smoke pattern.
[0161] The personnel monitoring adjustment module 204 is used to identify human body contour information and corresponding color areas based on the status information and the extracted time series, and obtain the personnel monitoring adjustment result.
[0162] The channel object adjustment module 205 is used to identify the object edge, object shape and preset area of the object according to the state information, and obtain the adjustment result of the channel object.
[0163] The alarm generation module 206 is used to generate alarm information based on the adjustment results of fire detection, personnel monitoring, and passageway objects.
[0164] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A 5S intelligent monitoring method based on multi-task image recognition, characterized in that, include: Acquire video and sensor information of the work area where visual interference sources exist; The visual interference sources include image interference sources generated by the welding process and the grinding process; The status information of the work area is determined based on the video information and sensor information; the status information refers to the parameters corresponding to the work mode determined by the visual interference image of the work area. Based on the aforementioned status information, the response sensitivity of the strong light point, the flame pattern, and the duration information are adjusted to obtain the adjustment results for fire detection; wherein, the flame pattern refers to the changes in the area of the flame region, the changes in the flame color, and the smoke pattern. Based on the status information and combined with the extracted time series, the human body contour information and corresponding color areas are identified to obtain the adjustment results of personnel monitoring; Based on the state information, the object's edge, shape, and preset area are identified to obtain the adjustment result of the channel object. Alarm information is generated based on the adjustment results of fire detection, personnel monitoring, and passageway objects; The sensor information includes thermal imaging information and auxiliary production control system data; determining the status information of the work area based on the video information and sensor information includes: The status information of the work area is determined based on the video information, the thermal imaging information, and the data from the auxiliary production control system. The step of adjusting the response sensitivity of the strong light point, the flame pattern, and the duration information based on the state information to obtain the adjustment result of fire detection includes: Based on the state information, preset rule information for material heat treatment of welding and grinding processes is obtained; Based on the preset rules for material heat treatment in the welding and grinding processes, the video information, the thermal imaging information, and the auxiliary production control system data are matched to obtain fire conclusion information. Based on the fire conclusion information, the response sensitivity of the strong light point, the flame pattern and duration information are adjusted to obtain the adjusted fire detection results.
2. The 5S intelligent monitoring method based on multi-task image recognition according to claim 1, characterized in that, The sensor information includes acoustic information, thermal imaging information, and chemical composition information; Determining the status information of the work area based on the video information and sensor information includes: The status information of the work area is determined based on the video information, the acoustic information, the thermal imaging information, and the chemical composition information. The step of adjusting the response sensitivity of the strong light point, the flame pattern, and the duration information based on the state information to obtain the adjustment result of fire detection includes: Preset rule information is obtained based on the state information; Based on the preset rule information, the video information, the acoustic information, the thermal imaging information, and the chemical composition information are matched to obtain fire conclusion information; the fire conclusion information refers to the conclusion of whether a fire has occurred. Based on the fire conclusion information, the response sensitivity of the strong light point, the flame pattern and duration information are adjusted to obtain the adjusted fire detection results.
3. The 5S intelligent monitoring method based on multi-task image recognition according to claim 1, characterized in that, The color regions include the safety helmet image in the head region and the work clothes and name tag image in the torso region; based on the status information and combined with the extracted time series, the human body contour information and corresponding color regions are identified to obtain the adjustment results of personnel monitoring, including: Based on the state information, contour separation is performed on the human body region in the video information to obtain multiple independent human body contours; For each individual human silhouette, identify the safety helmet image in the head region and the work clothes and name tag image in the torso region; By tracking each individual human body contour using time series data, images of safety helmets and work clothes / name tags are extracted, and the correlation between the safety helmet images, work clothes / name tag images, and human body contours is obtained. By using human posture information, adjacent or partially overlapping human body contours are distinguished, and the adjustment results of personnel monitoring are obtained.
4. The 5S intelligent monitoring method based on multi-task image recognition according to claim 1, characterized in that, The method includes: Based on the video information, detect object images appearing in the channel area to obtain the initial position information and initial shape information of the object images; If, after a preset time has elapsed, the initial position and shape information of the object image remain unchanged, then the object is determined to be a long-term object. Analyze the shape characteristics of the long-term object; Based on the shape characteristics of the long-term object, determine whether the shape characteristics of the long-term object match the preset shape of the items that are allowed to be temporarily placed. If the shape of the long-term object does not match the preset shape of the items that can be temporarily placed, a clutter alarm will be triggered.
5. The 5S intelligent monitoring method based on multi-task image recognition according to claim 1, characterized in that, The method further includes: Set up a preset safety zone within the work area based on video information; Monitor the location information of personnel within the preset safety area; When the time a person stays in the preset safe area reaches a preset threshold, and the location information of the person is still detected within the preset safe area, the activity status of the person is determined. If the activity status of the person is not accessible, an alarm message is generated.
6. The 5S intelligent monitoring method based on multi-task image recognition according to claim 5, characterized in that, The location information includes second location information and third location information; The location information of the monitoring personnel within the preset safety area includes: When there is overlap or occlusion between people, the position information of the people is corrected by using the features of the people who are not occluded and the motion trajectory information in continuous time frames to obtain the second position information; When the preset safety zone is located in a multi-story building, the spatial location information of the personnel is mapped and fused using the personnel location information from different levels of perspective, so as to determine the third location information of the personnel in three-dimensional space.
7. The 5S intelligent monitoring method based on multi-task image recognition according to claim 6, characterized in that, When there is overlap or occlusion between personnel, the position information of the personnel is corrected using the characteristics of the unoccluded personnel and the motion trajectory information in continuous time frames to obtain second position information, including: Extract the motion trajectory information of the personnel in consecutive time frames; Based on the movement trajectory information of the person, identify the person's current movement state; the movement state includes uniform movement, accelerated movement, or decelerated movement. Based on the current movement state of the person, the parameters for predicting the movement trajectory are adjusted to obtain the adjusted movement trajectory prediction parameters; Based on the adjusted motion trajectory prediction parameters, the position information of the person in the time frame is predicted to obtain the predicted position information; By using the characteristics of the person not being obscured and the predicted location information, the location information of the person is corrected to obtain the second location information.
8. A 5S intelligent monitoring method based on multi-task image recognition according to claim 6, characterized in that, The method further includes: Acquire visual, depth, and thermal information corresponding to the features of the unobstructed parts of a person; The visual information, depth information, and thermal imaging information are fused to obtain enhanced personnel feature information; Based on the enhanced personnel feature information and the personnel's motion trajectory information in continuous time frames, the personnel's position information is corrected to obtain second position information.
9. A 5S intelligent monitoring system based on multi-task image recognition, characterized in that, The system includes: The information acquisition module is used to acquire video information and sensor information of the work area where visual interference sources exist; the visual interference sources include image interference sources generated by welding and grinding processes. The status information recognition module is used to determine the status information of the work area based on the video information and sensor information; the status information refers to the parameters corresponding to the work mode determined by the visual interference image of the work area. The fire detection adjustment module is used to adjust the response sensitivity of the strong light point, the flame pattern and duration information according to the status information to obtain the adjustment result of fire detection; wherein, the flame pattern refers to the change in the area of the flame area, the change in the flame color and the smoke pattern. The personnel monitoring adjustment module is used to identify human body contour information and corresponding color areas based on the status information and the extracted time series, and obtain the personnel monitoring adjustment result. The channel object adjustment module is used to identify the object edge, object shape, and preset area of the object based on the state information, and obtain the adjustment result of the channel object. The alarm generation module is used to generate alarm information based on the adjustment results of fire detection, personnel monitoring, and passageway objects. The sensor information includes thermal imaging information and auxiliary production control system data; determining the status information of the work area based on the video information and sensor information includes: The status information of the work area is determined based on the video information, the thermal imaging information, and the data from the auxiliary production control system. The step of adjusting the response sensitivity of the strong light point, the flame pattern, and the duration information based on the state information to obtain the adjustment result of fire detection includes: Based on the state information, preset rule information for material heat treatment of welding and grinding processes is obtained; Based on the preset rules for material heat treatment in the welding and grinding processes, the video information, the thermal imaging information, and the auxiliary production control system data are matched to obtain fire conclusion information. Based on the fire conclusion information, the response sensitivity of the strong light point, the flame pattern and duration information are adjusted to obtain the adjusted fire detection results.