Personnel-perceived speed adaptive control method

By constructing a closed-loop interaction mechanism for the entire process through manifold regularized nonnegative matrix factorization and maximum entropy deep reinforcement learning decision model, the problem of fixed speed regulation mode and insufficient safety redundancy of passenger transport equipment is solved. It realizes full-scenario adaptive matching and multi-objective collaborative optimization, and improves the safety, efficiency and intelligence level of the equipment.

CN122018326AInactive Publication Date: 2026-05-12SICHUAN JINGZHUN SPECIAL EQUIP INSPECTION CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SICHUAN JINGZHUN SPECIAL EQUIP INSPECTION CO LTD
Filing Date
2026-04-09
Publication Date
2026-05-12
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing speed control technologies for passenger transport equipment suffer from fixed and singular speed regulation modes, insufficient depth and dimension of human perception, inability to adapt to changes in passenger flow load, insufficient safety redundancy, lack of multi-objective collaborative capability of control models, susceptibility to local optima, and lack of a closed-loop interaction mechanism throughout the entire process, resulting in low safety and efficiency.

Method used

The algorithm employs manifold regularized nonnegative matrix factorization to extract personnel perception features. Combined with a maximum entropy deep reinforcement learning decision model, a closed-loop interaction mechanism is constructed to generate low-dimensional state vectors and risk quantification indicators. Speed ​​decision constraints are dynamically generated, and iterative optimization is performed using Gaussian process regression anomaly detection to output optimized speed control commands.

Benefits of technology

It achieves full-scenario adaptive matching of conveying equipment, improves safety performance, realizes multi-objective collaborative optimization, reduces energy waste, extends equipment life, and improves the riding experience and system intelligence level.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122018326A_ABST
    Figure CN122018326A_ABST
Patent Text Reader

Abstract

The invention provides a personnel perception speed adaptive control method, and belongs to the technical field of intelligent adaptive control and personnel perception cross, and the method comprises the steps: S1, multi-modal full-dimension data collection and space-time standardization preprocessing; s2, personnel perception feature manifold embedding and risk index generation; s3, dynamic constraint boundary and multi-target weight prior generation; s4, carrying out the initial speed decision of the maximum entropy deep reinforcement learning with constraints; s5, calculating and feeding back a quaternary weighted dynamic reward function; s6, three-way interactive iterative optimization and optimal speed output are carried out; and S7, generating and executing a smooth speed control curve. According to the invention, multi-objective collaborative optimization of safety priority, energy conservation, high efficiency, comfortable experience and equipment life extension is finally achieved, and intelligent control technologies of various passenger conveying equipment are promoted to develop to a higher safety level and a higher intelligent level.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent adaptive control and human perception technology, specifically to a speed adaptive control method based on human perception. Background Technology

[0002] Passenger transport equipment is a core transportation facility in densely populated scenarios such as public transportation, commercial complexes, cultural and tourism scenic spots, transportation hubs, and industrial and mining areas. The appropriateness of its operating speed directly determines the equipment's operational safety, energy efficiency, passenger experience, and the lifespan of core components. Currently, speed control technology for various passenger transport equipment still suffers from many common industry pain points and technical deficiencies:

[0003] Speed ​​control modes are rigid and simplistic, with a severe lack of depth and dimension in human perception. The vast majority of passenger transport equipment operates at a fixed rated speed, which cannot adapt to dynamic changes in passenger flow, load, and passenger structure. During peak hours, this can easily lead to overcrowding and inefficient evacuation, while during off-peak hours, it results in significant energy waste and ineffective equipment wear and tear. The few devices with basic speed control functions rely solely on real-time passenger numbers and load weight as the sole basis for speed adjustment, failing to deeply integrate refined human perception characteristics such as passenger age, mobility, safe distance, and behavioral status. This makes it impossible to make adaptive speed adjustments for high-risk passenger groups, resulting in insufficient safety redundancy and extremely poor speed control accuracy and scenario adaptability.

[0004] Feature processing is disconnected from the decision-making process, and multiple algorithms operate in isolation without collaboration. In existing technologies, modules such as personnel perception data processing, speed decision-making, anomaly detection, and safety control are mostly designed separately and operate independently, without building a deep linkage interaction mechanism, resulting in serious data silos. The feature extraction process is not deeply bound to subsequent decision-making objectives, and key features related to personnel safety are easily lost during dimensionality reduction processing. Anomaly detection only has a post-event alarm function and does not integrate the detection results back into the decision model optimization, failing to form a complete closed loop of "perception-feature-decision-execution-feedback-optimization".

[0005] The control model lacks multi-objective coordination capabilities and is prone to getting trapped in local optima. Existing speed control logic is often heavily biased towards a single optimization objective. It either unilaterally pursues maximum energy saving and continuously reduces speed, resulting in excessively long passenger dwell time and a poor experience; or it excessively pursues carrying efficiency and operates at high speed for extended periods, ignoring the safety risks to high-risk passengers and the wear and tear on core equipment components. At the same time, conventional reinforcement learning control models suffer from an imbalance between exploration and utilization, easily converging to local optimal strategies, and lacking generalization ability and robustness under complex operating conditions. Summary of the Invention

[0006] This invention provides a speed adaptive control method based on human perception. Through in-depth representation of the core characteristics of human perception, in-depth integration and innovation of intelligent algorithms, and the construction of a closed-loop interaction mechanism throughout the entire process, it achieves full-scenario adaptive matching of the operating speed of the conveyor equipment to human characteristics, equipment status, and environmental conditions. Ultimately, it achieves multi-objective collaborative optimization that prioritizes safety, energy efficiency, user comfort, and equipment lifespan extension, thereby promoting the development of intelligent control technology for various passenger conveyor equipment towards higher safety levels and higher levels of intelligence.

[0007] A human-perceived speed adaptive control method includes:

[0008] S1. Collect personnel perception data, equipment operation data and environmental condition data of the target conveying equipment, and perform spatiotemporal alignment processing and standardization preprocessing to obtain a time-synchronized standardized multi-source dataset;

[0009] S2. The manifold regularized nonnegative matrix factorization algorithm is used to extract features and reduce dimensions of the standardized multi-source dataset to generate low-dimensional state vectors and a set of personnel risk quantification indicators.

[0010] S3. Based on the personnel risk quantification index set, complete the personnel risk level determination. According to the personnel risk level and the inherent rated parameters of the target conveying equipment, dynamically generate the hard constraint range of speed decision, the speed change rate constraint threshold and the multi-objective weight prior parameters.

[0011] S4. Construct a maximum entropy deep reinforcement learning decision model that integrates manifold regularization constraints. Input the low-dimensional state vector, hard constraint interval, and velocity change rate constraint threshold into the maximum entropy deep reinforcement learning decision model, and output the initial optimal target running speed within the constraint feasible region.

[0012] S5. Based on multi-objective weighted prior parameters, combined with personnel perception data, equipment operation data, environmental condition data and initial optimal target operating speed, a quaternary weighted reward function covering safety, energy saving, riding experience and equipment protection is constructed to calculate the real-time total reward value;

[0013] S6. The Gaussian process regression anomaly detection algorithm is adopted to complete the anomaly detection of the whole working condition based on the low-dimensional state vector and real-time operation monitoring data, and generate anomaly detection results. Combining the real-time total reward value, anomaly detection results and hard constraint interval, the manifold regularized non-negative matrix factorization algorithm, the maximum entropy deep reinforcement learning decision model and the Gaussian process regression anomaly detection algorithm are subjected to bidirectional interactive iterative update and parameter optimization, and the optimized optimal target running speed is output.

[0014] S7. Based on the optimized target operating speed and speed change rate constraint threshold, generate a smooth speed control curve, send speed control commands to the controller of the target conveying equipment, and complete speed control.

[0015] In this manual, the specific process of spatiotemporal alignment processing is as follows: using the unified clock of the acquisition system as a reference, millisecond-level timestamps are added to personnel perception data, equipment operation data, and environmental condition data respectively; the three types of data under the same timestamp are matched and aligned one by one, and invalid data with time deviations exceeding twice the acquisition cycle are removed; missing data are completed by linear interpolation using the average of historical operation data under the same equipment type and the same operating conditions; the standardization preprocessing adopts the min-max normalization method to uniformly map all data to the 0-1 interval, eliminating the difference in the dimensions of data of different dimensions.

[0016] In this specification, when using the manifold regularized nonnegative matrix factorization algorithm, the decision-related feature constraints output from S6 are received simultaneously. These constraints are generated by the policy gradient derivation after iterative updates of the maximum entropy deep reinforcement learning decision model in S6. The manifold regularization weight coefficients of the manifold regularized nonnegative matrix factorization algorithm are adjusted in conjunction with the decision-related feature constraints. The basis matrix and coefficient matrix in the manifold regularized nonnegative matrix factorization algorithm are then coordinated and fine-tuned through a multiplicative iterative update rule. This ensures that the generated low-dimensional state vectors preferentially retain the core risk characteristics of personnel and the strong correlation characteristics of speed decisions, thereby improving the relevance of feature representation.

[0017] In this manual, when generating the set of personnel risk quantification indicators, based on the passenger age, action status, safety protection compliance status, and adjacent distance data in the personnel perception data collected by S1, four types of core indicators are extracted: high-risk passenger ratio, passenger average safe distance deviation rate, safety protection device compliance rate, and passenger unstable state ratio. The high-risk passenger ratio is the sum of the proportions of elderly passengers, child passengers, and passengers with mobility impairments. All four types of core indicators are mapped to the 0-1 range through min-max normalization processing, with higher values ​​representing higher corresponding risks.

[0018] In this specification, the policy optimization objective function of the maximum entropy deep reinforcement learning decision model simultaneously introduces an entropy regularization term and a constraint violation penalty term. The entropy regularization term maintains the model's exploration capability by calculating the entropy value of the policy distribution, thus preventing the policy from getting trapped in local optima. The constraint violation penalty term applies a linear penalty to candidate speed actions that exceed the hard constraint range generated by S3 or the speed change rate constraint threshold. The penalty coefficient is positively correlated with the degree of constraint violation, ensuring that the output initial optimal target running speed strictly meets the constraint requirements.

[0019] In this specification, after the maximum entropy deep reinforcement learning decision model outputs the initial optimal target running speed, an additional constraint verification process is performed: the current actual running speed of the target conveying equipment is obtained through the equipment operation data collected by S1, and the difference between the initial optimal target running speed and the current actual running speed is calculated; if the difference exceeds the speed change rate constraint threshold generated by S3, then among all candidate speed actions that meet the speed change rate constraint threshold, the speed with the second best Q value is selected as the adjusted initial optimal target running speed to ensure that there are no stutters or impacts in the speed adjustment process.

[0020] In this manual, the dynamic adjustment process of the weights of the quaternary weighted reward function is as follows: If the personnel risk quantification indicators generated by S2 show that the proportion of high-risk passengers exceeds 40%, or if the environmental operating condition data collected by S1 shows that there is severe weather and / or the power grid voltage fluctuation exceeds the preset range, then the total proportion of the safety reward weight and the equipment protection reward weight will be increased to no less than 70%; if the equipment operation data collected by S1 shows that the wear coefficient of the core components is less than 0.2 and the operating condition is stable, then the energy-saving reward weight and the riding experience reward weight will each be increased by 5%-10%, and all weights will be normalized to keep the sum of 1 after adjustment.

[0021] In this specification, when the Gaussian process regression anomaly detection algorithm is running, it uses the low-dimensional state vector output by S2 as the core input feature, and simultaneously receives the initial optimal target running speed output by S4. Based on the initial optimal target running speed and the rated speed in the inherent rated parameters of the target conveying equipment, the speed ratio is calculated, and the anomaly detection threshold is dynamically adjusted according to the speed ratio. The higher the speed ratio, the lower the anomaly detection threshold.

[0022] In this specification, the specific process of bidirectional interactive iterative update is as follows: First, the real-time total reward value calculated by S5 is used as the core feedback signal for updating the parameters of the maximum entropy deep reinforcement learning decision model, and the anomaly detection result is used as a hard constraint for policy adjustment. The network parameters of the decision model are updated through the proximal policy optimization algorithm. Second, the optimal feature constraint is derived in reverse based on the policy gradient updated by the maximum entropy deep reinforcement learning decision model. The optimal feature constraint is fed back to S2 to adjust the objective function of the manifold regularized nonnegative matrix factorization algorithm and optimize the feature extraction direction. Third, the anomaly sample data marked by the anomaly detection result is added to the training set of the Gaussian process regression anomaly detection algorithm. The maximum likelihood estimation method is used to complete the incremental optimization of the algorithm hyperparameters and improve the detection accuracy of similar anomaly conditions.

[0023] In this specification, for different types of target conveying equipment, it is only necessary to adapt to the inherent rated parameters and constraint threshold benchmarks of the corresponding equipment, without reconstructing the core architectures of the manifold regularization non-negative matrix factorization algorithm, the maximum entropy deep reinforcement learning decision model, and the Gaussian process regression anomaly detection algorithm, and the implementation method can be adapted and applied to various target conveying equipment; the inherent rated parameters include rated speed, rated acceleration, rated load, and the tolerance threshold of core components.

[0024] This specification can at least achieve the following beneficial effects: Significantly improved safety performance: Deeply integrating refined personnel perception features, through pre-set safety constraints, full-condition anomaly detection, and real-time linkage disposal, realize the pre-control of personnel safety risks, greatly reduce risks such as crowded people falling and equipment overload failure, and improve the safety redundancy and compliance of equipment operation. Multi-objective collaborative optimization: Break through the limitation of single-objective preference in traditional control schemes, and can adaptively balance the four core objectives of safety guarantee, energy conservation, riding experience, and equipment protection based on real-time working conditions, so as to achieve the optimal comprehensive operation performance of the equipment. Excellent working condition adaptability: Through the deep integration of core algorithms and the full-process closed-loop design, it can be adapted to various complex scenarios such as sudden changes in passenger flow, changes in personnel structure, equipment deterioration, and environmental fluctuations in real time, with strong model generalization and operation robustness. Energy conservation and consumption reduction and equipment life extension: On the premise of ensuring safety and transportation efficiency, intelligently optimize the operating speed based on real-time working conditions to avoid ineffective high-power consumption operation; at the same time, reduce the mechanical losses caused by equipment overload and start-stop impacts, extend the service life of the whole equipment and core components, and reduce the operation and maintenance costs in the whole life cycle. Comprehensively optimized riding experience: Deeply integrate personnel perception features into the entire decision-making process, adaptively adjust the operating speed and acceleration and deceleration characteristics for different passenger groups, with smooth speed regulation and no jerks, taking into account both transportation efficiency and riding comfort. Strong universality in all scenarios: It is a general intelligent control method, without reconstructing the core architecture, and only needs to adapt the corresponding parameters to cover various application scenarios such as escalators, passenger ropeways, passenger amusement facilities, and industrial and mining passenger equipment, with high adaptability and promotion value. Full-link closed-loop collaborative evolution: Break the problem of isolated information islands of traditional technical modules, realize the two-way interaction and collaborative optimization of the three core links of feature extraction, decision optimization, and anomaly detection, and continuously improve the intelligent level and long-term operation stability of the system. Brief Description of the Drawings

[0025] Figure 1 It is a schematic diagram of the speed adaptive control method for personnel perception.

[0026] Figure 2 It is a schematic diagram of the perception-feature-constraint generation process.

[0027] Figure 3 It is a schematic diagram of the decision-evaluation-three-way iterative optimization process.

[0028] Figure 4Schematic diagram of the execution-monitoring-closed-loop feedback process. Detailed implementation manners

[0029] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0030] As Figure 1 shown, this embodiment provides a speed adaptive control method for personnel perception, which is applicable to conveying equipment with personnel carrying functions, including but not limited to escalators, moving walkways, passenger ropeways, passenger cable cars, large-scale amusement facilities for humans, passenger conveying equipment supporting rail transit, outdoor passenger carrying equipment in scenic spots, passenger equipment in aviation and water transportation passenger hubs, special passenger equipment in theme parks, passenger equipment supporting commercial super high-rise buildings, special barrier-free passenger conveying equipment, passenger conveying equipment in special industrial and mining areas, and special passenger conveying equipment for emergency rescue. The overall process refers to Figure 2 、 Figure 3 and Figure 4 .

[0031] S1. Multi-modal personnel-equipment-environment full-dimensional spatio-temporal synchronization data acquisition and spatio-temporal alignment preprocessing:

[0032] The core purpose of this step is to obtain the full amount of original data in the three core dimensions of personnel perception, equipment operation, and environmental conditions, solve the problems of asynchronous multi-source data sampling, inconsistent dimensions, and time series misalignment, and provide a precise and time series-aligned data source for subsequent full-process processing. The data sources of this step include various types of perception sensors supporting the target conveying equipment, the built-in data acquisition module of the equipment controller, the equipment factory technical documents, and the latest equipment operation and maintenance records.

[0033] Refer to Figure 2First, comprehensive data collection is performed, with all data collection frequencies uniformly set at 100 milliseconds per data point. For high-speed, heavy-duty equipment, this frequency can be dynamically increased to 50 milliseconds per data point based on operational characteristics. The collected content comprehensively covers four categories, as follows: The first category is basic personnel perception data, including real-time forward passenger volume, real-time reverse passenger volume, queue length at equipment entrances, evacuation and disembarkation speed at equipment exits, and real-time passenger count in enclosed cabin equipment. The second category is personnel perception characteristic data, including the proportion of passengers carrying large luggage, the proportion of elderly passengers, the proportion of children, the proportion of passengers with mobility impairments, the proportion of passengers in wheelchairs and strollers, the proportion of able-bodied adults, and the proportion of passengers wearing safety protective equipment in enclosed cabin equipment. The data includes: 1) Wearing rate, real-time data on passenger spacing, and passenger standing stability data; 2) Equipment operating status data, including current actual operating speed, real-time load rate of the equipment power system, wear coefficient of core load-bearing components, trigger count and trigger type of equipment safety protection devices, braking system status parameters, stress values ​​of core structures of heavy-duty high-speed equipment, real-time temperature of motor windings, actual energy consumption data, and real-time vibration data of core components; 3) Environmental operating condition data, including ambient temperature, ambient humidity, real-time fluctuation of grid voltage, lighting brightness of the area where the equipment is located, real-time wind force level, rain and snow intensity, ambient visibility, and slippery ground conditions of indoor equipment.

[0034] The hardware layout and acquisition method for data collection are as follows: Basic data and characteristic data of personnel perception are collected through binocular vision sensors. For open conveyor equipment, one binocular vision sensor is installed at the top of the entrance, the top of the exit, and on both sides of the middle of the equipment, for a total of four binocular vision sensors. For enclosed cabin equipment, one binocular vision sensor is installed at the entrance and exit of the equipment, and one is installed inside each cabin. Additional sensors are installed at key points along the equipment's operating route as needed. The collected image data is transmitted to the edge computing module and processed by a pre-trained image recognition algorithm to obtain the raw data. Passenger age is obtained through a facial feature recognition algorithm based on a convolutional neural network, passenger luggage status is obtained through an object recognition algorithm based on contour features, and passenger adjacent spacing and standing stability are obtained through a binocular vision three-dimensional ranging algorithm. In the equipment operation status data, the current actual operating speed is collected by an incremental encoder installed at the drive shaft end of the equipment's power system; the real-time load rate of the equipment's power system is collected by a Hall current sensor built into the motor controller, and calculated by dividing the actual current by the rated current and then multiplying by 100%; the wear coefficient of the core load-bearing components of the equipment is collected by corresponding dedicated sensors. For escalators and moving walkways, the force uniformity data is collected by eight pressure sensors evenly distributed on the bottom of the steps for calculation; for passenger ropeways and cable cars, the data is collected by steel cable tension and wear detection sensors; and for large amusement facilities for human passengers, the data is collected by guide rail wear sensors and structural stress sensors. Other equipment operation status data are directly collected and read by the equipment controller and its supporting dedicated sensors. In the environmental condition data, the ambient temperature and humidity are collected by a temperature and humidity sensor installed 1.5 meters above the ground next to the equipment; the real-time fluctuation value of the power grid voltage is collected by a voltage sensor connected in series in the equipment's power supply circuit; the lighting brightness is collected by a brightness sensor installed 2 meters above the ground on the left side of the equipment entrance; outdoor equipment is equipped with wind speed and direction sensors, rain and snow sensors, and visibility meters to collect corresponding environmental data; and the indoor ground slipperiness status is collected by a visual sensor or a ground humidity sensor.

[0035] After the raw data acquisition is completed, spatiotemporal alignment and standardization preprocessing are performed. First, spatiotemporal alignment is performed by adding millisecond-level timestamps to all raw data, using the unified clock of the acquisition system as a reference. Personnel perception data, equipment operating status data, and environmental condition data at the same timestamp are aligned and matched to form a time-synchronized multi-source data set. Invalid data with timestamp deviations exceeding twice the acquisition cycle are removed. Missing data is filled using linear interpolation based on the historical average of the same equipment type and operating conditions. Next, standardization preprocessing is performed. The time-synchronized multi-source data sets are grouped by data type and subjected to min-max normalization to eliminate dimensional differences between different dimensions of data. After normalization, the value range of all data is uniformly mapped to the 0-1 interval. Data such as the wear coefficient of core equipment components and the proportion of various passenger types, which are inherently in the 0-1 interval, only undergo outlier truncation. Values ​​exceeding the 0-1 interval are corrected according to the interval boundary values.

[0036] In some embodiments, the component wear coefficient is calculated by weighted fusion of multi-dimensional features. The wear coefficient ranges from 0 to 1, with higher values ​​indicating more severe wear on the core component. The calculation formulas for the wear coefficient of different types of target conveying equipment are differentiated.

[0037] For open conveyor systems such as escalators and moving walkways, the core components are the step / pedal chain, main drive shaft, and step rollers. The formula for calculating the wear coefficient is as follows: ;in, The normalized value for chain wear is calculated based on the ratio of the chain tension deviation value to the rated tension threshold, and the data is collected by a chain tension sensor. The normalized value for wear of the step rollers is calculated based on the force unevenness collected by the pressure sensor at the bottom of the pedal. The higher the force unevenness, the higher the corresponding value. The normalized wear value of the main drive shaft is calculated based on the ratio of the peak vibration acceleration of the drive shaft to the rated vibration threshold. The normalized value of structural stress is calculated based on the ratio of stress sensor data collected from key load-bearing structures to the rated stress threshold; all sub-parameters are mapped to the 0~1 range through min-max normalization.

[0038] For enclosed cabin equipment such as passenger ropeways and cable cars, the core components are the load-bearing steel cable, drive wheel, and cable grip. The formula for calculating the wear coefficient is as follows: ;in, The normalized value for steel cable wear is calculated by weighting the ratio of steel cable diameter wear, number of broken wires, and national standard allowable threshold. The normalized wear value of the cable gripper is calculated based on the ratio of the wear amount of the cable gripper jaws to the rated wear threshold. The normalized value for drive wheel liner wear is calculated based on the ratio of liner wear thickness to the rated thickness threshold; all sub-parameters are mapped to the 0~1 range through min-max normalization.

[0039] For large-scale amusement rides for human passengers, the core components are guide rails, wheels, and load-bearing structural parts. The formula for calculating the wear coefficient is as follows: ;in, The normalized value for guide rail wear is calculated based on the ratio of the wear amount on the side and top surfaces of the guide rail to the design allowable threshold. The normalized value for wheel wear is calculated based on the ratio of wheel diameter wear, radial runout, and rated threshold. The normalized wear value for load-bearing structural components is calculated based on the ratio of structural stress monitoring data, cumulative fatigue damage value, and design allowable threshold; all sub-parameters are mapped to the 0~1 range through min-max normalization.

[0040] In some embodiments, when the wear coefficient of the core component is less than 0.2, it is determined that the core component is in a state of slight wear and stable operation; when the wear coefficient is in the range of 0.2 to 0.8, it is determined to be in a state of normal wear; when the wear coefficient is greater than 0.8, it is determined to be in a state of severe wear, triggering the equipment protection weight adjustment mechanism.

[0041] In some embodiments, the passenger instability is determined by extracting key skeletal data from video sequences acquired by a binocular vision sensor. These key skeletal data include at least the head, shoulders, hips, knees, and ankles. Specific rules for determining passenger instability include: First category, standing instability: For standing passengers, if a horizontal offset of the hip key point exceeds 10cm for 5 consecutive frames or more, or if the angle between the upper torso and the vertical direction is consistently greater than 15°, it is determined to be a standing instability state. Second category, tilting tendency due to center of gravity imbalance: If a tilting tendency is detected for 3 consecutive frames or more, the center of gravity projection point exceeds... The first category is unbalanced and prone to falling, where the range of support for both feet or the torso tilt angle is greater than 30° and continues to increase. The second category is unsteady walking, where passengers walking on the equipment are detected to have a gait frequency fluctuation of more than 50% or a step length deviation of more than 20cm for 4 or more consecutive frames, accompanied by significant torso swaying. The third category is unsteady walking, where passengers sitting in enclosed cabins are detected to have their body deviate from the normal sitting position by more than 20cm for 5 or more consecutive frames, or significant abnormal displacement of the body restraint parts corresponding to the safety protection device.

[0042] In some embodiments, the percentage of passengers in unstable states is the ratio of the number of passengers identified as being in unstable states to the total number of passengers in the same captured image at the same time. This ratio is mapped to the 0~1 range through min-max normalization and incorporated into the personnel risk quantification index set.

[0043] In some embodiments, the abnormal passenger behavior is identified and defined based on image data collected by a binocular vision sensor. Specifically, it includes the following high-risk abnormal behavior types, covering core scenarios for personnel safety management: First, reverse movement: For unidirectional moving equipment such as escalators and moving walkways, if a passenger's direction of travel is opposite to the equipment's direction of travel, and the continuous movement distance exceeds 1 meter, it is defined as reverse movement. Second, climbing / crossing behavior: If a passenger is detected climbing the equipment's handrails, side panels, guardrails, or cabin edges, or crossing the equipment's safety protection structure, regardless of the duration, it is defined as climbing / crossing behavior. Third, falling / lying down behavior: If a passenger is detected falling / lying down while not in a sitting position... The following are considered abnormal behaviors: First, falling and lying down: An angle of less than 30° between the passenger and the ground / equipment contact surface, or the passenger lying down inside the equipment steps, pedals, or cabin. Second, pushing and shoving: Passengers are detected to be less than 0.2m apart, accompanied by limb collisions and rapid displacement, and this continues for 3 or more consecutive frames. Third, dangerous riding behavior: For escalators, passengers are detected pushing strollers, wheelchairs, or large luggage, or children running, jumping, or riding in the opposite direction on the steps. For enclosed cabin equipment, passengers are detected unauthorized removal of safety devices or extending body parts outside the cabin.

[0044] In some embodiments, each frame of the above-mentioned abnormal behavior is identified as a security anomaly event, and the corresponding reward is deducted in the calculation of the security reward item; if the abnormal behavior of the same passenger continues for more than 10 frames, or if abnormal behavior of 3 or more people is detected at the same time, the mechanism of lowering the anomaly detection threshold and tightening the speed constraint range is triggered.

[0045] After this step is completed, a spatiotemporally aligned and standardized multi-source dataset is generated, which will be used directly as the basic data source for the subsequent S2 feature processing step.

[0046] S2. Manifold embedding of personnel-perceived features and construction of low-dimensional state vectors based on manifold regularized nonnegative matrix factorization:

[0047] The core objective of this step is to use personnel perception characteristics as the central anchor point and employ a manifold-regularized nonnegative matrix factorization (MDF) algorithm to perform hierarchical feature extraction, dimensionality reduction, and manifold structure preservation on a high-dimensional spatiotemporally aligned and standardized multi-source dataset. This addresses the problems of loss of core personnel safety features, destruction of local neighborhood structures, and high feature redundancy in conventional dimensionality reduction algorithms. The output is a low-dimensional feature carrier that can be directly used for reinforcement learning decision-making and risk quantification. Manifold-regularized nonnegative matrix factorization is a niche algorithm in this field. Its core advantage lies in simultaneously preserving the nonnegativity of the data and the local manifold structure, making the features more physically interpretable. It can also accurately capture the inherent geometric structure of personnel-equipment-environment coupled data, avoiding the loss of key risk features caused by conventional dimensionality reduction algorithms, and significantly improving the safety and accuracy of subsequent decisions. The input to this step is the spatiotemporally aligned and standardized multi-source dataset output by S1, and the processing consists of four core stages: model building, model training, model application, and construction of a personnel risk quantification index set.

[0048] refer to Figure 2 First, we construct a model for the manifold regularized nonnegative matrix factorization algorithm. Let the spatiotemporally aligned and normalized multi-source dataset output by S1 be a matrix. The matrix dimension is ,in This represents the total feature dimension of the original data. Represents the number of data samples, matrix Each column in the matrix corresponds to a multi-source data sample at a given timestamp, and each row corresponds to a feature in one dimension. The core objective of manifold regularized nonnegative matrix factorization is to transform the matrix... It can be decomposed into the product of two nonnegative matrices, which are the basis matrices. sum coefficient matrix , where the basis matrix The dimension is coefficient matrix The dimension is , Let represent the lower-dimensional feature dimension after dimensionality reduction, and satisfy . At the same time, a manifold regularization term is introduced to preserve the local neighborhood structure of the data, ensuring that local features related to personnel safety are not lost.

[0049] To achieve the above objectives, an objective function with manifold regularization is constructed, as shown in the following formula: ; The Frobenius norm of a matrix is ​​used to measure the reconstruction error between the original and reconstructed data. The trace operation represents the sum of the elements on the main diagonal of a matrix. This represents the Laplacian matrix, used to characterize the local manifold structure among data samples; This represents the manifold regularization weight coefficient, used to balance reconstruction error and manifold structure preservation, and has a fixed value of 0.1. This represents the Frobenius norm regularization weight coefficient, used to prevent model overfitting, and has a fixed value of 0.01.

[0050] Laplace matrix The construction process is as follows: First, construct the adjacency matrix. The adjacency matrix has a dimension of 1. The K-nearest neighbors algorithm is used to determine the nearest neighbors of each sample point. The number of nearest neighbors is fixed at 10. Only when two sample points are each other's nearest neighbors are a connection edge established between them in the adjacency matrix. The edge weight is calculated using the hot kernel function, as shown in the following formula:

[0051] If the sample With sample If they are close neighbors, then ;otherwise ;

[0052] Representing the adjacency matrix The element in the i-th row and j-th column represents the edge weight between sample i and sample j. and Let represent the feature vectors of the i-th and j-th samples, respectively; The L2 norm of a vector is used to calculate the Euclidean distance between two samples. The bandwidth parameter represents the hot kernel function, and its value is the mean distance between all samples and the feature vector.

[0053] After constructing the adjacency matrix, construct the degree matrix. The dimension of the degree matrix is is a diagonal matrix, whose diagonal elements are the sum of the elements of the corresponding rows of the adjacent matrices, as shown in the following formula:

[0054] ; in the formula Degree matrix The diagonal element in the i-th row and i-th column.

[0055] Finally, calculate the Laplace matrix. The formula is as follows: After model construction is complete, model training is performed using the manifold regularized nonnegative matrix factorization algorithm. The core objective of model training is to minimize the objective function. While maintaining the basis matrix sum coefficient matrix To ensure the non-negativity of matrix elements, a multiplicative iterative update algorithm is used. This algorithm naturally preserves the non-negativity of matrix elements during iteration, eliminating the need for additional non-negativity constraints and offering superior computational efficiency and stability.

[0056] basis matrix The multiplicative iterative update rule is as follows: ; Representation of basis matrix The element in the d-th row and k-th column of the array; This represents the element in the d-th row and k-th column of the corresponding matrix.

[0057] coefficient matrix The multiplicative iterative update rule is as follows: ; Represents the coefficient matrix The element in the k-th row and n-th column; This represents the element in the k-th row and n-th column of the corresponding matrix.

[0058] The complete execution flow of model training is as follows: Step 1, initialize the basis matrix. sum coefficient matrix The first step is to use a random non-negative initialization method to ensure that all elements in the matrix are positive; the second step is to fix the coefficient matrix. According to the basis matrix The update rule updates the basis matrix; the third step is to fix the basis matrix. According to the coefficient matrix The fourth step is to update the coefficient matrix according to the update rules; the updated objective function is then calculated. The fifth step is to repeat steps two through four until the change in the objective function value is less than the preset threshold. If the number of iterations reaches a preset limit of 500, the iteration stops, and the optimal basis matrix after training is obtained. and the optimal coefficient matrix .

[0059] After model training is complete, the model application using the manifold regularized nonnegative matrix factorization algorithm is executed. The core of the model application is mapping the high-dimensional data samples output by S1 in real-time to a low-dimensional feature space to obtain a low-dimensional state vector. The specific execution process is as follows: First, the current-time samples of the spatiotemporally aligned and normalized multi-source dataset output by S1 in real-time are processed. As input, the sample dimension is Then, based on the optimal basis matrix after training... Solve for the low-dimensional coefficient vector corresponding to the sample at the current time. The coefficient vector dimension is The solution process employs the same multiplicative iterative update rule as the training phase, updating only the coefficient vector while keeping the basis matrix fixed. The iteration count is 100 to ensure rapid convergence. Finally, the obtained low-dimensional coefficient vector is... The low-dimensional state vector output in this step Among them, the low-dimensional feature dimension The value is fixed at 16 to balance feature representation capability and computational efficiency of subsequent algorithms.

[0060] The final step in this process involves constructing a set of risk quantification indicators for personnel, based on the optimal coefficient matrix obtained after training. and optimal basis matrix By combining the perception characteristics of people in the original data, a set of quantitative indicators for personnel risk was calculated, including the proportion of high-risk passengers, the deviation rate of the average safe distance of passengers, the level of personnel density, the compliance rate of safety protection devices, and the proportion of passengers in unstable states. All indicators were normalized to the range of 0 to 1, and the higher the value, the higher the personnel safety risk.

[0061] After this step is completed, the output low-dimensional state vector directly serves as the core input for the subsequent S3 constraint space generation, S4 velocity decision, and S6 anomaly detection stages; the output personnel risk quantification index set directly serves as the core input for the subsequent S3 constraint space generation and S5 reward function calculation stages; the output optimal basis matrix... It is directly used for algorithm interaction and feature optimization in subsequent S4 and S6 stages.

[0062] S3. Dynamic constraint space generation and adaptive calibration of action boundaries based on personnel risk levels:

[0063] The core objective of this step is to use the previously output set of personnel risk quantification indicators as the core driver to generate hard and soft constraint boundaries for speed decisions in advance. At the same time, it calibrates the weighted prior parameters for multi-objective optimization, solving the problems of decision overshooting, insufficient safety redundancy, and multi-objective imbalance caused by decision-making followed by verification in conventional technologies. This delineates a compliant and optimal feasible domain for subsequent speed decisions, fundamentally preventing decision results from violating safety regulations and equipment operation constraints.

[0064] The inputs for this step include the low-dimensional state vector output by S2, the set of personnel risk quantification indicators, and the inherent rated parameters of the target conveying equipment. The inherent rated parameters of the equipment are derived from the equipment's factory technical documents and the latest operation and maintenance records, including the equipment's rated operating speed, rated acceleration, rated load, the upper and lower speed limits specified by safety standards, and the design service life parameters of core components.

[0065] This step involves four core processes: personnel risk level classification and determination, dynamic speed constraint interval calibration, speed change rate constraint threshold calibration, and multi-objective weight prior parameter calibration.

[0066] refer to Figure 2First, personnel risk level classification is performed. Based on three core indicators in the personnel risk quantification index set—the proportion of high-risk passengers, the deviation rate of average safe distance between passengers, and the personnel density level—a weighted summation method is used to calculate the comprehensive personnel risk value, as shown in the following formula: ; The comprehensive personnel risk value ranges from 0 to 1; The percentage of high-risk passengers is the sum of the percentages of elderly passengers, child passengers, and passengers with mobility impairments. The average passenger safety distance deviation rate; This represents the population density level.

[0067] Based on the magnitude of the comprehensive personnel risk value, personnel risk levels are divided into five levels as follows: Level 1 (Level I, Low Risk): The comprehensive personnel risk value ranges from 0 to 0.2, with no significant safety risk. Priority should be given to balancing transport efficiency, energy saving, passenger experience, and equipment protection. Level 2 (Level II, Lower Risk): The comprehensive personnel risk value ranges from greater than 0.2 to less than or equal to 0.4, with slight safety risks. Safety should be the primary focus, while other optimization objectives should be considered. Level 3 (Level III, Medium Risk): The comprehensive personnel risk value ranges from greater than 0.4 to less than or equal to 0.6, with moderate safety risks. Safety and passenger experience should be prioritized, while other objectives should be considered to a certain extent. Level 4 (Level IV, Higher Risk): The comprehensive personnel risk value ranges from greater than 0.6 to less than or equal to 0.8, with higher safety risks. Safety should be the core focus, with strict constraints on the speed limit and speed variation range. Level 5 (Level V, High Risk): The comprehensive personnel risk value ranges from greater than 0.8 to less than or equal to 1, with extremely high safety risks. The strictest safety constraints should be implemented, and only basic transport functions should be retained.

[0068] Subsequently, dynamic speed constraint interval calibration is performed. Based on personnel risk level, equipment type, and inherent rated parameters of the equipment, the hard constraint interval for speed decision is calibrated. The speed decision result must fall within this interval; otherwise, it is considered an invalid decision. The interval calibration rules are as follows according to equipment type:

[0069] The first category is low-speed open conveyor equipment, with a rated speed limit of 1.0 m / s. This includes moving walkways and barrier-free conveyor equipment. The interval marking rules are as follows: Level I low risk corresponds to a speed constraint interval of 0.5 m / s to 1.0 m / s; Level II relatively low risk corresponds to a speed constraint interval of 0.4 m / s to 0.9 m / s; Level III medium risk corresponds to a speed constraint interval of 0.3 m / s to 0.7 m / s; Level IV relatively high risk corresponds to a speed constraint interval of 0.2 m / s to 0.5 m / s; and Level V high risk corresponds to a speed constraint interval of 0.1 m / s to 0.3 m / s.

[0070] The second category is medium-speed open conveyor equipment, with a rated speed limit of 1.5 meters per second, including escalators. The interval marking rules are as follows: Level I low risk corresponds to a speed constraint interval of 0.6 meters per second to 1.2 meters per second; Level II relatively low risk corresponds to a speed constraint interval of 0.5 meters per second to 1.0 meters per second; Level III medium risk corresponds to a speed constraint interval of 0.4 meters per second to 0.8 meters per second; Level IV relatively high risk corresponds to a speed constraint interval of 0.3 meters per second to 0.6 meters per second; and Level V high risk corresponds to a speed constraint interval of 0.2 meters per second to 0.4 meters per second.

[0071] The third category is high-speed equipment with enclosed cabins, including passenger ropeways and cable cars. The speed limit is set as a percentage of the equipment's rated speed, with the following rules: Level I (low risk) corresponds to a speed limit range of 80% to 100% of the rated speed; Level II (lower risk) corresponds to a speed limit range of 75% to 95% of the rated speed; Level III (medium risk) corresponds to a speed limit range of 70% to 85% of the rated speed; Level IV (higher risk) corresponds to a speed limit range of 60% to 75% of the rated speed; and Level V (high risk) corresponds to a speed limit range of 50% to 65% of the rated speed.

[0072] The fourth category is large-scale amusement rides for human passengers. The speed limit is set as a percentage of the equipment's rated operating speed. The setting rules are as follows: Level I (low risk) corresponds to a speed limit range of 80% to 100% of the rated speed; Level II (lower risk) corresponds to a speed limit range of 70% to 90% of the rated speed; Level III (medium risk) corresponds to a speed limit range of 60% to 80% of the rated speed; Level IV (higher risk) corresponds to a speed limit range of 50% to 70% of the rated speed; and Level V (high risk) corresponds to a speed limit range of 40% to 50% of the rated speed.

[0073] For outdoor equipment, if the environmental conditions data collected by S1 trigger severe weather conditions, including wind force greater than or equal to level 5, heavy rain, and visibility less than 100 meters, the speed limit will be further reduced by no less than 20% based on the speed range corresponding to the above-mentioned personnel risk level, to ensure operational safety in extreme environments.

[0074] Next, speed change rate constraint threshold calibration is performed. Based on personnel risk level and equipment type, acceleration and deceleration rate constraint thresholds are calibrated during speed adjustment to prevent passenger falls and equipment impacts caused by sudden speed changes. The constraint rules are as follows according to equipment type: Category 1 is low-speed open conveyor equipment, with the following constraints: For Level I low risk and Level II relatively low risk, the acceleration constraint threshold is less than or equal to 0.1 m / s², and the single speed adjustment amplitude is less than or equal to 0.2 m / s; for Level III medium risk, Level IV relatively high risk, and Level V high risk, the acceleration constraint threshold is less than or equal to 0.05 m / s², and the single speed adjustment amplitude is less than or equal to 0.1 m / s. Category 2 is medium-speed open conveyor equipment, with the following constraints: For Level I low risk and Level II relatively low risk, the acceleration constraint threshold is less than or equal to 0.15 m / s², and the single speed adjustment amplitude is less than or equal to 0.3 m / s; for Level III medium risk, Level IV relatively high risk, and Level V high risk, the acceleration constraint threshold is less than or equal to 0.08 m / s², and the single speed adjustment amplitude is less than or equal to 0.15 m / s. The third category is high-speed enclosed cabin equipment, with the following constraints: For Level I (low risk) and Level II (lower risk), the acceleration constraint threshold is less than or equal to 50% of the equipment's rated acceleration, and the single speed adjustment amplitude is less than or equal to 5% of the rated speed; for Level III (medium risk), Level IV (higher risk), and Level V (higher risk), the acceleration constraint threshold is less than or equal to 30% of the equipment's rated acceleration, and the single speed adjustment amplitude is less than or equal to 3% of the rated speed. The fourth category is large-scale amusement rides for human passengers, with the following constraints: For Level I (low risk) and Level II (lower risk), the acceleration constraint threshold is less than or equal to 40% of the equipment's rated acceleration, and the single speed adjustment amplitude is less than or equal to 4% of the rated speed; for Level III (medium risk), Level IV (higher risk), and Level V (higher risk), the acceleration constraint threshold is less than or equal to 20% of the equipment's rated acceleration, and the single speed adjustment amplitude is less than or equal to 2% of the rated speed.

[0075] Finally, multi-objective weight prior parameter calibration is performed. Based on the personnel risk level, the weight prior parameters of the four sub-items in the subsequent reward function—safety reward, energy-saving reward, passenger experience reward, and equipment protection reward—are calibrated. The sum of the weights is always equal to 1. The calibration rules are as follows: For Level I low risk, the weights are: safety reward 0.3, energy-saving reward 0.3, passenger experience reward 0.2, and equipment protection reward 0.2; for Level II lower risk, the weights are: safety reward 0.35, energy-saving reward 0.25, and passenger experience reward 0.2. The weighting for equipment protection rewards is 0.2; for Level III medium risk, the weightings are: safety reward weight 0.45, energy saving reward weight 0.15, passenger experience reward weight 0.25, and equipment protection reward weight 0.15; for Level IV higher risk, the weightings are: safety reward weight 0.6, energy saving reward weight 0.1, passenger experience reward weight 0.2, and equipment protection reward weight 0.1; for Level V high risk, the weightings are: safety reward weight 0.75, energy saving reward weight 0.05, passenger experience reward weight 0.15, and equipment protection reward weight 0.05.

[0076] If the wear coefficient of the core load-bearing component of the equipment in the low-dimensional state vector is greater than or equal to 0.8, or the real-time load rate of the power system is greater than or equal to 80%, then on the basis of the above weights, the weight of the equipment protection reward will be increased to no less than 0.4, and the weights of the energy-saving reward and passenger experience reward will be decreased simultaneously to ensure that the weight of the safety reward is not reduced; if the queue length at the equipment entrance is greater than or equal to 10 people, or the cabin load is close to the rated upper limit, then the weight of the passenger experience reward will be appropriately increased to prioritize the efficiency of transportation and evacuation.

[0077] After this step is completed, the output dynamic speed constraint range and speed change rate constraint threshold are directly used as hard constraint inputs for the subsequent S4 speed decision and S7 speed execution stages; the output multi-objective weight prior parameters are directly used as the core inputs for the subsequent S5 reward function calculation stage.

[0078] S4. Optimal velocity decision generation by fusing maximum entropy deep reinforcement learning with manifold regularization constraints:

[0079] The core objective of this step is to embed the hard constraint boundary generated in the preceding step and the manifold features output by S2 into the decision-making process of maximum entropy deep reinforcement learning. This allows for the output of the optimal target running speed within the constrained feasible region, while simultaneously achieving bidirectional interaction with the manifold regularized nonnegative matrix factorization algorithm. This addresses the problems of conventional deep reinforcement learning, such as decision-making prone to going out of bounds, imbalance between exploration and exploitation, and insufficient weighting of key features. Maximum entropy deep reinforcement learning is a niche algorithm in this field. Its core advantage lies in achieving the optimal balance between exploration and exploitation by maximizing policy entropy, avoiding the policy from getting trapped in local optima. It also better adapts to the dynamic changes in personnel-equipment-environment coupling conditions, exhibiting stronger decision robustness and generalization ability under complex conditions compared to conventional deep Q-networks.

[0080] The inputs to this step include the low-dimensional state vector, optimal basis matrix, and personnel risk quantification index set output by S2, and the dynamic velocity constraint interval and velocity change rate constraint threshold output by S3. The processing consists of four core parts: maximum entropy deep reinforcement learning model construction, bidirectional interaction mechanism with manifold regularized nonnegative matrix factorization, model training, and model application.

[0081] refer to Figure 3 First, we construct the model for the maximum entropy deep reinforcement learning algorithm. The core of maximum entropy deep reinforcement learning is to maximize the policy entropy while maximizing the expected reward, thereby maintaining the policy's exploratory ability and avoiding premature convergence to a suboptimal policy while ensuring the decision-making benefit. This scheme constructs a constrained maximum entropy policy gradient framework, using a stochastic policy network to output the probability distribution of actions instead of deterministic actions, to adapt to dynamically changing constraint boundaries and working conditions.

[0082] Define the policy network as ,in These are the trainable parameters of the policy network. The input is a low-dimensional state vector. The output of the policy network is the probability distribution of all candidate speed actions within the dynamic speed constraint interval, which is the speed action output.

[0083] To achieve constrained maximum entropy optimization, the objective function for policy optimization is constructed as follows:

[0084] ;

[0085] It represents the decision trajectory consisting of state, action, and reward; Indicates the maximum time length of the trajectory; Indicates the trajectory Obedience strategy Expectation calculation when the distribution is distributed; This represents the total reward value at time t, which can be calculated using the subsequent formula; Entropy is used to measure the randomness of a strategy; the higher the entropy value, the stronger the strategy's exploratory ability. This represents the entropy regularization weight coefficient, used to balance expected reward and policy entropy, with a fixed value of 0.1. This indicates a constraint violation penalty, used to penalize actions that exceed the dynamic speed constraint range and the speed change rate constraint threshold. This represents the constraint penalty weight coefficient, which is fixed at 10 to ensure the strictness of safety constraints and prevent decisions from going out of bounds.

[0086] Policy Entropy The calculation formula is as follows: ; This represents the set of all candidate speed actions within the dynamic speed constraint range. The discretization step size of the candidate speed is determined according to the equipment type. For open conveyor equipment, the discretization is performed within the dynamic speed constraint range in steps of 0.01 meters per second. For high-speed equipment with enclosed cabins, the discretization is performed within the dynamic speed constraint range in steps of 0.5% of the rated speed. This ensures that all candidate speeds fall within the dynamic speed constraint range, thereby reducing the occurrence of constraint violations from the root. Represents a set One of the candidate speed actions; This represents the probability of the candidate speed action output by the policy network.

[0087] Constraints and penalties The calculation formula is as follows: If ,but ;like ,but ;like ,but In other cases... ; This represents the speed of action at time t; This represents the actual speed at which the device operates at time t-1. and These represent the upper and lower limits of the dynamic speed constraint range output by S3, respectively. This represents the threshold value for the single speed adjustment range output by S3.

[0088] The policy network's structural design and constraint embedding are deeply integrated. The network structure consists of four layers: input layer, attention layer, dual-branch hidden layer, and output layer, as follows: The input layer has a fixed dimension of 16, perfectly matching the low-dimensional state vector dimension of the S2 output. It also incorporates an auxiliary feature vector composed of a set of personnel risk quantification indicators to ensure that the input features cover all core decision-making criteria. The attention layer weights the input features, applying weights to enhance safety-related core features such as the proportion of high-risk passengers and the wear coefficient of core equipment components, thereby increasing the impact of key features on decision-making. The dual-branch hidden layer has two parallel fully connected branches: a safety decision branch and an optimization decision branch. The safety decision branch focuses on safety constraint-related features to ensure compliance, while the optimization decision branch focuses on multi-objective balance to achieve synergistic optimization of energy saving, user experience, and equipment protection. Both branches have two fully connected layers: the first layer has 32 neurons, and the second layer has 16 neurons, both using the ReLU activation function. The output layer uses the Softmax activation function to output the probability distribution of all candidate speed actions within the dynamic speed constraint interval, ensuring that the sum of the probabilities of all actions is 1.

[0089] The core innovation of this step is then constructed, namely the bidirectional interaction mechanism between manifold regularized nonnegative matrix factorization and maximum entropy deep reinforcement learning. The two algorithms achieve deep bidirectional interaction, breaking the information silo problem of feature extraction and decision-making algorithms running independently in conventional technologies, and realizing the co-evolution of feature representation and decision optimization. The specific interaction mechanism is divided into two directions: positive feature transfer interaction and reverse manifold constraint interaction.

[0090] The first approach is positive feature propagation interaction. The low-dimensional state vector output by S2 is directly input into the policy network of the maximum entropy deep reinforcement learning system as the core input of the policy network. At the same time, the personnel risk quantification index set output by S2 is used as an auxiliary feature input into the attention module of the policy network to strengthen the influence of key risk features on decision-making. The interaction formula is as follows: ; This represents the final input feature vector of the policy network; This represents the low-dimensional state vector output by S2; Indicates feature concatenation operation; This represents an auxiliary feature vector composed of a set of quantitative indicators of personnel risk. Through this interaction, the policy network can directly obtain low-dimensional features optimized by manifold regularization, while focusing on core information related to personnel risk, significantly improving the safety and accuracy of decision-making.

[0091] The second approach involves inverse manifold constraint interaction. The action probability distribution output by maximum entropy deep reinforcement learning is fed back into the manifold regularized nonnegative matrix factorization algorithm to optimize the basis matrix and coefficient matrix. This makes the low-dimensional state vector better reflect the features related to the optimal velocity decision, allowing the feature representation to better serve the decision objective. Specifically, this is achieved by adjusting the objective function of the manifold regularized nonnegative matrix factorization algorithm and introducing decision-related constraint terms. The adjusted objective function formula is as follows:

[0092] ;

[0093] Let represent the adjusted manifold regularized nonnegative matrix factorization objective function; The original objective function defined in S2; This represents the reverse constraint weight coefficient, with a fixed value of 0.05; Represents the low-dimensional state vector Follows the optimal coefficient matrix Expectation calculation when the distribution is distributed; This represents the optimal low-dimensional state vector corresponding to the optimal speed action, which is derived by back-deriving the feature gradient corresponding to the optimal action output by the policy network.

[0094] Based on the adjusted objective function, the same multiplicative iterative update rule as in S2 is used to update the optimal basis matrix. and the optimal coefficient matrix Fine-tuning is performed, with 50 iterations, to ensure a stronger correlation between the low-dimensional state vector and the optimal decision, thereby achieving bidirectional co-evolution of feature representation and decision optimization.

[0095] After completing the model construction and interaction mechanism design, the model training of the maximum entropy deep reinforcement learning algorithm is performed. The model training is divided into two stages: offline pre-training and online fine-tuning, to ensure that the model first has basic decision-making ability and then continuously adapts to the dynamic changes of real-time working conditions.

[0096] The first stage is the offline pre-training stage. The training input uses historical operating data of the corresponding type of target conveying equipment to construct a pre-training dataset. The dataset size is no less than 1 million valid samples. Each sample is in the format of a historical low-dimensional state vector, historical velocity action, historical reward value, and the next time step historical low-dimensional state vector. The training objective is to minimize the negative value of the constrained maximum entropy objective function, that is, to maximize... This allows the policy network to learn decision policies that meet safety constraints and multi-objective optimization requirements. The training algorithm uses a proximal policy optimization algorithm, which is suitable for handling constrained policy optimization problems and has significantly better training stability than conventional policy gradient algorithms, effectively avoiding training crashes caused by excessively large policy update steps. The training process is as follows: First, initialize the parameters of the policy network. The system is initialized using a random normal distribution; then, batch samples are sampled from the pre-training dataset with a batch size of 64; the policy gradient is calculated based on the batch samples; the policy network parameters are updated using the pruning objective function of the near-end policy optimization algorithm; the sampling, gradient calculation, and parameter update steps are repeated until the expected reward value on the validation set no longer increases after 100 consecutive iterations, or the number of iterations reaches 5000, at which point pre-training stops, and the optimal parameters of the pre-trained policy network are obtained. .

[0097] The second stage is the online fine-tuning stage. The training input is real-time collected experience samples. A new set of experience samples is constructed and stored in the experience replay pool every second. Every second, 32 sets of samples are randomly sampled from the experience replay pool for parameter updates. The training objective is to enable the policy network to continuously adapt to the dynamic changes of real-time conditions, maintain decision accuracy and scenario adaptability, and maintain sufficient policy entropy to achieve a balance between exploration and exploitation. The training algorithm uses the same proximal policy optimization algorithm as the offline pre-training stage. The training process is as follows: Real-time experience samples are stored in the experience replay pool, and the most... The maximum capacity is set to 1 million samples. When the maximum capacity is exceeded, the oldest samples are replaced according to the first-in-first-out principle. A batch of samples is randomly sampled from the experience replay pool. The policy gradient is calculated based on the batch of samples. At the same time, a manifold regularization constraint term is introduced into the gradient calculation to ensure that the manifold structure of the low-dimensional state vector is not destroyed. The policy network parameters are updated by pruning the objective function. Meanwhile, after every 10 parameter updates, a reverse interaction with the manifold regularization nonnegative matrix factorization is triggered to fine-tune the basis matrix and coefficient matrix. The above process is continuously executed to realize the online real-time fine-tuning of the policy network.

[0098] After model training is completed, the maximum entropy deep reinforcement learning algorithm is applied to the model. The core is to output the optimal target running speed based on the real-time low-dimensional state vector. The specific execution process is as follows: First, the low-dimensional state vector output by S2 in real-time and the auxiliary feature vector composed of the personnel risk quantification index set are concatenated into the final input feature vector of the policy network. Then, the input feature vector is input into the trained policy network. After forward propagation, the probability distribution of all candidate speed actions within the dynamic speed constraint interval is output. Next, the final initial optimal target running speed is selected based on the output probability distribution. During the model training exploration phase, a random sampling method is used to randomly sample a speed action from the probability distribution to achieve better environmental exploration. During the stable operation phase, the maximum probability selection method is used to select the speed action with the highest probability as the initial optimal target running speed. Finally, a secondary constraint verification is performed on the initial optimal target running speed to check whether the difference between it and the current actual running speed of the equipment is less than or equal to the single speed adjustment amplitude threshold output by S3. If this is satisfied, it is determined to be a valid decision, and the final optimal target running speed is output. If the condition is not met, then among all candidate speeds that satisfy the difference constraint, the speed with the second highest probability is selected as the final optimal target running speed. .

[0099] After this step is completed, the output is the optimal target running speed. It is directly used as the core input for the subsequent S5 reward function calculation and S7 speed execution stage; the output policy network parameters and candidate speed action probability distribution are used for the subsequent model iteration update in the S6 stage; the output base matrix and coefficient matrix after reverse interaction adjustment are synchronously fed back to the S2 stage to update the parameters of the feature processing model.

[0100] S5. Calculation of a multi-objective dynamic weighted reward function based on the coupling characteristics of personnel, equipment, and environment:

[0101] The core objective of this step is to construct a four-element weighted reward function that balances safety, energy efficiency, passenger experience, and equipment protection, using the personnel risks, constraint parameters, and decision results from previous steps as core inputs. This function dynamically adjusts the weights based on real-time operating conditions, providing accurate and condition-aligned feedback signals for the iterative updates of the reinforcement learning model. This addresses the problems of fixed weights, imbalances in multi-objective collaboration, and disconnection from previous decision-making processes inherent in conventional reward functions. The reward function is the core feedback signal in reinforcement learning, and its accuracy directly determines the learning direction of the policy network and the final decision-making effect.

[0102] The inputs to this step include the spatiotemporally aligned and normalized multi-source dataset output by S1, the personnel risk quantification index set output by S2, the multi-objective weight prior parameters output by S3, the optimal target running speed output by S4, and the real-time running monitoring dataset fed back by S7. The real-time running monitoring dataset is only used for non-first-time calculation scenarios and is not involved in the first calculation.

[0103] This step involves four core processes: determining the core framework of the reward function, adjusting dynamic weights in real time, calculating individual rewards, calculating the total reward value, and correcting anomalies.

[0104] refer to Figure 3 First, the core framework of the reward function is determined. The reward function adopts a quaternary weighted summation structure, and the total reward value is the sum of the products of the four sub-rewards and their corresponding dynamic weights. The core formula is as follows: ; This is the total reward value, ranging from 0 to 1; As a sub-category of safety rewards, the core measure is the safety compliance of personnel and equipment operations; As a sub-item of energy-saving rewards, the core measure is the energy-saving effect of equipment operation; The passenger experience reward categories primarily measure passenger comfort and transport efficiency. The equipment protection reward category primarily measures the wear and load of the equipment's core components. , , , These are the dynamic weights corresponding to the four sub-rewards, satisfying... All weights range from 0 to 1.

[0105] Subsequently, dynamic weight adjustments are performed in real time. Based on the multi-objective weight prior parameters output by S3, and combined with the real-time collected personnel-equipment-environment coupled operating conditions, the weights are dynamically fine-tuned to ensure that the weights are fully matched with the core control objectives of the current operating conditions. The fine-tuning rules are as follows: If sudden changes in passenger flow, abnormal equipment trends, or severe outdoor weather conditions are detected, the safety reward weight and equipment protection reward weight are increased first, with the total weight of the two not less than 0.8, while the energy reward and passenger experience reward weights are decreased simultaneously; if the wear coefficient of the core load-bearing components of the equipment increases by more than 0.01 per month, or the power system load rate increases continuously... If the percentage of passengers with high risk exceeds 80% for more than 5 consecutive seconds, the weight of the equipment protection reward will be increased by at least 0.2, while the weight of the energy reward will be decreased simultaneously. If the percentage of high-risk passengers exceeds 40%, the weight of the safety reward and passenger experience reward will be increased, with the total weight of the two weights not less than 0.7, while the weight of the energy reward and equipment protection reward will be decreased simultaneously. If the operating conditions are stable for a long period of time, with no risk triggers or parameter changes for 30 consecutive seconds, and the personnel risk level is Level I low risk or Level II low risk, the weight of the energy reward can be moderately increased, not exceeding 0.4, while other weights are slightly adjusted simultaneously to ensure that the safety reward weight is not less than 0.25.

[0106] After the weights are adjusted, they are normalized to ensure that the sum of the four weights is always equal to 1. The weight adjustment results are recorded to provide a basis for subsequent sub-item reward calculations.

[0107] Next, the calculation of four sub-rewards will be performed. All sub-reward calculations will use the parameters output from the previous steps and the real-time collected data, without introducing any parameters out of thin air. The complete calculation rules are explained below:

[0108] The first sub-item is the safety reward sub-item. The core metrics include personnel safety status, safety protection compliance, and the triggering of safety incidents. The initial and subsequent calculations use the same calculation logic, as shown in the following formula:

[0109] ;

[0110] The average safe distance deviation value for passengers is calculated by taking the distance data between adjacent passengers collected by binocular vision for open equipment and calculating the mean of the absolute values ​​of the deviations of each distance from the standard safe distance. The standard safe distance is determined according to the type of equipment and national standards. For open conveyor equipment, it is 0.5 meters, and for enclosed cabin equipment, it is the per capita safe space of the cabin's rated passenger capacity. The maximum permissible safety clearance deviation is fixed at 0.3 meters, ensuring that the value of this item is within the range of 0 to 1; The cumulative number of monitored safety anomalies, including passenger crowding, falls, wrong-way movement, climbing, failure to wear safety protection devices in compliance with regulations, abnormal opening of cabin doors, and equipment overload, will deduct 0.2 reward points for each anomaly. If a serious safety incident occurs that triggers the safety protection device, the safety reward points will be directly reduced to 0.

[0111] The second sub-item is the energy-saving incentive sub-item. The core measure is the energy-saving effect of equipment operation. The initial calculation is based on theoretical energy consumption, while subsequent calculations are based on actual energy consumption, according to the following rules:

[0112] In the initial calculation, based on the optimal target operating speed and the rated power of the equipment's power system, the theoretical energy consumption at the current speed is calculated and compared with the theoretical energy consumption at the rated speed, as shown in the following formula:

[0113] ;

[0114] The theoretical energy consumption at the equipment's rated speed is obtained by multiplying the equipment's rated power by the calculation period; The theoretical energy consumption at the optimal target operating speed is calculated based on the relationship between the equipment's operating resistance characteristics and speed. The max function is used to ensure that the calculated result is directly taken as 0 when it is less than 0, thus guaranteeing that the reward value is non-negative.

[0115] For subsequent calculations, the actual energy consumption data of the device, based on feedback from S7, is calculated using the following formula:

[0116] ;

[0117] In the formula, The actual energy consumption of the equipment during the calculation cycle is calculated by real-time power integration collected by the power sensor built into the power system; if the actual energy consumption is higher than the theoretical energy consumption at the rated speed, the bonus value is directly recorded as 0 points.

[0118] The third sub-item is the passenger experience reward sub-item. The core metrics are passenger comfort and transport efficiency, dynamically adjusted based on passenger risk characteristics. The initial and subsequent calculations use the same logic, as shown in the following formula:

[0119] ;

[0120] For the first calculation, the deviation value of passenger running time is the ideal running time calculated based on the optimal target running speed and the effective transport stroke of the equipment, and the deviation value of the standard running time at the rated speed of the equipment; for subsequent calculations, it is the absolute value of the deviation between the actual average passenger dwell time fed back by S7 and the ideal running time. The maximum allowable running time deviation is dynamically determined based on the equipment stroke length, with a minimum value of 5 seconds. The passenger discomfort coefficient is calculated based on a set of quantitative indicators of personnel risk, using the following formula: ,in For the proportion of passengers carrying large luggage, For the proportion of elderly passengers, The percentage of child passengers; if the calculated result is less than 0, the passenger experience reward score will be 0 points.

[0121] The fourth sub-item is the equipment protection reward sub-item. The core measurement measures the wear and load on key components during equipment operation to prevent overload and exceeding limits. The initial and subsequent calculations use the same logic, as shown in the following formula:

[0122] ;

[0123] The real-time load rate of the equipment power system is derived from the equipment operating status data collected by S1 or the real-time monitoring data fed back by S7. The wear coefficient of the core load-bearing components of the equipment is derived from the equipment operation status data collected by S1 or the real-time monitoring data fed back by S7. Each of the two indicators accounts for 0.5 of the weight, which can be fine-tuned according to the equipment type to ensure balanced protection of the core components of the equipment. If the calculation result is less than 0, the equipment protection bonus score will be directly recorded as 0.

[0124] Finally, the total reward value is calculated and anomaly correction is performed. First, according to the core formula of the reward function, the four sub-rewards are multiplied by their adjusted dynamic weights and then summed to obtain the initial total reward value. If anomaly detection results are output in the subsequent S6 stage, the total reward value needs to be corrected. The correction formula is as follows: ;In the formula, This is the adjusted total reward value after the anomaly correction. To ensure the rationality of reward feedback under abnormal conditions and guide the policy network to learn safer decision-making strategies, the abnormal score output by the Gaussian process regression anomaly detection algorithm in the subsequent S6 stage is used to ensure that the corrected total reward value is truncated to ensure that the final output total reward value is fixed in the range of 0 to 1, which fully meets the input requirements of the reinforcement learning model for reward value.

[0125] After this step is completed, the total reward value output will be directly used as the core feedback input for the subsequent S6 stage reinforcement learning model iteration and update; the output of each sub-item reward value and dynamic weight adjustment results will provide a reference for subsequent abnormal working condition handling and model optimization.

[0126] S6. Iterative optimization of the model integrating Gaussian process regression anomaly detection and maximum entropy strategy updating:

[0127] The core objective of this step is to use the state, action, and reward data output from previous steps as core input, and employ a Gaussian process regression anomaly detection algorithm to achieve anomaly detection across all operating conditions. Simultaneously, the anomaly detection results are deeply integrated with a maximum entropy deep reinforcement learning algorithm to achieve phased iterative updates of the model. This addresses the problems of insufficient uncertainty estimation, low accuracy in small-sample detection, and disconnection from the decision model in conventional anomaly detection algorithms. Furthermore, it forms a three-way interactive closed loop with the preceding manifold regularized nonnegative matrix factorization and maximum entropy deep reinforcement learning algorithms, achieving synergistic optimization throughout the entire process. Gaussian process regression anomaly detection is a niche algorithm in this field. Its core advantage lies in providing uncertainty estimation for anomaly scores, achieving significantly higher detection accuracy for small-sample anomalies than conventional algorithms such as isolated forests, and naturally integrating with probabilistic maximum entropy deep reinforcement learning models to achieve deep linkage between anomaly detection and decision optimization.

[0128] The inputs to this step include the low-dimensional state vector and optimal basis matrix output by S2, the optimal target running speed and candidate speed action probability distribution output by S4, the total reward value output by S5, and the real-time running monitoring dataset fed back by S7. The processing consists of five core steps: building a Gaussian process regression anomaly detection model, building a three-way interaction mechanism, training the Gaussian process regression model, applying the model, and iteratively updating the maximum entropy deep reinforcement learning model in stages.

[0129] refer to Figure 3 First, we construct the Gaussian process regression anomaly detection algorithm model. Gaussian process regression is a nonparametric Bayesian model. Its core idea is to fit the distribution of data through Gaussian process priors. For samples under normal operating conditions, the model can provide accurate predictions with relatively small uncertainties; for abnormal samples, the model's prediction uncertainty will increase significantly, thereby achieving anomaly detection.

[0130] Define the training dataset under normal operating conditions as follows: ,in Let be the low-dimensional state vector of the i-th normal sample. This corresponds to the normal label value, which is uniformly set to 0. This represents the number of normal samples.

[0131] The Gaussian process regression model is entirely defined by the mean function and the covariance function. In this scheme, the mean function is a zero-mean function, that is, it is assumed that the prior mean of the data is 0; the covariance function is the squared exponential covariance function, which has the characteristics of good smoothness and strong fitting ability, and is very suitable for processing continuous working condition data. The formula is as follows:

[0132] ;

[0133] and Let represent the low-dimensional state vectors of the i-th and j-th samples, respectively; This represents the signal variance, used to control the overall fluctuation range of the function, and its initial value is 1. This represents the length scale, used to control the rate at which the correlation of the covariance function decays, i.e., how far apart two samples are before the correlation significantly decreases; the initial value is 1. This represents the noise variance, used to simulate noise in the observed data, with an initial value of 0.01. This represents the Kronecker function, which takes the value 1 when i equals j, and 0 otherwise.

[0134] For the sample to be tested Gaussian process regression models can output their predicted mean. and prediction variance The prediction variance represents the model's uncertainty regarding the sample. A larger variance indicates that the sample deviates more from normal operating conditions, and the higher the probability of anomalies. The formulas for calculating the prediction mean and prediction variance are as follows:

[0135] ;

[0136] ;

[0137] Indicates the sample to be tested The covariance vector between the sample and all normal training samples has a dimension of . ; This represents the covariance matrix among normal training samples, with dimension 1. The element in the i-th row and j-th column is ; Represents the identity matrix, with dimensions AND Consistent; This represents the label vector of a normal training sample, with all elements being 0.

[0138] Based on prediction variance Calculate anomaly scores The anomaly score ranges from 0 to 1, with higher values ​​indicating a higher probability of the sample being an anomaly. The formula is as follows:

[0139] ;

[0140] In the formula, This represents the variance of the anomaly threshold, determined based on the predicted variance distribution of normal training samples. It ensures that over 95% of normal samples have anomaly scores below a preset threshold. The preset threshold is 0.65 to 0.75 for low-speed devices, 0.7 to 0.8 for medium-speed devices, and 0.6 to 0.7 for high-speed devices. The anomaly score is compared to the preset threshold. If the anomaly score is greater than the threshold, it is considered an anomaly sample, and an anomaly signal and score are output, with an anomaly label of 1. If the score is less than or equal to the threshold, it is considered a normal sample, and a normal signal and anomaly score are output, with an anomaly label of 0.

[0141] The core innovation of this step is then constructed: a three-way interaction mechanism among Gaussian process regression anomaly detection, manifold regularized nonnegative matrix factorization, and maximum entropy deep reinforcement learning. These three algorithms achieve deep bidirectional interaction, forming a complete closed-loop optimization system. This breaks the information silo problem of independent operation of algorithms in conventional techniques, achieving the co-evolution of feature representation, decision optimization, and anomaly detection. The specific interaction mechanism is divided into three groups, fully explained below:

[0142] The first group is a two-way interaction between Gaussian process regression and manifold regularized nonnegative matrix factorization, which is divided into two directions: positive interaction and negative interaction.

[0143] Forward interaction: The low-dimensional state vector output by S2 is directly input into the Gaussian process regression anomaly detection model as the feature vector of the sample to be detected. At the same time, the optimal basis matrix output by S2 after training is used for covariance function preprocessing of the Gaussian process regression model, further enhancing the manifold structure characteristics of the features and improving the anomaly detection accuracy. The interaction formula is as follows: ; This represents the final input feature vector of the Gaussian process regression model; This represents the optimal basis matrix output by S2; This represents the low-dimensional state vector output by S2. Through linear transformation of the optimal basis matrix, core features related to personnel safety and equipment operation are further highlighted, redundant information is suppressed, and the sensitivity and accuracy of anomaly detection are improved.

[0144] Reverse interaction: The anomaly score output by the Gaussian process regression anomaly detection model is fed back to the manifold regularized nonnegative matrix factorization algorithm to adjust the manifold regularization weight coefficients in the objective function. This allows the feature representation to focus more on key risk features under abnormal conditions, improving the accuracy of feature representation under abnormal conditions. The interaction formula is as follows: ;In the formula, This represents the adjusted manifold regularization weight coefficient; These are the original manifold regularization weights defined in S2; This represents the outlier score output by the Gaussian process regression model. When the outlier score is high, the manifold regularization weights increase accordingly, strengthening the ability to preserve the local manifold structure of the low-dimensional state vector, allowing the feature representation to better capture key changes under abnormal conditions. Based on the adjusted weight coefficients, the same multiplicative iterative update rule as in S2 is used to update the optimal basis matrix. and the optimal coefficient matrix Fine-tuning was performed, with 30 iterations.

[0145] The second group is the bidirectional interaction between Gaussian process regression and maximum entropy deep reinforcement learning, which is also divided into two directions: forward interaction and reverse interaction.

[0146] Forward interaction: The anomaly score output by the Gaussian process regression anomaly detection model is directly input into the objective function of the maximum entropy deep reinforcement learning model. This is used to adjust the weight coefficients of the constraint violation penalty term and the entropy regularization weight coefficients, allowing the policy network to prioritize the safety of decisions under abnormal conditions, reducing the randomness of the policy, and avoiding safety risks caused by exploration under abnormal conditions. The interaction formula is as follows: ; ; This represents the adjusted constraint penalty weight coefficient; These are the original constraint penalty weight coefficients defined in S4; This represents the adjusted entropy regularization weight coefficient; These are the original entropy regularization weight coefficients defined in S4; This represents the outlier score output by the Gaussian process regression model. When the outlier score is high, the constraint penalty weights increase significantly to ensure that decisions more strictly adhere to safety constraints; the entropy regularization weights decrease simultaneously to reduce the randomness of the policy, resulting in more conservative and safer decision actions and avoiding risks under abnormal conditions. Based on the adjusted weight coefficients, the same proximal policy optimization algorithm as in S4 is used to fine-tune the policy network parameters in an emergency.

[0147] Reverse interaction: The optimal target running speed output by the maximum entropy deep reinforcement learning is fed back into the Gaussian process regression anomaly detection model to dynamically adjust the anomaly threshold variance. This enables adaptive adjustment of anomaly detection sensitivity at different running speeds, as higher device running speeds pose greater safety risks and require higher anomaly detection sensitivity to avoid missed detections. The interaction formula is as follows:

[0148] ;

[0149] This represents the adjusted variance of the outlier threshold; This represents the variance of the original outlier threshold. The optimal target running speed output by S4; The rated operating speed of the equipment is used. When the optimal target operating speed is higher, the variance of the anomaly threshold will decrease, and the corresponding anomaly score threshold will decrease, thereby improving the detection sensitivity of anomalies under high-speed operating conditions and avoiding missed detection of safety risks.

[0150] The third group is the bidirectional interaction between manifold regularized nonnegative matrix factorization and maximum entropy deep reinforcement learning. This interaction mechanism has been explained in detail in S4. The three form a complete three-way interactive closed loop. The output of any one link will optimize the model of the other two links, realizing the co-evolution of the whole process and greatly improving the decision accuracy, anomaly detection capability and working condition adaptation capability of the entire system.

[0151] After completing the model construction and interaction mechanism design, the Gaussian process regression anomaly detection algorithm is trained. The model training is divided into three stages: initial training, incremental training, and emergency optimization, to ensure that the model can accurately identify normal working conditions and quickly adapt to the feature changes of abnormal working conditions.

[0152] The first stage is the initial training stage. The training input uses historical data from the normal operating conditions of the corresponding type of target conveying equipment to construct the initial training set. The training set samples are low-dimensional state vectors output by S2, with a sample size of no less than 100,000, ensuring coverage of all normal operating scenarios. The training objective is to learn the distribution of data under normal operating conditions and optimize the hyperparameters of the Gaussian process regression model, including signal variance, length scale, and noise variance. The training algorithm uses the maximum likelihood estimation method to optimize the hyperparameters by maximizing the logarithmic marginal likelihood function of the training data to find the optimal combination of hyperparameters. The training process is as follows: First, initialize the hyperparameters using default initial values; then calculate the logarithmic marginal likelihood function based on the initial training set; maximize the logarithmic marginal likelihood function using the gradient ascent method to optimize the hyperparameters; repeat the calculation of the logarithmic marginal likelihood function and the gradient ascent update steps until the change in the logarithmic marginal likelihood function value is less than a preset threshold. The initial parameters of the trained Gaussian process regression model are obtained when the number of iterations reaches the preset upper limit of 100.

[0153] The second stage is the incremental training stage. The training input is an incremental training set constructed hourly using normal operating data from the past hour, with a sample size of no less than 1000. The training objective is to enable the model to continuously adapt to slow changes in normal operating conditions, such as parameter drift caused by equipment aging and environmental parameter changes caused by seasonal variations, to avoid model performance degradation. The training algorithm uses the same maximum likelihood estimation method as the initial training stage, but only fine-tunes the hyperparameters, with no more than 20 iterations. The training process is as follows: add the incremental training set to the initial training set to form a new training set; fine-tune the hyperparameters based on the new training set; update the covariance matrix and model parameters; complete the incremental training.

[0154] The third stage is the emergency optimization stage, triggered by the detection of abnormal operating conditions and an abnormal score greater than 0.9, or after the abnormal operating conditions have been handled. The training input consists of abnormal samples from the current abnormal operating conditions and samples of similar abnormalities from the past. The training objective is to improve the model's detection accuracy for similar abnormalities and avoid subsequent missed detections. The training process is as follows: add abnormal samples to the training set and set the label value of the abnormal samples to 1; retrain the Gaussian process regression model and adjust the hyperparameters to adapt to the distribution of abnormal samples; update the abnormal threshold variance to ensure that similar abnormalities can be accurately detected; and complete the emergency optimization.

[0155] After model training is completed, the Gaussian process regression anomaly detection algorithm is applied to the model. The core is to output anomaly scores and anomaly signals based on real-time low-dimensional state vectors. The specific execution process is as follows: First, the low-dimensional state vector output by S2 in real time is linearly transformed through the optimal basis matrix to obtain the final input feature vector of the Gaussian process regression model. Then, the input feature vector is input into the trained Gaussian process regression model to calculate the predicted mean and predicted variance. Next, the anomaly score is calculated based on the predicted variance. Finally, the anomaly score is compared with a preset threshold, and the corresponding anomaly signal, anomaly score, and anomaly label are output.

[0156] Finally, this step performs a phased iterative update of the maximum entropy deep reinforcement learning model. Combined with the results of Gaussian process regression anomaly detection, the policy network is updated and optimized in a targeted manner. Specifically, it is divided into two stages: online fine-tuning under normal operating conditions and emergency optimization under abnormal operating conditions.

[0157] The first stage is online fine-tuning under normal operating conditions. The trigger condition is that the abnormal score is less than or equal to a preset threshold, which is judged as a normal operating condition. The update input is real-time collected experience samples. A new set of experience samples is built and stored in the experience replay pool every 1 second. 32 sets of samples are randomly sampled from the experience replay pool every 1 second. The update goal is to make the policy network continuously adapt to the dynamic changes of normal operating conditions, maintain decision accuracy and scenario adaptability, and maintain sufficient policy entropy. The update algorithm adopts the same near-end policy optimization algorithm as in S4, using the conventional weight coefficients before adjustment. The update process is as follows: store real-time experience samples in the experience replay pool; randomly sample batch samples from the experience replay pool; calculate the policy gradient based on the batch samples; update the policy network parameters using the pruning objective function; after every 10 parameter updates, trigger a reverse interaction with manifold regularized nonnegative matrix factorization; continue to execute the above process.

[0158] The second stage is emergency optimization under abnormal operating conditions. The trigger condition is that the abnormal score is greater than a preset threshold, which is judged as an abnormal operating condition. The update input is the emergency sample corresponding to the abnormal operating condition and historical similar abnormal samples, with a batch size of 16. The update goal is to quickly optimize the policy network parameters to adapt to the decision-making needs under abnormal operating conditions and improve the decision-making safety under abnormal operating conditions. The update algorithm adopts the same near-end policy optimization algorithm as in S4, but uses adjusted weight coefficients. The update process is as follows: the learning rate is temporarily increased to twice the normal online learning rate; training is focused on abnormal samples, with no less than 10 iterations; until the decision result corresponding to the abnormal operating condition meets the safety constraints and the reward value returns to the normal range; after optimization, the learning rate is restored to the normal value, and the emergency sample is stored in the abnormal sample exclusive storage area of ​​the experience replay pool; after the abnormal operating condition is resolved, online fine-tuning under normal operating conditions is resumed.

[0159] After this step is completed, the updated maximum entropy deep reinforcement learning policy network parameters are synchronously fed back to the decision network in S4 to achieve real-time iterative optimization of the decision model; the output anomaly detection results are synchronously fed back to the feature processing stage in S2, the constraint space generation stage in S3, the reward function calculation stage in S5, and the speed execution stage in S7; the updated manifold regularized nonnegative matrix decomposition basis matrix and coefficient matrix are synchronously fed back to the feature processing stage in S2; and the updated Gaussian process regression anomaly detection model parameters are used for subsequent real-time anomaly detection.

[0160] S7. Optimal speed smooth execution and closed-loop monitoring of the entire operational status:

[0161] The core objective of this step is to achieve smooth and precise speed execution based on the optimal speed decision results and constraint thresholds output in the preceding steps. At the same time, it enables closed-loop monitoring of equipment operation, personnel status, and environmental conditions throughout the entire process, providing real-time feedback data for all preceding steps. This establishes a closed loop throughout the entire process of perception, decision-making, execution, and feedback, solving the problems of insufficient execution accuracy, untimely monitoring data feedback, and broken closed-loop links in conventional technologies. This step is the final execution stage of the solution, completing the closed-loop control of the entire process.

[0162] The inputs for this step include the optimal target operating speed output by S4, the speed change rate constraint threshold and single speed adjustment amplitude threshold output by S3, and the anomaly detection results output by S6. The processing consists of four core components: optimal speed smooth execution control, closed-loop monitoring of the entire operating status, real-time transmission of monitoring data and closed-loop feedback of the entire system, and real-time linkage handling of abnormal signals.

[0163] refer to Figure 4 First, optimal speed smooth execution control is implemented. The dedicated programmable logic controller (PLC) of the target conveyor receives the optimal target operating speed and generates a smooth speed control curve based on the speed change rate constraint threshold. This ensures that the speed adjustment process is smooth and shock-free, avoiding passenger falls or mechanical shocks to the equipment caused by sudden speed changes. The specific execution process is as follows: The first step is to plan the acceleration and deceleration curve. Based on the difference between the optimal target operating speed and the current actual operating speed of the equipment, combined with the acceleration constraint threshold output by S3, a linear acceleration and deceleration curve is planned. The acceleration during the acceleration and deceleration process remains constant and does not exceed the constraint threshold, ensuring a smooth speed transition. The speed adjustment time is calculated using the following formula: ,in The optimal target running speed is output by S4. This represents the current actual operating speed of the equipment. The first step is to set the acceleration constraint threshold for S3 output; the second step is to generate and issue control commands. The programmable logic controller (PLC) generates real-time speed control commands based on the planned acceleration / deceleration curve and issues them to the frequency converter or speed controller of the equipment's power system every 10 milliseconds to ensure the real-time performance and accuracy of speed control; the third step is to perform speed closed-loop feedback verification. While the frequency converter or speed controller executes the speed control commands, it also collects the actual operating speed of the equipment every 10 milliseconds through the incremental encoder at the drive shaft end of the power system, and verifies the deviation from the target speed. The deviation calculation formula is... ,in This refers to the actual operating speed. The target speed is set at the current moment. If the deviation exceeds the threshold, which is 0.5% of the rated speed of the equipment and no less than 0.05 meters per second, the proportional-integral-derivative fine-tuning is immediately triggered to adjust the output frequency of the frequency converter until the deviation between the actual operating speed and the target speed meets the requirements, thus ensuring the accuracy of speed execution.

[0164] Subsequently, closed-loop monitoring of the entire operational status was implemented, establishing a comprehensive monitoring system covering three dimensions: personnel safety, equipment operation, and environmental conditions. All monitoring data was appended with millisecond-level timestamps to maintain consistency with the time base of the preceding data collection stages, ensuring data time alignment. The monitoring content and data usage are fully explained below:

[0165] The first category is equipment operation status monitoring. The monitoring content includes the actual operating speed of the equipment, the real-time load rate of the power system, the temperature of the motor windings, the status parameters of the braking system, the wear coefficient of the core load-bearing components, the actual energy consumption data of the equipment, the trigger status of the safety protection device, the vibration data of the core components, and the real-time fluctuation value of the grid voltage. The acquisition frequency is set to 10 milliseconds per acquisition for the equipment operating speed and 1 second per acquisition for the other data. For high-speed and heavy-load equipment, the acquisition frequency can be increased to 500 milliseconds per acquisition. The acquired data is fed back to the preprocessing stage of S1 in real time to update the spatiotemporally aligned and standardized multi-source dataset. It is also input into the Gaussian process regression anomaly detection model of S6 for abnormal operating condition identification and provides actual operating data for the reward function calculation of S5.

[0166] The second category is personnel safety status monitoring. The monitoring content includes the real-time number and distribution of passengers, the distance between adjacent passengers, the stability of passengers standing, abnormal passenger behavior, changes in the proportion of high-risk passengers, the wearing status of safety protection devices for passengers in enclosed cabin equipment, and the closure status of cabin doors. The data collection frequency is once every 1 second, consistent with the frequency of the preceding personnel perception data collection. The collected data is fed back to the preprocessing stage of S1 in real time to update the personnel perception feature data, and is synchronously input into the feature processing stage of S2 to update the personnel risk quantification index set. At the same time, it provides real-time data for the safety reward calculation of S5 and provides personnel safety dimension monitoring data for anomaly detection of S6.

[0167] The third category is environmental condition monitoring, which includes environmental temperature and humidity, wind force level of outdoor equipment, rain and snow intensity, environmental visibility, indoor floor slipperiness, and lighting brightness. The data collection frequency is once every 1 second, which can be increased to once every 500 milliseconds under severe weather conditions. The collected data is fed back to the preprocessing stage of S1 in real time to update the environmental condition data, and is synchronously input into the constraint space generation stage of S3 to dynamically adjust the velocity constraint interval and weight parameters. At the same time, it provides environmental dimension monitoring data for anomaly detection in S6.

[0168] Next, real-time transmission of monitoring data and end-to-end closed-loop feedback are implemented. All monitoring data is transmitted in real time via industrial Ethernet. High-security equipment employs redundant dual-link transmission to ensure the reliability and continuity of data transmission. Cyclic redundancy check is used during transmission to ensure data integrity and avoid decision-making errors caused by data transmission mistakes. The feedback link of monitoring data is divided into three links, each corresponding to the preceding steps without any link breaks, achieving a closed-loop iteration throughout the entire process: The first link transmits data to the S1 data cache and preprocessing module, used to update the spatiotemporally aligned and standardized multi-source dataset for the next iteration, providing the latest basic data for feature processing, decision-making, and reward calculation throughout the entire process; the second link transmits data to the S5 reward function calculation module, used for itemized reward calculation in non-first-time scenarios, providing real-time and accurate reward feedback signals for model training; the third link transmits data to the S6 Gaussian process regression anomaly detection module and experience replay pool, used for real-time anomaly detection and experience sample construction, supporting iterative updates and anomaly identification of the model.

[0169] Finally, real-time linkage handling of abnormal signals is executed. Based on real-time monitoring data and combined with the abnormal detection results output by S6, real-time linkage feedback of abnormal signals is completed, synchronizing abnormal information to all preceding stages across the entire chain. This enables real-time strategy adjustment under abnormal operating conditions. The specific rules are as follows: When monitoring data triggers a preset safety red line, including safety protection device triggering, passenger fall, equipment overload, core component parameter exceeding limits, or extreme weather, a safety braking command is immediately generated. The programmable logic controller executes a smooth emergency stop operation, and simultaneously feeds back the abnormal signal to stages S3, S4, and S6 across the entire chain, triggering emergency tightening of constraint boundaries. The system pauses decision-making strategies and locks model parameter updates to ensure personnel and equipment safety. When monitoring data shows an abnormal trend but does not trigger the safety red line, or when S6 outputs an abnormal signal but does not reach the level of a serious anomaly, the abnormal signal is immediately fed back to S3, S4, and S6, triggering dynamic constraint range adjustment, decision-making strategy optimization, and emergency fine-tuning of model parameters. Simultaneously, the monitoring frequency of the corresponding dimension is strengthened to eliminate abnormal risks and prevent the anomaly from escalating. All abnormal signals and related data from the handling process are simultaneously stored in the experience playback pool and model training set for iterative optimization of the preceding algorithm models, achieving adaptive closed-loop control throughout the entire process.

[0170] After this step is completed, the generated actual operation control commands for the equipment are directly sent to the equipment power system to achieve precise and smooth speed control; the collected real-time operation monitoring dataset realizes closed-loop feedback across the entire chain, supporting real-time iterative optimization of each preceding stage; the linkage handling of abnormal signals completes real-time management and control of safety risks, and finally realizes a closed-loop control of the entire process of adaptive speed of the conveyor driven by personnel perception.

Claims

1. A human-perceived speed adaptive control method, characterized in that, include: S1. Collect personnel perception data, equipment operation data and environmental condition data of the target conveying equipment, and perform spatiotemporal alignment processing and standardization preprocessing to obtain a time-synchronized standardized multi-source dataset; S2. The manifold regularized nonnegative matrix factorization algorithm is used to extract features and reduce dimensions of the standardized multi-source dataset to generate low-dimensional state vectors and a set of personnel risk quantification indicators. S3. Based on the personnel risk quantification index set, complete the personnel risk level determination. According to the personnel risk level and the inherent rated parameters of the target conveying equipment, dynamically generate the hard constraint range of speed decision, the speed change rate constraint threshold and the multi-objective weight prior parameters. S4. Construct a maximum entropy deep reinforcement learning decision model that integrates manifold regularization constraints. Input the low-dimensional state vector, hard constraint interval, and velocity change rate constraint threshold into the maximum entropy deep reinforcement learning decision model, and output the initial optimal target running speed within the constraint feasible region. S5. Based on multi-objective weighted prior parameters, combined with personnel perception data, equipment operation data, environmental condition data and initial optimal target operating speed, a quaternary weighted reward function covering safety, energy saving, riding experience and equipment protection is constructed to calculate the real-time total reward value; S6. The Gaussian process regression anomaly detection algorithm is adopted to complete the anomaly detection of the whole working condition based on the low-dimensional state vector and real-time operation monitoring data, and generate anomaly detection results. Combining the real-time total reward value, anomaly detection results and hard constraint interval, the manifold regularized non-negative matrix factorization algorithm, the maximum entropy deep reinforcement learning decision model and the Gaussian process regression anomaly detection algorithm are subjected to bidirectional interactive iterative update and parameter optimization, and the optimized optimal target running speed is output. S7. Based on the optimized target operating speed and speed change rate constraint threshold, generate a smooth speed control curve, send speed control commands to the controller of the target conveying equipment, and complete speed control.

2. The speed adaptive control method based on human perception according to claim 1, characterized in that, The specific process of spatiotemporal alignment is as follows: using the unified clock of the acquisition system as a reference, millisecond-level timestamps are added to personnel perception data, equipment operation data and environmental condition data respectively; the three types of data under the same timestamp are matched and aligned one by one, and invalid data with time deviations exceeding twice the acquisition cycle are removed; missing data are completed by linear interpolation using the average of historical operation data under the same equipment type and the same operating conditions. The standardization preprocessing uses the min-max normalization method to uniformly map all data to the 0-1 interval, eliminating the dimensional differences between data of different dimensions.

3. The speed adaptive control method based on human perception according to claim 1, characterized in that, When using the manifold regularized nonnegative matrix factorization algorithm, the decision-related feature constraints from the reverse output of S6 are received simultaneously. These constraints are generated by the policy gradient derivation after iterative updates of the maximum entropy deep reinforcement learning decision model in S6. The manifold regularization weight coefficients of the manifold regularized nonnegative matrix factorization algorithm are adjusted in conjunction with these constraints. A multiplicative iterative update rule is used to perform collaborative fine-tuning of the basis matrix and coefficient matrix in the manifold regularized nonnegative matrix factorization algorithm, ensuring that the generated low-dimensional state vector preferentially retains the core risk characteristics of personnel and the speed-related decision-making characteristics. Improve the targeting of feature representation.

4. The speed adaptive control method based on human perception according to claim 1, characterized in that, When generating a set of quantitative indicators for personnel risk, four core indicators are extracted based on the passenger age, movement status, safety protection compliance status, and adjacent distance data in the personnel perception data collected by S1: the proportion of high-risk passengers, the deviation rate of the average safe distance between passengers, the compliance rate of safety protection devices, and the proportion of passengers in unstable states. The proportion of high-risk passengers is the sum of the proportions of elderly passengers, child passengers, and passengers with mobility impairments. All four core indicators are mapped to the 0-1 range through min-max normalization, with higher values ​​representing higher risks.

5. The speed adaptive control method based on human perception according to claim 1, characterized in that, In the policy optimization objective function of the maximum entropy deep reinforcement learning decision model, an entropy regularization term and a constraint violation penalty term are introduced simultaneously. The entropy regularization term maintains the model's exploration capability by calculating the entropy value of the policy distribution, thus preventing the policy from getting trapped in local optima. The constraint violation penalty term applies a linear penalty to candidate speed actions that exceed the hard constraint range generated by S3 or the speed change rate constraint threshold. The penalty coefficient is positively correlated with the degree of constraint violation, ensuring that the output initial optimal target running speed strictly meets the constraint requirements.

6. The speed adaptive control method based on human perception according to claim 1, characterized in that, After the maximum entropy deep reinforcement learning decision model outputs the initial optimal target running speed, an additional constraint verification process is executed: the current actual running speed of the target conveying equipment is obtained through the equipment operation data collected by S1, and the difference between the initial optimal target running speed and the current actual running speed is calculated; if the difference exceeds the speed change rate constraint threshold generated by S3, then among all candidate speed actions that meet the speed change rate constraint threshold, the speed with the second best Q value is selected as the adjusted initial optimal target running speed to ensure that there are no stutters or impacts in the speed adjustment process.

7. The speed adaptive control method based on human perception according to claim 1, characterized in that, The dynamic adjustment process of the weights in the quaternary weighted reward function is as follows: If the personnel risk quantification indicators generated by S2 show that the proportion of high-risk passengers exceeds 40%, or if the environmental operating condition data collected by S1 shows that there is severe weather and / or the power grid voltage fluctuation exceeds the preset range, then the total proportion of safety reward weight and equipment protection reward weight will be increased to no less than 70%; if the equipment operation data collected by S1 shows that the wear coefficient of core components is less than 0.2 and the operating condition is stable, then the energy-saving reward weight and the riding experience reward weight will each be increased by 5%-10%, and all weights will be normalized to keep the sum of 1 after adjustment.

8. The speed adaptive control method based on human perception according to claim 1, characterized in that, When the Gaussian process regression anomaly detection algorithm is running, it uses the low-dimensional state vector output by S2 as the core input feature, and simultaneously receives the initial optimal target running speed output by S4. Based on the initial optimal target running speed and the rated speed in the inherent rated parameters of the target conveying equipment, the speed ratio is calculated, and the anomaly detection threshold is dynamically adjusted according to the speed ratio. The higher the speed ratio, the lower the anomaly detection threshold.

9. The speed adaptive control method based on human perception according to claim 1, characterized in that, The specific process of bidirectional interactive iterative update is as follows: First, the real-time total reward value calculated by S5 is used as the core feedback signal for updating the parameters of the maximum entropy deep reinforcement learning decision model, and the anomaly detection result is used as a hard constraint for policy adjustment. The network parameters of the decision model are updated through the proximal policy optimization algorithm. Second, the optimal feature constraint is derived in reverse based on the policy gradient after the maximum entropy deep reinforcement learning decision model is updated. The optimal feature constraint is fed back to S2 to adjust the objective function of the manifold regularized nonnegative matrix factorization algorithm and optimize the feature extraction direction. Third, the anomaly sample data marked by the anomaly detection result is added to the training set of the Gaussian process regression anomaly detection algorithm. The maximum likelihood estimation method is used to complete the incremental optimization of the algorithm hyperparameters and improve the detection accuracy of similar anomaly conditions.

10. The speed adaptive control method based on human perception according to claim 1, characterized in that, For different types of target conveying equipment, it is only necessary to adapt the inherent rated parameters and constraint threshold benchmarks of the corresponding equipment, without reconstructing the core architecture of the manifold regularized nonnegative matrix factorization algorithm, the maximum entropy deep reinforcement learning decision model, and the Gaussian process regression anomaly detection algorithm, so as to realize the adaptable application of the method on various target conveying equipment; the inherent rated parameters include rated speed, rated acceleration, rated load, and core component tolerance threshold.