Resource scheduling method and system based on multi-modal cooperation and deep reinforcement learning

By constructing a digital twin environment at the station, integrating audiovisual data, and using deep reinforcement learning to predict passenger flow diffusion, an optimal resource scheduling strategy is generated. This solves the problems of slow resource scheduling response speed and inconsistent instructions in existing technologies, and achieves efficient and low-energy resource scheduling.

CN121809984BActive Publication Date: 2026-05-12SUZHOU BOYUAN RONGTIAN INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SUZHOU BOYUAN RONGTIAN INFORMATION TECH CO LTD
Filing Date
2026-03-06
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

The existing station resource scheduling system is unable to adapt to the complex and changing passenger flow environment, resulting in slow resource scheduling response speed, low diversion efficiency, energy waste, and inconsistent instructions from various subsystems, making it impossible to form a precise synergy at the same time.

Method used

A digital twin simulation environment for the station is constructed, and a multimodal state matrix is ​​generated by integrating visual passenger flow density and auditory environmental noise. A deep reinforcement learning model is used to predict the passenger flow diffusion trend, and a collaborative strategy of broadcast volume gain and directional screen path pointing is generated. The optimal resource scheduling instruction is verified through physical space mapping.

Benefits of technology

It enables rapid response to complex passenger flow environments, dynamically balances traffic management efficiency and energy consumption, eliminates discrete instructions, and enhances decision-making consistency and on-site guidance synergy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121809984B_ABST
    Figure CN121809984B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of resource scheduling, in particular to a resource scheduling method and system based on multi-modal cooperation and deep reinforcement learning, which comprises the following steps: constructing a digital twin environment and fusing the collected visual passenger flow density and acoustic environment noise in real time to generate a multi-modal state matrix; predicting a passenger flow diffusion trend and fusing the same with the multi-modal state matrix to generate a composite feature vector; inputting the composite feature vector into a deep cooperative scheduling model, using a strategy network with double constraints of maximum evacuation efficiency and minimum energy consumption to generate a primary cooperative strategy containing broadcast volume gain and guiding screen path pointing; performing secondary constraint verification on the primary cooperative strategy to determine an optimal resource scheduling instruction; and issuing the optimal resource scheduling instruction to drive the on-site broadcast and guiding screen to execute actions. The application can realize rapid response, dynamically balance evacuation efficiency and operation energy consumption, effectively eliminate instruction dispersion caused by independent decision-making, and thus improve the decision consistency of resource scheduling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of resource scheduling technology, and in particular to a resource scheduling method and system based on multimodal collaboration and deep reinforcement learning. Background Technology

[0002] Railway stations are crucial nodes connecting urban transportation networks. Their internal spatial structure is vast and their functional areas complex, bearing the burden of handling massive daily passenger volumes. To ensure passengers receive clear route guidance and maintain efficient flow, stations are equipped with comprehensive information display systems and public address systems. These two types of information guidance resources are not only the main channels for passengers to obtain information on entering and exiting the station, waiting directions, and transfer routes, but also the core control tools for station operators to maintain order and optimize passenger flow distribution. The rationality of their scheduling directly affects the overall traffic efficiency and operating costs of the station.

[0003] Existing station resource scheduling mainly adopts either a timetable-based preset rule mode or a passive response mode based on manual monitoring. In the preset rule mode, the content displayed on the wayfinding screens and the frequency of broadcasts strictly correspond to the train arrival and departure timetables. Regardless of whether there is actual passenger congestion in the current area, the system executes the broadcasting task according to predetermined logic. In the manual response mode, staff in the control room observe passenger flow within the station through video surveillance. When they visually detect excessive crowd density or congestion in a certain area, they manually operate the broadcast control terminal to provide voice prompts or switch the display interface of the wayfinding screens to change the directional guidance.

[0004] However, the aforementioned scheduling methods, which rely on human experience or static logic, are ill-suited to the non-linear fluctuations in actual passenger flow distribution at stations. Because fixed rules cannot detect real-time differences in passenger density, they often lead to over-utilization of resources and wasted energy during periods of low passenger flow, while providing insufficient guidance during periods of high passenger flow. Furthermore, manual decision-making is limited by the attention span and reaction speed of staff, resulting in a significant time lag between the detection of congestion and the execution of scheduling instructions. Moreover, since directional displays and broadcasting systems are typically controlled as two separate subsystems, it is difficult to create a precise and consistent guiding force for specific areas simultaneously, ultimately preventing optimal passenger flow management efficiency. In summary, existing technologies, when facing the complex and changing passenger flow environment of stations, suffer from slow resource scheduling response, an inability to balance management efficiency and operational energy consumption, and fragmented guidance instructions and poor decision-making consistency due to the independent operation of each subsystem. Summary of the Invention

[0005] This application provides a resource scheduling method and system based on multimodal collaboration and deep reinforcement learning, which can achieve rapid response to complex changes in stations. While dynamically balancing traffic flow efficiency and operational energy consumption, it effectively eliminates instruction dispersion caused by independent decision-making, thereby improving the consistency of resource scheduling decisions and the synergy of on-site guidance. This application provides the following technical solutions:

[0006] In a first aspect, this application provides a resource scheduling method based on multimodal cooperation and deep reinforcement learning, the method comprising:

[0007] A digital twin simulation environment for the station scene is constructed, and the visual passenger flow density and auditory environmental noise collected on site are integrated in real time to generate a multimodal state matrix that represents the spatial congestion status and information transmission interference level at the current time.

[0008] Based on the inertial motion law of passenger flow, the trend of passenger flow diffusion under non-intervention conditions is predicted, and the trend is fused and serialized with the multimodal state matrix to generate a composite feature vector containing the current state and future inertia.

[0009] The composite feature vector is input into a preset deep collaborative scheduling model. The model uses a strategy network embedded in the model that is based on the dual constraints of maximizing evacuation efficiency and minimizing energy consumption to generate a primary collaborative strategy that includes broadcast volume gain and path pointing to the guide screen.

[0010] The primary coordination strategy is subjected to secondary constraint verification. Based on the physical space mapping relationship, the geometric visibility and physical connectivity between the broadcast sound field and the screen field of view are verified and the parameters are corrected to determine the final optimal resource scheduling instruction.

[0011] The optimal resource scheduling command is issued to drive the on-site broadcasting and guidance screen to perform actions, and the passenger flow evacuation rate and actual energy consumption after execution are monitored in real time. Based on the monitoring results, the feedback deviation is calculated to update the network parameters of the deep collaborative scheduling model online.

[0012] In a specific feasible implementation, the construction of a digital twin simulation environment for the station scene, and the real-time fusion of visual passenger flow density and auditory environmental noise collected on-site, to generate a multimodal state matrix characterizing the spatial congestion state and information transmission interference level at the current time period, includes:

[0013] The physical space of the station is divided into a set of uniformly sized two-dimensional discrete grid cells. ,set up Includes OK List a grid of cells, and denote any one of the grid cells as . ,in Indicates row index, Indicates column index;

[0014] Calculate the grid by capturing real-time video streams using a camera. Average visual passenger flow density Noise sensors are used to collect ambient sound data, and grid calculations are performed. Average ambient noise intensity ;

[0015] Define grid Information transmission interference at the location As shown below:

[0016] ;

[0017] in, The reference noise threshold; This is the noise sensitivity coefficient;

[0018] Visual passenger flow density calculated for each grid cell Normalization process is performed to obtain Normalized visual passenger flow density and information transmission interference Together, they form a dimension. Multimodal state matrix Any position in this matrix The numerical definition is: .

[0019] In a specific feasible implementation, the step of predicting the passenger flow diffusion trend under no-intervention conditions based on the inertial motion law of passenger flow, and fusing and serializing this trend with the multimodal state matrix to generate a composite feature vector containing the current state and future inertia includes:

[0020] For each grid cell Define the grid Passenger flow inertial drift intensity The calculation formula is as follows:

[0021] ;

[0022] in, The normalized visual passenger flow density for the current grid; Indicated by grid The set of neighborhood grids centered on; The number of neighboring grid cells; This is the sum of the normalized densities of the neighborhood grid; The diffusion inertia coefficient;

[0023] Following the spatial scanning order of the grid, extract three feature values ​​at each grid location sequentially. Connect the first and last parts of the vector to form a composite feature vector.

[0024] In one specific implementation scheme, the step of inputting the composite feature vector into a preset deep cooperative scheduling model, and using the policy network embedded in the model based on the dual constraints of maximizing evacuation efficiency and minimizing energy consumption, to generate a primary cooperative strategy including broadcast volume gain and guide screen path direction includes:

[0025] The deep collaborative scheduling model includes an input reconstruction layer, a global feature encoder, and a dual-constraint policy output layer. The dual-constraint policy output layer includes an evacuation guidance branch and an energy consumption control branch.

[0026] The input reconstruction layer uses inverse dimension mapping to transform the composite feature vector. Restored to the station grid OK The three-dimensional spatial feature body corresponding to the column The calculation logic is as follows:

[0027] ;

[0028] in, Indicates the dimensional inverse mapping operator;

[0029] The global feature encoder is composed of stacked residual modules, and the three-dimensional spatial feature volume After inputting the global feature encoder, convolutional feature extraction is performed layer by layer. For the first layer of the network... Each residual module outputs feature data. The calculation is as follows:

[0030] ;

[0031] in, For the output features of the previous layer, Characterizes the mapping of convolution operations within this residual module; and This is the set of weights and biases for each convolutional layer within the module; The ReLU activation function is used; after deep feature extraction, the global feature encoder performs bilinear interpolation upsampling to restore the spatial resolution of the feature map to [value missing]. After full-layer processing, the latent feature map is output. .

[0032] In one specific implementation scheme, the step of inputting the composite feature vector into a preset deep cooperative scheduling model, and using the policy network embedded in the model based on the dual constraints of maximizing evacuation efficiency and minimizing energy consumption to generate a primary cooperative strategy including broadcast volume gain and guide screen path direction, further includes:

[0033] The latent feature map It is simultaneously distributed to two branches, which respectively compute the guiding instructions and broadcast parameters;

[0034] The evacuation guidance branches are used to achieve the constraint of maximizing evacuation efficiency. The density gradient around the mesh is analyzed through a fully connected layer, and the probability of gain for each guide screen's indicated direction is calculated for the mesh. Calculate its path pointing probability vector :

[0035] ;

[0036] in, and The weights and biases of this branch; The function maps the output to a probability distribution of four categories: go straight, turn left, turn right, and stop. The category with the highest probability value is selected as the path pointing to the guide screen for that grid. ;

[0037] The energy consumption control branch is used to achieve the energy consumption minimization constraint. It uses regression analysis to determine whether broadcasting must be enabled in the current region, specifically for the grid. Calculate the broadcast volume gain :

[0038] ;

[0039] in, and For the parameters of this branch; This is the preset silent bias threshold;

[0040] Summarize the broadcast volume gain calculated from all grids Path pointing to the guide screen The two sets of parameters are matched and combined one-to-one according to their spatial location in the grid to form a primary cooperative strategy. .

[0041] In a specific feasible implementation, the secondary constraint verification of the primary cooperative strategy, based on the physical space mapping relationship, verifies and corrects the geometric visibility and physical connectivity between the broadcast sound field and the screen view, and determines the final optimal resource scheduling instruction, including:

[0042] With each grid cell The center point is taken as the observation origin. A line-of-sight vector is formed by connecting this origin with the geometric center of the guide screen display surface. The angle between the line-of-sight vector and the normal to the guide screen surface is calculated. Simultaneously, the straight-line Euclidean distance from the origin to the guide screen is measured. ;

[0043] The actual observed solid angle of the current grid relative to the guide screen is calculated using the solid angle approximation formula. :

[0044] ;

[0045] in, Let be the physical display area of ​​the guide screen. If there are physical obstacles on the line-of-sight vector path, then let Set reference solid angle As a normalization benchmark, Defined as the optimal viewing distance that meets the standards of human visual ergonomics. And the ideal solid angle value corresponding to when facing the screen is calculated as follows: ; Calculate the geometric visibility coefficient :

[0046] .

[0047] In a specific implementation scheme, the secondary constraint verification of the primary coordination strategy, the verification and parameter correction of the geometric visibility and physical connectivity between the broadcast sound field and the screen field of view based on the physical space mapping relationship, and the determination of the final optimal resource scheduling instruction further include:

[0048] Construct the following correction formula:

[0049] ;

[0050] in, The corrected broadcast volume gain; This is the original gain; Geometric visibility coefficient; To normalize visual passenger flow density; Basic penalty factor; This is a risk amplification factor.

[0051] Read path pointing to Based on the station map, a topological adjacency matrix of the entire station grid is constructed. The matrix records the physical access status between each grid and its eight neighboring grids, i.e., connectivity or obstruction. The verification logic is as follows: Check instructions. If the indicated movement direction is marked as blocked in the topological adjacency matrix, then the primary strategy is deemed invalid, and a local neighborhood optimization algorithm is initiated: using the current grid... Centered on a given point, traverse all its topologically connected neighboring grids; calculate the actual path distance from the center point of each neighboring grid to the nearest safe exit; select the direction corresponding to the neighboring grid with the smallest distance and that is passable as the corrected optimal path direction. ;

[0052] Corrected broadcast volume gain Path pointing after physical connectivity verification and correction Repackage and generate the final optimal resource scheduling instruction.

[0053] Secondly, this application provides a resource scheduling system based on multimodal collaboration and deep reinforcement learning, employing the following technical solution:

[0054] A resource scheduling system based on multimodal collaboration and deep reinforcement learning includes:

[0055] The data acquisition module is used to construct a digital twin simulation environment of the station scene and integrate the visual passenger flow density and auditory environmental noise collected on site in real time to generate a multimodal state matrix that represents the spatial congestion status and information transmission interference level at the current time.

[0056] The vector generation module is used to predict the passenger flow diffusion trend under non-intervention conditions based on the passenger flow inertial motion law, and to fuse and serialize the trend with the multimodal state matrix to generate a composite feature vector containing the current state and future inertia.

[0057] The strategy generation module is used to input the composite feature vector into a preset deep cooperative scheduling model, and use the strategy network embedded in the model based on the dual constraints of maximizing evacuation efficiency and minimizing energy consumption to generate a primary cooperative strategy that includes broadcast volume gain and guide screen path direction.

[0058] The strategy correction module is used to perform secondary constraint verification on the primary cooperative strategy, verify and correct the geometric visibility and physical connectivity of the broadcast sound field and the screen field of view based on the physical space mapping relationship, and determine the final optimal resource scheduling instruction.

[0059] The parameter update module is used to issue the optimal resource scheduling command to drive the on-site broadcast and guidance screen to perform actions, and to monitor the passenger flow evacuation rate and actual energy consumption after execution in real time. Based on the monitoring results, the feedback deviation is calculated to update the network parameters of the deep collaborative scheduling model online.

[0060] Thirdly, this application provides an electronic device, the device including a processor and a memory; the memory stores a program, the program being loaded and executed by the processor to implement a resource scheduling method based on multimodal collaboration and deep reinforcement learning as described in the first aspect.

[0061] Fourthly, this application provides a computer-readable storage medium storing a program that, when executed by a processor, is used to implement a resource scheduling method based on multimodal cooperation and deep reinforcement learning as described in the first aspect.

[0062] This application discloses a resource scheduling method based on multimodal collaboration and deep reinforcement learning. This method constructs a digital twin environment, fuses audiovisual data in real time, and generates composite feature vectors by combining passenger flow inertia prediction. It then uses a deep collaborative model with embedded constraints of maximizing evacuation efficiency and minimizing energy consumption to generate a primary strategy, and performs secondary verification of geometric visibility and connectivity based on physical spatial mapping relationships. This scheme effectively overcomes the shortcomings of existing technologies in dealing with slow response speed and perception lag when facing complex passenger flow environments by utilizing multimodal fusion and trend prediction mechanisms. Through the dual constraint strategy embedded in the deep model, it directly solves the contradiction between evacuation efficiency and operational energy consumption at the decision-making source. In particular, through secondary correction based on physical geometric relationships, it forcibly resolves the problems of scattered guidance instructions, audiovisual mismatch, and poor decision consistency caused by the independent operation of the broadcast and guidance subsystems. Finally, combined with an online feedback update mechanism, it achieves efficient collaborative and adaptive evolution of resource scheduling.

[0063] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, the preferred embodiments of this application are described in detail below with reference to the accompanying drawings. Attached Figure Description

[0064] Figure 1 This is a flowchart illustrating the resource scheduling method based on multimodal collaboration and deep reinforcement learning in the embodiments of this application.

[0065] Figure 2 This is a schematic diagram of the overall process of the resource scheduling method based on multimodal collaboration and deep reinforcement learning in the embodiments of this application.

[0066] Figure 3 This is a structural block diagram of the resource scheduling system based on multimodal collaboration and deep reinforcement learning in the embodiments of this application.

[0067] Figure 4 This is a block diagram of an electronic device based on multimodal collaboration and deep reinforcement learning in an embodiment of this application. Detailed Implementation

[0068] The specific embodiments of this application will be described in further detail below with reference to the accompanying drawings and examples. The following examples are used to illustrate this application, but are not intended to limit the scope of this application.

[0069] Optionally, this application uses the resource scheduling method based on multimodal collaboration and deep reinforcement learning provided in various embodiments as an example for application in an electronic device. The electronic device is a terminal or a server. The terminal can be a computer, tablet computer, etc. This embodiment does not limit the type of electronic device.

[0070] Reference Figure 1 This is a flowchart illustrating a resource scheduling method based on multimodal collaboration and deep reinforcement learning provided in an embodiment of this application. The method includes at least the following steps:

[0071] Step S101: Construct a digital twin simulation environment for the station scene, and integrate the visual passenger flow density and auditory environmental noise collected on site in real time to generate a multimodal state matrix that represents the spatial congestion status and information transmission interference level at the current time.

[0072] In step S101, the aim is to establish a digital mapping foundation connecting the physical station and the algorithmic decision-making system, and to standardize and fuse multi-dimensional sensory data within the station. Existing scheduling systems often only focus on the visual passenger flow, neglecting the physical constraints of the acoustic environment on the effectiveness of broadcast information transmission, resulting in broadcast instructions not being effectively received by passengers in noisy environments. To address this, this step first constructs a gridded digital twin environment based on the station's spatial structure data, then acquires visual distribution data and auditory noise data in the same time and space through a synchronous acquisition mechanism, converts physical decibel values ​​into probabilistic information transmission interference, and finally fuses the two into a multimodal state matrix, thereby achieving a full-dimensional digital representation of the station's operational status.

[0073] Specifically, the first step is to construct a digital twin simulation environment for the station scene. This involves acquiring the station's architectural information model data, extracting the geometric outlines of all passenger-accessible functional areas within the station, and dividing this physical space into a set of uniformly sized two-dimensional discrete grid cells. ,set up Includes OK List a grid of cells, and denote any one of the grid cells as . ,in Indicates row index, Indicates the column index. For each grid cell... The system maps the associated physical attributes, which include the spatial coordinates of the grid center point, the broadcast partition index to which the grid belongs, and the corresponding guide screen device number within the visible range of the grid.

[0074] Secondly, data acquisition and processing based on a preset time window are performed. A synchronous sampling time window is set for the system. Within each time window, real-time video streams are captured using surveillance cameras distributed throughout the station. Existing computer vision crowd counting algorithms are used to identify the number of passengers in each frame of the image, and the identified passenger locations are mapped to corresponding grid cells. In the middle, calculate the grid within this time window. Average visual passenger flow density The unit is people per square meter. Simultaneously, noise sensors deployed in each broadcast zone continuously collect on-site sound data, calculating the corresponding grid. Average ambient noise intensity within this time window The unit is decibel. Based on this, in order to accurately measure the obstruction effect of environmental noise on the transmission of broadcast voice commands, this application does not directly use the decibel value, but instead constructs a nonlinear mapping model based on a sigmoid function to calculate the information transmission interference degree. Considering that the masking effect of human ear on noise has nonlinear characteristics—that is, when the noise is below a certain threshold, its impact on hearing is small, but when the noise exceeds a certain threshold, the interference degree will rise sharply and tend to saturate—a grid is defined. Information transmission interference at the location As shown below:

[0075] ;

[0076] in, is the base of the natural logarithm; This represents the average ambient noise intensity of the grid within the current time window. To ensure speech clarity, a baseline noise threshold is preset based on the rated power of the station's public address system and the characteristics of human hearing. This is the noise sensitivity coefficient, used to adjust the steepness of the gradient of interference degree as a function of noise intensity. It is calculated using this formula. It is a dimensionless value between 0 and 1. The closer the value is to 0, the quieter the environment and the smoother the broadcast information is transmitted; the closer the value is to 1, the stronger the environmental noise interference and the more difficult it is for passengers to clearly identify the broadcast information.

[0077] Finally, a multimodal state matrix is ​​generated based on the above calculation results. The visual passenger flow density is calculated for each grid cell. Normalization process is performed to obtain A pre-defined limit capacity density value is set for each grid cell. The average visual passenger flow density calculated for each grid cell is divided by this limit capacity density value to obtain the normalized visual passenger flow density. Then, the normalized visual passenger flow density and information transmission interference were compared. Together, they form a dimension. The three-dimensional data structure, namely the multimodal state matrix Any position in this matrix The numerical definition is:

[0078] ;

[0079] Among them, the first channel data of the matrix The second channel data of the matrix represents the physical congestion status of the station space during the current time period. This characterizes the level of information transmission interference on the station channel during the current time period. By constructing this multimodal state matrix, the passenger flow distribution characteristics and sound field environment characteristics of the station are simultaneously locked in a unified data structure, completing the accurate data solidification that includes both visual and auditory dimensions.

[0080] Step S102: Predict the passenger flow diffusion trend under no-intervention conditions based on the passenger flow inertial motion law, and fuse and serialize the trend with the multimodal state matrix to generate a composite feature vector containing the current state and future inertia.

[0081] In step S102, this step aims to introduce physical evolution information in the time dimension to compensate for the lag in purely static perception data. Although step S101 obtains the current congestion and noise status, in a real station scenario, passenger flow has a fluid-like physical inertia; that is, without external broadcasts or guidance intervention, dense crowds will naturally spread to surrounding open areas. If this natural evolutionary trend is not considered, the scheduling system may issue erroneous strong intervention instructions for areas that are about to naturally ease in the next second. Therefore, this step uses the diffusion principle in physics to predict the passive drift direction of passenger flow and deeply integrates this prediction information with the current audiovisual state to construct a composite feature vector that includes both the current situation and future inertia.

[0082] Specifically, the first step is to calculate the passenger flow diffusion trend under uninterrupted conditions based on the inertial motion law of passenger flow. Data from the first channel of the multimodal state matrix, i.e., the visual passenger flow density distribution, is read. Based on the potential energy driving principle in fluid dynamics, it is assumed that people will be instinctively driven to permeate from high-density areas to low-density areas without guidance. To quantify this physical trend, data is analyzed for each grid cell. The density change potential at the next time step is calculated using a method based on the local density potential energy difference. A grid is defined. Passenger flow inertial drift intensity This indicator reflects the accumulation or dissipation trend of passenger flow in the area under the influence of inertia. The calculation formula is designed as follows:

[0083] ;

[0084] in, The normalized visual passenger flow density for the current grid; Indicated by grid The neighborhood grid set centered on the application embodiment An eight-neighborhood structure is adopted, which includes up to eight grid cells that are adjacent to the grid in the horizontal, vertical and diagonal directions; for boundary grids or cases where there are physical obstacles, only the actually connected adjacent grids are selected into the set. The number of neighboring grid cells; This is the sum of the normalized densities of the neighborhood grid; The diffusion inertia coefficient characterizes the activity level of crowd movement. It can be determined through offline statistical analysis or simulation calibration based on historical passenger flow data at the station, or optimized as an adjustable hyperparameter during model training. The physical meaning of this formula is: if the current grid density is significantly higher than the average density of the surrounding neighborhood, then... A positive value indicates that passenger flow in the area will spread outwards to release pressure; conversely, if the current grid density is lower than the surrounding area, then... A negative value indicates that surrounding passenger flow will converge and fill the area. A set of meshes is generated by traversing all mesh cells and calculating the corresponding drift intensity. A two-dimensional data set with consistent dimensions, namely, a passenger flow diffusion trend matrix.

[0085] Secondly, multi-dimensional fusion and serialization processing of the data is performed. To achieve unified encoding of the current congestion status, environmental noise interference, and future evolution trends, features from different dimensions need to be spatially aligned. A fusion vector is constructed, retaining the normalized visual passenger flow density from step S101. and information transmission interference And the newly calculated passenger flow inertial drift intensity It is added as a third feature dimension.

[0086] Finally, a composite feature vector is generated. Since the one-dimensional input structure facilitates transmission, the spatially fused data containing three dimensions is flattened. Following the spatial scanning order of the grid, three feature values ​​are extracted sequentially from each grid location. The first and last parts are concatenated to form a long one-dimensional numerical sequence. This sequence is the composite feature vector, which is three times the total number of grid cells. This vector completely encodes the current spatiotemporal physical state and sound field environment information of the station.

[0087] Step S103: Input the composite feature vector into the preset deep collaborative scheduling model, and use the strategy network embedded in the model based on the dual constraints of maximizing evacuation efficiency and minimizing energy consumption to generate a primary collaborative strategy that includes broadcast volume gain and path pointing to the guide screen.

[0088] In step S103, this step aims to utilize the massive parallel computing power of artificial intelligence algorithms to simulate the decision-making logic of human experts, transforming the abstract perception data generated in step S102 into specific equipment control commands. Traditional rule-based systems struggle to handle massive concurrent states and cannot simultaneously balance evacuation speed and energy consumption in a short time. Therefore, this step introduces a deep collaborative scheduling model, which is essentially a pre-trained deep neural network capable of understanding the spatiotemporal relationships in composite feature vectors and directly outputting the optimal control parameters for each broadcast partition and guide screen throughout the station.

[0089] Specifically, the deep cooperative scheduling model in this application is built on a Multi-TaskResNet-50 architecture, which includes three core components: an input reconstruction layer, a global feature encoder, and a dual-constraint policy output layer. The input reconstruction layer is used to restore the spatial structure of the data, the global feature encoder is used to extract deep semantics, and the dual-constraint policy output layer contains two parallel branches: evacuation guidance and energy consumption control. The composite feature vector is input into the model, processed through various layers, and finally decoded into a primary cooperative policy.

[0090] First, spatial reconstruction of the input data is performed. Step S102: Input composite feature vector As a one-dimensional sequence, it lacks spatial adjacency information between grids. The input reconstruction layer uses a dimensionality inverse mapping operation to restore this vector to its relationship with the station grid. OK The three-dimensional spatial feature body corresponding to the column The calculation logic of this process is as follows:

[0091] ;

[0092] in, This represents the inverse dimension mapping operator, which backfills linear data into three-dimensional space according to the grid index order. The generated three-dimensional spatial feature volume includes three physical channels: normalized visual passenger flow density, information transmission interference degree, and passenger flow diffusion trend.

[0093] It should be noted that this application employs a residual network architecture based on convolutional operations. The feature extraction mechanism of this architecture relies on the adjacency relationships of data in physical space; that is, it uses convolutional kernels to slide across a two-dimensional plane to capture passenger flow density gradients and noise diffusion patterns within local areas. Because the one-dimensional vector form of the composite feature vector disrupts the original row and column topology of the grid cells, the model cannot directly perceive the adjacency and continuity characteristics in physical space. Therefore, it is necessary to inversely map the linear composite feature vector through an input reconstruction layer and restore it to a three-dimensional spatial feature volume, recovering its original physical field properties. This allows the model to perform effective feature extraction and decision-making inference based on the real spatial structure of the station.

[0094] Secondly, a global feature encoder is used to extract feature data. The global feature encoder uses ResNet-50 or its variants as the backbone network, consisting of a series of stacked residual convolutional modules. In implementation, ResNet-50 is used as a fully convolutional backbone network, removing the original fully connected classification layers and pooling classification heads. The downsampling stride is configured to allow the feature map to be restored to the same scale as the input grid. Subsequently, upsampling is used to align the output to... Three-dimensional spatial feature volume After inputting into this backbone network, convolutional feature extraction is performed layer by layer. For the first layer of the network... Each residual module outputs feature data. The calculation is as follows:

[0095] ;

[0096] in, For the output features of the previous layer, Characterizes the mapping of convolution operations within this residual module; and This is the set of weights and biases for each convolutional layer within the module; It is the ReLU activation function; The direct addition embodies the residual learning mechanism; when the dimension of the feature map changes, the residual learning mechanism is applied. Perform a linear projection to match the dimensions.

[0097] Furthermore, to ensure that the output semantic features correspond one-to-one with the original grid space and to address the spatial information loss caused by deep network downsampling, the backbone network performs bilinear interpolation upsampling (or transposed convolution) after deep feature extraction to restore the spatial resolution of the feature map to [original grid space value]. After processing by the entire ResNet-50 layer, the final latent feature map is output. This latent feature map is the feature data. Each pixel no longer represents a specific physical value, but rather encodes high-dimensional semantic information about the passenger flow congestion pattern and sound field interference pattern at that location, serving as a common input source for subsequent bi-branch decision-making.

[0098] Finally, a primary cooperative policy is generated from the output layer using a dual-constraint policy. (Latent feature map) The calculations are simultaneously distributed to two independent parallel branches, which respectively compute the guidance instructions and broadcast parameters. The evacuation guidance branch is used to achieve the constraint of maximizing evacuation efficiency. This branch analyzes the density gradient around the mesh using a fully connected layer and calculates the benefit probability of each guide screen's indicated direction. (The last sentence appears to be incomplete and possibly refers to a specific mesh configuration.) Calculate its path pointing probability vector :

[0099] ;

[0100] in, and The weights and biases of this branch; The function maps the output to a probability distribution for four categories: go straight, turn left, turn right, and stop. The model selects the category with the highest probability value as the path direction indicated by the guide screen for that grid. The Softmax function used in this branch is essentially a competitive normalization activation mechanism. Physically, the values ​​in the latent feature map represent the passenger flow pressure gradient around the grid. Through the Softmax operation, the model forces the four discrete movement directions to compete in the same probability space; only the category that follows the downward direction of the passenger flow gradient and has the least movement resistance receives the highest probability value. Therefore, the model selects the path direction with the highest probability. Mathematically, this is equivalent to finding the current optimal flow path within a complex passenger flow potential energy field. This decision-making method based on maximum likelihood probability can guide the crowd to move along the direction with the fastest local flow velocity, thereby ensuring the maximum efficiency of evacuation and passage throughout the station on a macroscopic level.

[0101] The energy consumption control branch is used to achieve the energy consumption minimization constraint. This branch uses regression analysis to determine whether broadcasting must be enabled in the current region. (For the grid...) Calculate the broadcast volume gain :

[0102] ;

[0103] in, and For the parameters of this branch; This is the preset silent bias threshold. The formula uses the Sigmoid activation function and introduces a negative bias, so that when the feature strength is insufficient (i.e., no people or no noise), the output... It automatically approaches 0, outputting high gain only when necessary, thus forcing energy saving at the algorithm's underlying level. The above formula constructs a nonlinear activation structure with a physical suppression threshold. A negative bias is introduced. Its purpose is to construct an energy dead zone. This is used when the strength of the input latent features (i.e., the weighted sum of passenger flow density or environmental noise) fails to exceed this threshold. When the exponential term approaches infinity, the denominator increases dramatically, forcing the output to be suppressed to near zero. This means the model defaults to an inertial state that rejects high energy consumption; neurons are only activated and output high gain when the congestion or noise characteristics of the scene are significant enough to offset the inhibitory effect of the negative bias. This mathematical construction of not activating unless necessary and not responding unless the target is met directly forces the minimization of operating energy consumption at the algorithm's fundamental structural level.

[0104] Summarize the broadcast volume gain calculated from all grids across the entire site. Path pointing to the guide screen These two sets of parameters are matched and combined one-to-one according to their spatial location in the grid to form a control parameter matrix covering the entire station, which is the primary cooperative strategy. It should be noted that the broadcasting system uses broadcast zones as control objects, while the guide screens use devices as control objects. For multiple grid cells belonging to the same broadcast zone, their corresponding broadcast volume gains are aggregated to generate the final gain value for that broadcast zone. Aggregation can be achieved by using the maximum value, the average value, or a weighted sum based on the grid passenger flow density. For multiple grid cells associated with the same guide screen device, their path pointing probability vectors are aggregated, and the direction with the highest probability after aggregation is selected as the path pointing for that guide screen. This strategy directly reflects the original scheduling scheme derived by the deep collaborative scheduling model based on the current perception data. This step sets this dual constraint to solve the multi-objective optimization problem between safety assurance and operational efficiency in the station resource scheduling process. Maximizing evacuation efficiency, as a hard constraint at the safety level, forces the model's evacuation guidance branch to calculate the path of least resistance in a dynamically changing passenger flow field, ensuring that people can be guided to open areas as quickly as possible when there is high-density passenger flow congestion, thereby ensuring the safety of passengers. Minimizing energy consumption, as a constraint from both economic and environmental perspectives, compels the model's energy control branch to possess sparse activation characteristics. This means it automatically suppresses device power output when unnecessary, which not only reduces the system's operating electricity costs but, more importantly, reduces noise pollution caused by ineffective broadcasts to the acoustic environment, thereby improving the signal-to-noise ratio of effective information. By embedding these two mutually restraining constraints into the model structure, it is possible to ensure that the generated scheduling strategy achieves precise and intensive resource utilization while meeting extreme evacuation needs.

[0105] Step S104: Perform secondary constraint verification on the primary coordination strategy. Based on the physical space mapping relationship, verify and correct the geometric visibility and physical connectivity between the broadcast sound field and the screen field of view, and determine the final optimal resource scheduling instruction.

[0106] In step S104, this step aims to introduce hard geometric constraints from the physical world to verify the effectiveness of the initial cooperative strategy generated in step S103. Although the deep cooperative scheduling model achieves convergence of dual optimization at the numerical level, the neural network lacks an intuitive understanding of the occlusion relationship of physical entities and the blind spots of equipment coverage. This may lead to a logical paradox in the strategy that is high in energy consumption but low in return. For example, if a broadcast is played at full power in a blind spot where the guide screen is completely blocked by a pillar, highly concentrated passengers can only hear the sound but cannot obtain specific route guidance, which can easily cause panic or disorderly flow. To address this, this step uses line-of-sight analysis technology based on computational geometry to construct a mapping relationship between sound and light in the physical space. The broadcast gain is corrected by computational geometry occlusion parameters to ensure that every instruction issued is physically executable and has audiovisual consistency.

[0107] Specifically, the geometric visibility parameters of the entire site's grid are first calculated. Based on the static building model in the digital twin environment, the geometric optics calculation logic is constructed. This is done for each grid cell. The center point is taken as the observation origin, and a line-of-sight vector is formed by connecting this origin with the geometric center of the guide screen display surface. The angle between this line-of-sight vector and the normal to the guide screen surface is calculated. Simultaneously, the straight-line Euclidean distance from the origin to the guide screen is measured. To objectively quantify visual acuity, this application introduces the relative solid angle ratio method to calculate the geometric visibility coefficient. .

[0108] The actual observed solid angle of the current grid relative to the guide screen is calculated using the solid angle approximation formula. :

[0109] ;

[0110] in, This refers to the physical display area of ​​the guide screen. If there are physical obstacles (such as pillars) in the line-of-sight vector path, then directly set... Set a reference solid angle. As a normalization benchmark, Defined as the optimal viewing distance that meets the standards of human visual ergonomics. (For example, set to be 5 meters away from the screen) and directly facing the screen (i.e.) The ideal solid angle value corresponding to the given time is calculated as follows: Calculate the geometric visibility coefficient. :

[0111] ;

[0112] The physical meaning of this formula lies in comparing the current actual observation effect with the ideal optimal observation effect. If the ratio equals 1, it means that the current location is within the golden observation zone; if the ratio is less than 1, it means that the screen's proportion in the field of view is reduced due to excessive distance or an excessively biased observation angle, resulting in a decrease in visual effect. Through this dimensionless ratio calculation, the physical fact of "how much screen content the human eye can actually see clearly at the current location" is precisely quantified.

[0113] Secondly, the broadcast volume gain in the primary cooperative strategy is corrected using a dynamic hyperbolic attenuation formula based on the risk of audiovisual misalignment. This application constructs the following correction formula:

[0114] ;

[0115] in, The corrected broadcast volume gain; The original gain output in step S103; Geometric visibility coefficient; The normalized visual passenger flow density obtained in step S101; Basic penalty factor; This is the risk amplification factor. and Take a non-negative value, and , This ensures the correction factor , making In Within the range. The design logic of the above formula is based on the physical law that crowding exacerbates the risk of audiovisual mismatch. In low-density scenarios ( Smaller), even if the grid is in a visual blind spot ( At lower density traffic conditions, passengers still have sufficient spatial freedom to locate the sound source. In this case, the error tolerance of the broadcast guidance is higher, therefore the penalty term in the correction formula increases more slowly, allowing for the retention of a certain broadcast volume. However, in high-density congestion scenarios (…), (The density is relatively high), passengers' freedom of movement is restricted and their psychological anxiety is increased. If they are in a state of audiovisual separation where they can hear announcements but cannot see signs, it is very easy for crowds to rush into each other or remain stuck in place while searching for directions. Therefore, this application utilizes the density term As a risk amplifier, this factor forces... The function rapidly saturates and approaches 1, leading to the calculation result... Quickly return to zero. This design mathematically enforces a safety strategy that necessitates silence to prevent panic, especially in congested and chaotic environments.

[0116] Simultaneously, physical connectivity is verified for the path pointing on the guide screen. The path pointing output in step S103 is read. A topological adjacency matrix of the entire station grid is constructed based on the station map. This matrix records the physical connectivity status (connectivity or obstruction) between each grid and its eight neighboring grids. The verification logic is as follows: Check instructions. The indicated movement direction is marked as blocked in the topological adjacency matrix (e.g., pointing to a wall, pillar, or non-operational area). If marked as blocked, the primary strategy is deemed invalid, and a local neighborhood optimization algorithm is initiated: using the current grid... Centered on a given point, traverse all its topologically connected neighboring grids; calculate the actual path distance from the center point of each of these neighboring grids to the nearest safe exit; select the direction corresponding to the neighboring grid with the smallest distance and that is passable as the corrected optimal path direction. This process ensures that all directional commands are absolutely passable in physical space, avoiding the issuance of infinite loops or wall-hitting commands.

[0117] Finally, the optimal resource scheduling instruction is determined. The broadcast volume gain, corrected using the dynamic hyperbolic attenuation formula, is then used. Path pointing after physical connectivity verification and correction The system is repackaged to generate the final optimal resource scheduling instructions. This instruction set not only inherits the advantages of artificial intelligence algorithms in spatiotemporal dynamic extrapolation, but also ensures the executability and security of the scheduling scheme in a real physical environment through rigorous physical geometry verification.

[0118] Step S105: Issue the optimal resource scheduling command to drive the on-site broadcast and guidance screen to perform actions, and monitor the passenger flow evacuation rate and actual energy consumption after execution in real time. Calculate the feedback deviation based on the monitoring results to update the network parameters of the deep collaborative scheduling model online.

[0119] In step S105, the aim is to establish a closed-loop control mechanism, transforming unidirectional command issuance into bidirectional feedback learning. Due to the high uncertainty of passenger flow dynamics at the station, statically trained models may not perfectly adapt to all unforeseen scenarios. Therefore, this step quantifies the actual performance of evacuation efficiency and energy consumption by real-time monitoring of the physical effects after command execution. This performance is then compared with the theoretical optimal value to calculate the deviation. The deviation value is used to fine-tune the model's weight parameters online, enabling the deep collaborative scheduling model to possess the self-evolutionary ability to continuously adapt to changes in the on-site environment.

[0120] Specifically, the optimal resource scheduling command is first issued and executed. The optimal resource scheduling command determined in step S104 is issued to the field equipment layer through the IoT control gateway. For the wayfinding screens, the path pointing command is parsed into a specific graphic display signal; for the broadcasting system, the volume gain command is converted into the control level of the power amplifier. Each physical broadcasting zone and wayfinding screen executes actions synchronously according to the received command, thereby exerting physical intervention on the on-site passenger flow.

[0121] Secondly, implement real-time monitoring and quantification of execution effectiveness. This involves a preset observation period after the instruction is executed. Within this process, the sensing network from step S101 is used to collect on-site data again, and the actual passenger flow evacuation rate and actual operating energy consumption are calculated respectively. The actual passenger flow evacuation rate is defined as follows: This indicator represents the rate of decrease of normalized passenger flow density across the entire station per unit time. Its calculation formula is as follows:

[0122] ;

[0123] in, The normalized visual passenger flow density of all grids in the entire station at the time of instruction execution; Time interval Normalized visual passenger flow density after normalization; and These represent the number of rows and columns of the grid, respectively. This formula quantifies the current evacuation effectiveness by calculating the differential change in the total passenger flow density of the entire station; a larger value indicates a faster evacuation.

[0124] Redefining actual operating energy consumption This indicator represents the total electrical energy consumed by all broadcasting equipment during the observation period, and its calculation formula is as follows:

[0125] ;

[0126] in, The rated reference power for a single broadcast zone; The corrected broadcast volume gain (range 0-1) issued in step S104; since power is proportional to the square of the amplitude, the square term of the gain is used. Perform the calculation.

[0127] Based on this, a dynamically weighted loss function based on congestion urgency is constructed to calculate feedback bias. Traditional fixed-weight loss functions are difficult to adapt to changing on-site conditions. This application proposes a loss calculation logic with adaptive adjustment capabilities: under high congestion conditions, the weight of evacuation rate is exponentially amplified while energy consumption costs are ignored; while under low congestion conditions, the weight of energy consumption costs is linearly amplified. Loss Function The definition is as follows:

[0128] ;

[0129] in, It is a natural constant; The sensitivity coefficient is dynamically adjusted to control the degree of weight change with density. The average normalized passenger flow density of the entire station represents the overall congestion urgency. The theoretical maximum evacuation rate designed for the station is used for... Perform dimensionless processing; The set energy consumption tolerance limit is used to... Dimensionless processing is performed. The design of this formula utilizes an exponential function. and A set of mutually exclusive dynamic weights was constructed: when When the congestion approaches 1 (extreme congestion), the weight of the preceding term... A dramatic increase means that as long as the evacuation rate... If the maximum value is not reached, the system will generate a huge penalty value, and at this time the subsequent weights... Approaching 0 means that the system can completely tolerate high energy consumption at this point; when When the flow approaches zero (sparse passenger flow), the weight of the first term tends to level off, while the weight of the second term increases, forcing the model to shift its optimization focus to reducing energy consumption. Through this mathematical construction, the model is forced to learn the intelligent decision-making logic of prioritizing safety at all costs during crises and carefully calculating energy consumption reduction during normal times.

[0130] Finally, the parameters of the deep cooperative scheduling model are updated online. The calculated loss function values ​​are then... As a benchmark for measuring model decision error, minimizing this loss value is used as the optimization objective. The error signal is then propagated back to the deep collaborative scheduling model via backpropagation. Based on a preset learning rate, the weight parameters of each layer in the model are directly fine-tuned. This process guides the model's parameter values ​​towards reducing feedback bias, thus solidifying the experience of each successful scheduling at the parameter level and ensuring that the model can automatically output better strategies when facing similar scenarios in the future.

[0131] In summary, combining Figure 2 This application discloses a resource scheduling method based on multimodal collaboration and deep reinforcement learning. The method first constructs a digital twin simulation environment of a station scene, fusing real-time visual passenger flow density and auditory environmental noise collected on-site to generate a multimodal state matrix representing the current spatial congestion state and information transmission interference. Then, based on the inertial motion law of passenger flow, it predicts the passenger flow diffusion trend under no-intervention conditions and fuses this trend with the multimodal state matrix to serialize it into a composite feature vector containing the current state and future inertia. Next, this vector is input into a preset deep collaborative scheduling model, utilizing a strategy network embedded in the model based on dual constraints of maximizing evacuation efficiency and minimizing energy consumption to generate a primary collaborative strategy including broadcast volume gain and directional screen path pointing. Based on this, a secondary constraint verification is performed on the primary collaborative strategy, verifying and correcting the geometric visibility and physical connectivity of the broadcast sound field and the screen's visual field based on physical spatial mapping relationships to determine the final optimal resource scheduling instruction. Finally, instructions are issued to drive on-site equipment execution, and the passenger flow evacuation rate and actual energy consumption after execution are monitored in real time. Feedback deviations are calculated based on the monitoring results to update the network parameters of the deep collaborative scheduling model online.

[0132] This method first constructs a digital twin environment and integrates visual passenger flow density with auditory environmental noise. It then generates a composite feature vector based on passenger flow inertia patterns. This multimodal, forward-looking perception mechanism effectively overcomes the response lag caused by traditional scheduling relying on human experience, enabling rapid capture of dynamic passenger flow trends. Building upon this, a primary strategy is generated using a deep collaborative scheduling model with embedded constraints of maximizing evacuation efficiency and minimizing energy consumption. This directly seeks a balance between safe evacuation and energy conservation at the algorithmic level, solving the problem that a single rule cannot simultaneously consider both evacuation efficiency and operational energy consumption. Furthermore, a secondary verification mechanism based on physical space mapping is introduced to correct the geometric visibility and physical connectivity of the broadcast sound field and screen view, breaking down the barriers to independent subsystem operation and ensuring high coordination and consistency between visual guidance and auditory broadcasting in physical space, eliminating invalid or contradictory guidance instructions. Finally, by monitoring the evacuation rate and energy consumption after execution in real time and updating model parameters online using feedback deviations, the system is endowed with adaptive evolution capabilities, fundamentally improving the resource scheduling response speed and decision-making accuracy in complex passenger flow environments.

[0133] Figure 3 This is a structural block diagram of a resource scheduling system based on multimodal collaboration and deep reinforcement learning, provided in one embodiment of this application. The system includes at least the following modules:

[0134] The data acquisition module is used to construct a digital twin simulation environment of the station scene and integrate the visual passenger flow density and auditory environmental noise collected on site in real time to generate a multimodal state matrix that represents the spatial congestion status and information transmission interference level at the current time.

[0135] The vector generation module is used to predict the passenger flow diffusion trend under non-intervention conditions based on the inertial motion law of passenger flow, and to fuse and serialize the trend with the multimodal state matrix to generate a composite feature vector containing the current state and future inertia.

[0136] The strategy generation module is used to input composite feature vectors into a preset deep collaborative scheduling model. It uses the strategy network embedded in the model based on the dual constraints of maximizing evacuation efficiency and minimizing energy consumption to generate a primary collaborative strategy that includes broadcast volume gain and path pointing to the guide screen.

[0137] The strategy correction module is used to perform secondary constraint verification on the primary coordination strategy, and to verify and correct the geometric visibility and physical connectivity between the broadcast sound field and the screen field of view based on the physical space mapping relationship, so as to determine the final optimal resource scheduling instruction.

[0138] The parameter update module is used to issue optimal resource scheduling instructions to drive the on-site broadcasting and guidance screens to perform actions and monitor the passenger flow evacuation rate and actual energy consumption in real time after execution. Based on the monitoring results, it calculates the feedback deviation to update the network parameters of the deep collaborative scheduling model online.

[0139] For relevant details, please refer to the above method implementation examples.

[0140] Figure 4 This is a block diagram of an electronic device provided in one embodiment of this application. The device includes at least a processor 401 and a memory 402.

[0141] Processor 401 may include one or more processing cores, such as a quad-core processor or an octa-core processor. Processor 401 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 401 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 401 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, processor 401 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.

[0142] Memory 402 may include one or more computer-readable storage media, which may be non-transitory. Memory 402 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in memory 402 is used to store at least one instruction, which is executed by processor 401 to implement the resource scheduling method based on multimodal cooperation and deep reinforcement learning provided in the method embodiments of this application.

[0143] In some embodiments, the electronic device may also optionally include: a peripheral device interface and at least one peripheral device. The processor 401, memory 402, and peripheral device interface can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface via a bus, signal line, or circuit board. Indicatively, peripheral devices include, but are not limited to: radio frequency circuits, touch displays, audio circuits, and power supplies.

[0144] Of course, electronic devices may also include fewer or more components, and this embodiment does not limit this.

[0145] Optionally, this application also provides a computer-readable storage medium storing a program that is loaded and executed by a processor to implement the resource scheduling method based on multimodal cooperation and deep reinforcement learning in the above-described method embodiments.

[0146] Optionally, this application also provides a computer product including a computer-readable storage medium storing a program, which is loaded and executed by a processor to implement the resource scheduling method based on multimodal cooperation and deep reinforcement learning described in the above method embodiments.

[0147] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0148] The above embodiments merely illustrate several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A resource scheduling method based on multimodal cooperation and deep reinforcement learning, characterized in that, The method includes: A digital twin simulation environment for the station scene is constructed, and the visual passenger flow density and auditory environmental noise collected on site are integrated in real time to generate a multimodal state matrix that represents the spatial congestion status and information transmission interference level at the current time. Based on the inertial motion law of passenger flow, the trend of passenger flow diffusion under non-intervention conditions is predicted, and the trend is fused and serialized with the multimodal state matrix to generate a composite feature vector containing the current state and future inertia. The composite feature vector is input into a preset deep collaborative scheduling model. The model uses a strategy network embedded in the model that is based on the dual constraints of maximizing evacuation efficiency and minimizing energy consumption to generate a primary collaborative strategy that includes broadcast volume gain and path pointing to the guide screen. The primary coordination strategy undergoes secondary constraint verification. Based on the physical space mapping relationship, the geometric visibility and physical connectivity between the broadcast sound field and the screen view are verified and parameters are corrected to determine the final optimal resource scheduling instruction, including: With each grid cell The center point is taken as the observation origin. A line-of-sight vector is formed by connecting this origin with the geometric center of the guide screen display surface. The angle between the line-of-sight vector and the normal to the guide screen surface is calculated. Simultaneously, the straight-line Euclidean distance from the origin to the guide screen is measured. The actual observed solid angle of the current grid relative to the guide screen is calculated using the solid angle approximation formula. : ; in, Let be the physical display area of ​​the guide screen. If there are physical obstacles on the line-of-sight vector path, then let Set reference solid angle As a normalization benchmark, Defined as the optimal viewing distance that meets the standards of human visual ergonomics. And the ideal solid angle value corresponding to when facing the screen is calculated as follows: ; Calculate the geometric visibility coefficient : ; Construct the following correction formula: ; in, The corrected broadcast volume gain; This is the original gain; Geometric visibility coefficient; To normalize visual passenger flow density; Basic penalty factor; This is a risk amplification factor. Read path pointing to Based on the station map, a topological adjacency matrix of the entire station grid is constructed. The matrix records the physical access status between each grid and its eight neighboring grids, i.e., connectivity or obstruction. The verification logic is as follows: Check instructions. If the indicated movement direction is marked as blocked in the topological adjacency matrix, then the primary cooperative strategy is deemed invalid, and a local neighborhood selection algorithm is initiated: using the current grid... Centered on a given point, traverse all its topologically connected neighboring grids; calculate the actual path distance from the center point of each neighboring grid to the nearest safe exit; select the direction corresponding to the neighboring grid with the smallest distance and that is passable as the corrected optimal path direction. ; Corrected broadcast volume gain Path pointing after physical connectivity verification and correction Repackage and generate the final optimal resource scheduling instruction; The optimal resource scheduling command is issued to drive the on-site broadcasting and guidance screen to perform actions, and the passenger flow evacuation rate and actual energy consumption after execution are monitored in real time. Based on the monitoring results, the feedback deviation is calculated to update the network parameters of the deep collaborative scheduling model online.

2. The resource scheduling method based on multimodal collaboration and deep reinforcement learning according to claim 1, characterized in that, The constructed digital twin simulation environment for the station scene, which integrates real-time visual passenger flow density and auditory environmental noise collected on-site, generates a multimodal state matrix representing the spatial congestion status and information transmission interference level at the current time period, including: The physical space of the station is divided into a set of uniformly sized two-dimensional discrete grid cells. ,set up Includes OK List a grid of cells, and denote any one of the grid cells as . ,in Indicates row index, Indicates column index; Calculate the grid by capturing real-time video streams using a camera. Average visual passenger flow density Noise sensors are used to collect ambient sound data, and grid calculations are performed. Average ambient noise intensity ; Define grid Information transmission interference at the location As shown below: ; in, The reference noise threshold; This is the noise sensitivity coefficient; Visual passenger flow density calculated for each grid cell Normalization process is performed to obtain Normalized visual passenger flow density and information transmission interference Together, they form a dimension. Multimodal state matrix Any position in this matrix The numerical definition is: .

3. The resource scheduling method based on multimodal cooperation and deep reinforcement learning according to claim 2, characterized in that, The method of predicting the passenger flow diffusion trend under no-intervention conditions based on the inertial motion law of passenger flow, and fusing and serializing this trend with the multimodal state matrix to generate a composite feature vector containing the current state and future inertia includes: For each grid cell Define the grid Passenger flow inertial drift intensity The calculation formula is as follows: ; in, The normalized visual passenger flow density for the current grid; Indicated by grid The set of neighborhood grids centered on; The number of neighboring grid cells; This is the sum of the normalized densities of the neighboring grid; The diffusion inertia coefficient; Following the spatial scanning order of the grid, extract three feature values ​​at each grid location sequentially. Connect the first and last parts of the vector to form a composite feature vector.

4. The resource scheduling method based on multimodal collaboration and deep reinforcement learning according to claim 2, characterized in that, The step of inputting the composite feature vector into a preset deep cooperative scheduling model, and using the policy network embedded in the model based on the dual constraints of maximizing evacuation efficiency and minimizing energy consumption, generates a primary cooperative strategy that includes broadcast volume gain and guide screen path direction, including: The deep collaborative scheduling model includes an input reconstruction layer, a global feature encoder, and a dual-constraint policy output layer. The dual-constraint policy output layer includes an evacuation guidance branch and an energy consumption control branch. The input reconstruction layer uses inverse dimension mapping to transform the composite feature vector. Restored to the station grid OK The three-dimensional spatial feature body corresponding to the column The calculation logic is as follows: ; in, Indicates the dimensional inverse mapping operator; The global feature encoder is composed of stacked residual modules, and the three-dimensional spatial feature volume After inputting the global feature encoder, convolutional feature extraction is performed layer by layer. For the first layer of the network... Each residual module outputs feature data. The calculation is as follows: ; in, For the output features of the previous layer, Characterizes the mapping of convolution operations within this residual module; and This is the set of weights and biases for each convolutional layer within the module; The ReLU activation function is used; after deep feature extraction, the global feature encoder performs bilinear interpolation upsampling to restore the spatial resolution of the feature map to [value missing]. After full-layer processing, the latent feature map is output. .

5. The resource scheduling method based on multimodal collaboration and deep reinforcement learning according to claim 4, characterized in that, The step of inputting the composite feature vector into a preset deep cooperative scheduling model, and using the policy network embedded in the model based on the dual constraints of maximizing evacuation efficiency and minimizing energy consumption to generate a primary cooperative strategy that includes broadcast volume gain and guide screen path direction, further includes: The latent feature map It is simultaneously distributed to two branches, which respectively compute the guiding instructions and broadcast parameters; The evacuation guidance branches are used to achieve the constraint of maximizing evacuation efficiency. The density gradient around the mesh is analyzed through a fully connected layer, and the probability of gain for each guide screen's indicated direction is calculated for the mesh. Calculate its path pointing probability vector : ; in, and The weights and biases of this branch; The function maps the output to a probability distribution of four categories: go straight, turn left, turn right, and stop. The category with the highest probability value is selected as the path pointing to the guide screen for that grid. ; The energy consumption control branch is used to achieve the energy consumption minimization constraint. It uses regression analysis to determine whether broadcasting must be enabled in the current region, specifically for the grid. Calculate the broadcast volume gain : ; in, and For the parameters of this branch; This is the preset silent bias threshold; Summarize the broadcast volume gain calculated from all grids Path pointing to the guide screen The two sets of parameters are matched and combined one-to-one according to their spatial location in the grid to form a primary cooperative strategy. .

6. A resource scheduling system based on multimodal collaboration and deep reinforcement learning, characterized in that, include: The data acquisition module is used to construct a digital twin simulation environment of the station scene and integrate the visual passenger flow density and auditory environmental noise collected on site in real time to generate a multimodal state matrix that represents the spatial congestion status and information transmission interference level at the current time. The vector generation module is used to predict the passenger flow diffusion trend under non-intervention conditions based on the passenger flow inertial motion law, and to fuse and serialize the trend with the multimodal state matrix to generate a composite feature vector containing the current state and future inertia. The strategy generation module is used to input the composite feature vector into a preset deep cooperative scheduling model, and use the strategy network embedded in the model based on the dual constraints of maximizing evacuation efficiency and minimizing energy consumption to generate a primary cooperative strategy that includes broadcast volume gain and guide screen path direction. The strategy correction module is used to perform secondary constraint verification on the primary cooperative strategy, verify and correct the geometric visibility and physical connectivity between the broadcast sound field and the screen view based on the physical space mapping relationship, and determine the final optimal resource scheduling instruction, including: With each grid cell The center point is taken as the observation origin. A line-of-sight vector is formed by connecting this origin with the geometric center of the guide screen display surface. The angle between the line-of-sight vector and the normal to the guide screen surface is calculated. Simultaneously, the straight-line Euclidean distance from the origin to the guide screen is measured. The actual observed solid angle of the current grid relative to the guide screen is calculated using the solid angle approximation formula. : ; in, Let be the physical display area of ​​the guide screen. If there are physical obstacles on the line-of-sight vector path, then let Set reference solid angle As a normalization benchmark, Defined as the optimal viewing distance that meets the standards of human visual ergonomics. And the ideal solid angle value corresponding to when facing the screen is calculated as follows: ; Calculate the geometric visibility coefficient : ; Construct the following correction formula: ; in, The corrected broadcast volume gain; This is the original gain; Geometric visibility coefficient; To normalize visual passenger flow density; Basic penalty factor; This is a risk amplification factor. Read path pointing to Based on the station map, a topological adjacency matrix of the entire station grid is constructed. The matrix records the physical access status between each grid and its eight neighboring grids, i.e., connectivity or obstruction. The verification logic is as follows: Check instructions. If the indicated movement direction is marked as blocked in the topological adjacency matrix, then the primary cooperative strategy is deemed invalid, and a local neighborhood selection algorithm is initiated: using the current grid... Centered on a given point, traverse all its topologically connected neighboring grids; calculate the actual path distance from the center point of each neighboring grid to the nearest safe exit; select the direction corresponding to the neighboring grid with the smallest distance and that is passable as the corrected optimal path direction. ; Corrected broadcast volume gain Path pointing after physical connectivity verification and correction Repackage and generate the final optimal resource scheduling instruction; The parameter update module is used to issue the optimal resource scheduling command to drive the on-site broadcast and guidance screen to perform actions, and to monitor the passenger flow evacuation rate and actual energy consumption after execution in real time. Based on the monitoring results, the feedback deviation is calculated to update the network parameters of the deep collaborative scheduling model online.

7. An electronic device, characterized in that, The device includes a processor and a memory; the memory stores a program, which is loaded and executed by the processor to implement a resource scheduling method based on multimodal collaboration and deep reinforcement learning as described in any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, The storage medium stores a program, which, when executed by a processor, is used to implement a resource scheduling method based on multimodal collaboration and deep reinforcement learning as described in any one of claims 1 to 5.