A safety warning robot control method and system with intelligent voice broadcasting
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 武汉船舶职业技术学院
- Filing Date
- 2026-05-08
- Publication Date
- 2026-08-07
AI Technical Summary
[0004]针对以上问题,本申请提供一种具有智能语音广播的安全预警机器人控制方法及系统,用于解决复杂环境中多源数据难以融合和关联导致对危险的预警出现偏差问题
Smart Images

Figure FT_1 
Figure FT_2
Abstract
Description
Technical Field
[0001] This application belongs to the field of robot control technology, specifically a safety early warning robot control method and system with intelligent voice broadcasting. Background Technology
[0002] In modern society, safety early warning technology has become a core pillar for protecting people's lives and property, especially in high-risk areas such as public places like shopping malls and concert venues, as well as chemical plants in industrial zones, where its role is increasingly prominent. With the popularization of intelligent devices, safety early warning systems are becoming a focus of technological research, aiming to prevent accidents by monitoring potential risks in real time and issuing alerts.
[0003] However, while existing safety warning devices have made progress in basic functions, they struggle to integrate and accurately analyze multi-source, heterogeneous information. For example, changes in crowd density captured by video surveillance and screams or abnormal buzzing detected by audio sensors often originate from devices from different manufacturers. The data collected by video surveillance and audio sensors differ significantly in sampling frequency, transmission protocols, and accuracy, leading to timestamp misalignment or format incompatibility, making it impossible to seamlessly stitch together a complete risk profile. Simultaneously, user behavior information, such as an individual's fall posture, frantic running trajectory, or unusual dwell time, typically comes from wearable devices or behavior recognition algorithms, resulting in a disconnect from the data collected by video surveillance and audio sensors. Therefore, in complex environments, the difficulty in fusing and correlating multi-source data leads to inaccurate warnings of danger. Summary of the Invention
[0004] To address the above issues, this application provides a safety warning robot control method and system with intelligent voice broadcasting, which solves the problem of deviations in warnings of danger caused by the difficulty in integrating and correlating multi-source data in complex environments.
[0005] To achieve the above objectives, the technical solution adopted in this application is as follows:
[0006] The first aspect of this application provides a method for controlling a safety warning robot with intelligent voice broadcasting, including: Acquire video data and emotional fluctuation data, and synchronize the video data and emotional fluctuation data with timestamps to obtain synchronized video data and synchronized emotional data; Object position features, person position features, and person behavior features are extracted from the synchronized video data. The object position features, person position features, and person behavior features are then standardized and corrected to obtain a multidimensional information set. Construct an environmental and behavioral feature map based on the multidimensional information set and synchronized sentiment data; Based on the environmental and behavioral feature maps, potential conflict areas are identified, and abnormal events in the potential conflict areas are analyzed to obtain abnormal statistical results. Based on the abnormal statistical results and historical accidents, the potential hazard type of the abnormal event is determined. A risk index is obtained based on the potential hazard type and risk level standard of the abnormal event. If the risk index of the abnormal event exceeds the risk threshold, an early warning notification is generated.
[0007] In some embodiments, object position features, person position features, and person behavior features are extracted from the synchronized video data, including: A shared encoder is used to extract shared features from the synchronized video data; Convolution is performed on the shared features to obtain object position features and person position features; The location features of the person are mapped to shared features to obtain the region of interest. The region of interest is then pooled and convolved to obtain the behavioral features of the person.
[0008] In some embodiments, the standardization and data correction of the object position features, person position features, and person behavior features to obtain a multidimensional information set includes: The Z-score normalization algorithm was used to standardize the object position features, person position features, and person behavior features to obtain a standard dataset. Acquire light intensity change data, adjust the correction factor based on the light intensity change data, and correct the standard data set based on the adjusted correction factor to obtain a multidimensional information set.
[0009] In some embodiments, constructing an environmental and behavioral feature map based on the multidimensional information set and synchronized sentiment data includes: Calculate the density features of the people based on their location features in the multidimensional information set; The synchronous sentiment data is classified using a support vector machine to determine the correlation between the synchronous sentiment data and the density features of the individuals. Based on the aforementioned relationships, the object location features, person location features, and person behavior features in the multidimensional information set are overlaid to obtain an environmental and behavioral feature map.
[0010] In some embodiments, the step of identifying potential conflict areas based on the environmental and behavioral feature map, analyzing abnormal events in the potential conflict areas, and obtaining abnormal statistical results includes: The environmental and behavioral feature maps were analyzed using a density clustering algorithm to identify potential conflict regions. Obtain the individual trajectory of the person, which is a coordinate sequence with a timestamp, and use the individual trajectory in the potential conflict area as the target trajectory; Extract the path features of the target trajectory, perform time series analysis on the path features to obtain feature trends, take the feature trends that exceed the preset behavior deviation threshold as target feature trends, take the target trajectory corresponding to the target feature trends as abnormal trajectories, count the abnormal proportion of abnormal trajectories in the potential conflict area, and mark the potential conflict area as having an abnormal event when the abnormal proportion in the potential conflict area exceeds the proportion threshold. A sliding window method was used to statistically analyze the frequency, duration, and affected area of abnormal events within potential conflict zones, yielding anomaly statistics.
[0011] In some embodiments, the anomaly statistics include the frequency of occurrence of the anomaly, the duration of the anomaly, and the area affected by the anomaly. Determining the potential hazard type of the anomaly based on the anomaly statistics and historical incidents includes: A historical accident database is pre-established. The historical accidents in the database include: abnormal event types and hazard types. In the historical accident database, the occurrence frequency of each hazard type is counted for each abnormal event type. The similarity between each abnormal event is calculated based on the occurrence frequency, duration, and affected area. The abnormal events are then classified according to the similarity to obtain the abnormal event type. Based on the abnormal event type of each abnormal event, the hazard type that occurs most frequently in historical accidents under each abnormal event type is taken as the potential hazard type of each abnormal event.
[0012] In some embodiments, obtaining the risk index based on the potential hazard type and risk level criteria of the abnormal event includes: Determine the severity index based on the type of potential hazard; Based on the frequency, duration, and affected area of the aforementioned abnormal events, the risk level standard is queried to obtain the risk level; The risk index is calculated based on the severity index and risk level.
[0013] In some embodiments, generating an early warning notification when the risk index of the abnormal event exceeds a risk threshold includes: Anomalies whose risk index exceeds the risk threshold are identified as target anomalies, and the target locations and target intensities of multiple target sensors corresponding to the detected target anomalies are extracted. A triangulation algorithm is used to calculate the locations and intensities of multiple targets to obtain the coordinates of the hazard source, and high-risk areas are determined based on the coordinates of the hazard source. Extract the population gathering points in the high-risk area, calculate the initial paths from the population gathering points to all refuge points, eliminate the initial paths with congestion coefficients exceeding the congestion threshold to obtain candidate paths, and determine the evacuation path for each candidate path by weighted calculation based on travel time and safety factor. Early warning notifications are generated based on the coordinates of the hazard source, high-risk areas, and evacuation routes.
[0014] In some embodiments, after generating the warning notification, the following is included: A continuous monitoring mechanism is used to iteratively analyze video data and emotional fluctuation data. After a new abnormal event is detected, the abnormal statistics are updated based on the new abnormal event, and the early warning notification is updated based on the updated abnormal statistics and historical incidents.
[0015] The second aspect of this application provides a safety warning robot control system with intelligent voice broadcasting, including: The synchronization module is used to acquire video data and emotional fluctuation data, and to synchronize the video data and emotional fluctuation data with timestamps to obtain synchronized video data and synchronized emotional data. The extraction and processing module is used to extract object position features, person position features, and person behavior features from the synchronized video data, and to perform standardization processing and data correction on the object position features, person position features, and person behavior features to obtain a multi-dimensional information set. The graph construction module is used to construct an environmental and behavioral feature graph based on the multidimensional information set and synchronized sentiment data. The identification and analysis module is used to identify potential conflict areas based on the environmental and behavioral feature map, analyze abnormal events in the potential conflict areas, and obtain abnormal statistical results. The judgment and early warning module is used to determine the potential danger type of the abnormal event based on the abnormal statistical results and historical accidents, obtain a risk index based on the potential danger type and risk level standard of the abnormal event, and generate an early warning notification when the risk index of the abnormal event exceeds the risk threshold.
[0016] The technical solution provided in this application has the following advantages and effects: It integrates video data and emotional fluctuation data synchronously through timestamps, then extracts object location features, person location features, and person behavior features from the video data to construct a standardized multi-dimensional information set. Based on this multi-dimensional information set and synchronized emotional data, it constructs an environmental and behavioral feature map, and then accurately identifies potential conflict areas and abnormal events. Combined with historical incident comparisons, it automatically generates early warning notifications when the risk index of an abnormal event exceeds a risk threshold. Attached Figure Description
[0017] Figure 1 A flowchart illustrating the control method for a safety warning robot with intelligent voice broadcasting provided in this application.
[0018] Figure 2 This is a structural block diagram of the safety early warning robot control system with intelligent voice broadcasting provided in this application. Detailed Implementation
[0019] To enable those skilled in the art to better understand the technical solution, the present application will be described in detail below with reference to the embodiments. The description in this section is only exemplary and explanatory, and should not be used to limit the scope of protection of the present application in any way.
[0020] like Figure 1 As shown, this embodiment provides a method for controlling a safety warning robot with intelligent voice broadcasting, including the following steps S1-S5:
[0021] Step S1: Obtain video data and emotional fluctuation data, and synchronize the video data and emotional fluctuation data with timestamps to obtain synchronized video data and synchronized emotional data.
[0022] In practical applications, video data can be captured via a camera, and audio data can be captured via a microphone array. Preprocessing of the audio data, such as frame segmentation and noise reduction, is performed. Audio features (such as fundamental frequency, energy, and speech rate) are extracted frame by frame from the preprocessed audio data to obtain emotional fluctuation data. During the acquisition of both video and emotional fluctuation data, a timestamp is assigned to each frame. For example, during the frame-by-frame extraction of video images, a timestamp is assigned to each image to obtain the video data's timestamp. Similarly, during the acquisition of audio data, a timestamp is assigned to each audio frame to obtain the emotional fluctuation data's timestamp. The timestamps of both the video and emotional fluctuation data are then unified to a UTC timestamp with a precision of 1ms. A linear interpolation algorithm is used to adjust the offset. For example, if the video data's timestamp is t1 = 1000ms and the emotional fluctuation data's timestamp is t2 = 1005ms, a 5ms difference is calculated and interpolated to ensure a synchronization error of less than 2ms, achieving timestamp synchronization between the video and emotional fluctuation data, resulting in synchronized video and emotional data.
[0023] Step S2: Extract object position features, person position features, and person behavior features from the synchronized video data, and perform standardization and data correction on the object position features, person position features, and person behavior features to obtain a multi-dimensional information set.
[0024] In practical applications, object and person position features are obtained by visual inspection of synchronized video data, and person behavior features are obtained by behavioral feature recognition of synchronized video data. Then, the object, person, and behavior features are standardized and corrected to facilitate subsequent processing.
[0025] Specifically, object position features, person position features, and person behavior features are extracted from the synchronized video data, including:
[0026] A shared encoder is used to extract shared features from the synchronized video data;
[0027] Convolution is performed on the shared features to obtain object position features and person position features;
[0028] The location features of the person are mapped to shared features to obtain the region of interest. The region of interest is then pooled and convolved to obtain the behavioral features of the person.
[0029] In practical applications, pre-trained feature extraction models can be used to extract object location features, person location features, and person behavior features. The feature extraction model includes a shared feature extraction module, an object estimation branch, a behavior estimation branch, and an output module. The shared feature extraction module constructs a shared encoder based on a lightweight backbone network. It can use feature extraction networks such as ResNet or MobileNet as the basic backbone network, removing the fully connected classification layer. This module performs global feature extraction on each frame of the video data, obtaining general visual features, i.e., shared features, which are shared by the location estimation and behavior estimation branches. After obtaining the shared features, they are input into the location estimation and behavior estimation branches respectively. The location estimation branch includes a multi-scale feature fusion layer, a location regression convolutional layer, a classification convolutional layer, a confidence convolutional layer, and a first output layer. The multi-scale feature fusion layer includes a 1×1 convolution to reduce the dimensionality of the shared features. The location regression convolutional layer includes a 3×3 convolution, batch normalization, and a ReLU activation function to perform regression prediction of the bounding box coordinates. The classification convolutional layer includes 3×3 convolutions, batch normalization, and ReLU activation functions. It outputs the class probability of each bounding box belonging to its respective category, enabling the discrimination and classification of objects within the bounding boxes, such as distinguishing between buildings and streets, or equipment and people. The confidence convolutional layer includes 3×3 convolutions and a Sigmoid activation function. It outputs the confidence score of whether an object exists within each bounding box. The output layer includes a concatenation layer, which concatenates the bounding box coordinates, class probabilities, and confidence scores to obtain location features. Based on the class probabilities, the location features are divided into object location features and person location features. The behavior estimation branch includes a ROI feature pooling layer, a behavior feature encoding layer, a behavior classification head module, and a second output layer. The ROI feature pooling layer extracts the corresponding region from the shared features based on the bounding box coordinates output by the person location features, obtaining the region of interest. Bilinear interpolation is performed on the ROI to calculate feature values, and the recalculated feature values are pooled and output to obtain fixed-size features. The behavior feature encoding layer includes 1×1 convolutions, 3×3 convolutions, and global pooling, used to reduce the dimensionality of fixed-size features, extract behavior space features, and perform global average pooling to obtain a fixed-length feature vector. The behavior classification head module includes a first fully connected layer, a second fully connected layer, a Dropout layer, and a Softmax activation function. The first fully connected layer is used to reduce the dimensionality and refine the features output by the ROI feature pooling layer. The second fully connected layer is used to map the features output by the first fully connected layer to a dimension equal to the number of behavior categories, obtaining the action amplitude and behavior prediction score. The Dropout layer is used to suppress overfitting during training by randomly deactivating neurons. The Softmax activation function is used to convert the behavior prediction score into a probability distribution between 0 and 1, obtaining a probabilistic output for multi-behavior classification, and the category with the highest action amplitude and probability is taken as the human behavior feature.
[0030] For training the feature extraction model, the training image is first input into the shared feature extraction module to obtain a shared feature map. Then, the shared features are input into the location estimation branch and the behavior estimation branch respectively, and the task losses for each branch are calculated separately. These losses are then weighted and summed to obtain the total loss. The parameters of the shared feature extraction module, the location estimation branch, and the behavior estimation branch are backpropagated. Training is iterated until the losses for each task converge, ultimately yielding the trained feature extraction model.
[0031] Specifically, the standardization and data correction of the object position features, person position features, and person behavior features yield a multi-dimensional information set, including:
[0032] The Z-score normalization algorithm was used to standardize the object position features, person position features, and person behavior features to obtain a standard dataset.
[0033] Acquire light intensity change data, adjust the correction factor based on the light intensity change data, and correct the standard data set based on the adjusted correction factor to obtain a multidimensional information set.
[0034] In practical applications, the Z-score normalization algorithm converts the bounding box coordinates of object position features, the bounding box coordinates of person position features, and the amplitude of person's behavioral features into a unified format with a mean of 0 and a standard deviation of 1. For example, sofa coordinates are normalized from (0.2, 0.3) to (-1.5, -1.2), and the amplitude of a person's limb movements, such as a 45-degree angle, is normalized to a normalized value of 0.8. The Z-score normalization algorithm involves calculating the mean and standard deviation of the entire dataset, and then subtracting the mean and dividing by the standard deviation for each item, achieving a unified vector representation of heterogeneous information such as object position, person position, and amplitude of movement. By acquiring light intensity change data through a light sensor, adjusting the correction factor based on the light intensity change data, and correcting the standard dataset according to the adjusted correction factor, the offset of bounding box coordinates and behavioral movements caused by light intensity changes can be corrected. Specifically, the correction factor can be calculated using a correction formula, which is:
[0035]
[0036] Where 'a' represents the correction factor, 'L' represents the light intensity variation data, and 'α' represents the sensitivity coefficient. The sensitivity coefficient is determined through calibration experiments. Visual coordinates of the same object are collected under different light intensities, the corresponding correction factor is calculated, and a scatter plot of the correction factor and light intensity variation data is plotted. The sensitivity coefficient is then fitted to this plot. The correction factor is multiplied by the bounding box coordinates and action amplitude in the standard dataset to obtain the corrected bounding box coordinates and action amplitude, thus yielding the multidimensional information set.
[0037] Step S3: Construct an environmental and behavioral feature map based on the multidimensional information set and synchronized emotional data.
[0038] Specifically, the step of constructing an environmental and behavioral feature map based on the multidimensional information set and synchronized sentiment data includes:
[0039] Calculate the density features of the people based on their location features in the multidimensional information set;
[0040] The synchronous sentiment data is classified using a support vector machine to determine the correlation between the synchronous sentiment data and the density features of the individuals.
[0041] Based on the aforementioned relationships, the object location features, person location features, and person behavior features in the multidimensional information set are overlaid to obtain an environmental and behavioral feature map.
[0042] In practical applications, a kernel density estimation method is used, with a bandwidth parameter set to 0.5 meters. The density distribution is calculated based on the bounding box coordinates of the person's position features to obtain the person's density features. A pre-determined sentiment category, such as positive, neutral, and negative, is used to train a support vector machine (SVM) to learn the sentiment classification. A speech set is used for training, and the speech set is divided into training and testing sets using an 8:2 ratio. The SVM extracts MFCC features, first-order difference features, and second-order difference features of the speech from the speech set to obtain high-dimensional feature vectors. These high-dimensional feature vectors are standardized. An RBF kernel-based SVM is used to achieve sentiment classification by maximizing the margin. The penalty coefficient and function coefficients of the SVM are optimized using a grid search. After training, a trained SVM is obtained, achieving the classification of positive, neutral, and negative sentiment categories. Synchronous sentiment data is input into the SVM to obtain the corresponding sentiment category. The synchronized sentiment data and the density features of individuals are already aligned temporally. The next step is to align them spatially. This is achieved by acquiring the location coordinates of the synchronized sentiment data using microphones. Based on these coordinates, the corresponding sentiment categories are matched with the corresponding density regions in the density features. The probability distribution of sentiment categories across different density ranges (low, medium, and high) in the density features is then calculated. For example, if the proportion of negative sentiment categories increases to 40% in the high-density region, it indicates increased stress among the population, thus establishing a correlation between the synchronized sentiment data and the density features. Finally, the correlation between the synchronized sentiment data and the density features is quantified using the Pearson correlation coefficient. The weights of object location features and human behavior features are adjusted based on the correlation between the synchronous emotional data and the human density features. If the correlation between the synchronous emotional data and the human density features is high, the weights of human behavior features and human location features are increased, while the weights of object location features are decreased. If the correlation between the synchronous emotional data and the human density features is low, the weights of human behavior features and human location features are decreased, while the weights of object location features are increased. Based on the adjusted weights, the object location features, human location features, and human behavior features are weighted and superimposed to obtain an environmental and behavioral feature map.
[0043] By classifying synchronous sentiment data using support vector machines, the sentiment categories of the synchronous sentiment data are obtained. The sentiment categories of the synchronous sentiment data are then matched with the density features of the people to obtain the correlation between the synchronous sentiment data and the density features of the people. Based on the correlation between the synchronous sentiment data and the density features of the people, the fusion ratio of the three types of features—object position, person position, and person behavior—can be dynamically and adaptively adjusted. This achieves deep coupling of multi-dimensional information such as group sentiment, crowd distribution, object environment, and person behavior. Furthermore, the feature superposition logic is optimized through correlation constraints to reduce feature redundancy and subjective weight errors, thereby strengthening the risk representation capability of the feature map. This can provide reliable data support for subsequent density clustering to accurately mine potential conflict areas.
[0044] Step S4: Identify potential conflict areas based on the environmental and behavioral feature map, analyze abnormal events in the potential conflict areas, and obtain abnormal statistical results.
[0045] Specifically, the step of identifying potential conflict areas based on the environmental and behavioral feature map, analyzing abnormal events in the potential conflict areas, and obtaining abnormal statistical results includes:
[0046] The environmental and behavioral feature maps were analyzed using a density clustering algorithm to identify potential conflict regions.
[0047] Obtain the individual trajectory of the person, which is a coordinate sequence with a timestamp, and use the individual trajectory in the potential conflict area as the target trajectory;
[0048] Extract the path features of the target trajectory, perform time series analysis on the path features to obtain feature trends, take the feature trends that exceed the preset behavior deviation threshold as target feature trends, take the target trajectory corresponding to the target feature trends as abnormal trajectories, count the abnormal proportion of abnormal trajectories in the potential conflict area, and mark the potential conflict area as having an abnormal event when the abnormal proportion in the potential conflict area exceeds the proportion threshold.
[0049] A sliding window method was used to statistically analyze the frequency, duration, and affected area of abnormal events within potential conflict zones, yielding anomaly statistics.
[0050] In practical applications, the radius parameter of the density clustering algorithm can be set to 5 meters and the minimum number of points to 10. These parameters can be adjusted based on the spatial scale and the density of people. By analyzing the environmental and behavioral feature maps using the density clustering algorithm with the set radius and minimum number of points, multiple clusters and their densities are obtained. Regions corresponding to clusters with densities exceeding a density threshold are identified as potential conflict areas. The density of all clusters is statistically analyzed, and the upper quantile (e.g., 85%) is selected as the base density threshold. This threshold is then adjusted based on the correlation between synchronized sentiment data and person density, as well as the relationship between person behavioral characteristics and object location characteristics, ultimately determining the final density threshold.
[0051] After obtaining the location features of the characters, their individual trajectories can be derived based on these features. Each character typically has a unique identifier. The bounding box coordinates of each character are sorted according to their timestamps to obtain their individual trajectories, i.e., a coordinate sequence with timestamps, such as Ti={(x1,y1,t1),(x2,y2,t2),…,(x k ,y k ,t kLet Ti represent the individual trajectory of the i-th person. Individual trajectories within the potential conflict area are extracted as target trajectories. Path features are calculated for each target trajectory using a fixed time window, which can be set to 0.5 seconds. Path features include: instantaneous velocity, instantaneous acceleration, motion direction angle, rate of change of direction, jerkiness, and dwell time. The path features are sorted chronologically to obtain a time series of path features. A linear model is fitted using the least squares method to obtain the characteristic trend of the path features. The characteristic trend includes trend direction (e.g., upward, downward, stable) and trend value. The trend value is compared with a behavior deviation threshold. If the trend value exceeds the behavior deviation threshold, the target trajectory corresponding to that trend value is considered an abnormal trajectory. To determine the behavioral deviation threshold, a large number of normal trend values of trajectory features under normal scenarios can be collected in advance. Normal scenarios refer to scenarios without conflict, with low density, and stable emotions. An empirical distribution is calculated for the trend values of each path feature. The lower quantile and upper quantile of the distribution are taken as the lower bound threshold and upper bound threshold, respectively. For example, the lower quantile can be 2%, and the upper quantile can be 98%. The interval between the lower and upper bound thresholds is used as the behavioral deviation threshold. To determine the proportion threshold, the proportion of abnormal trajectories in potential conflict areas during a large number of time periods without abnormal events can be collected in advance. The distribution of the abnormal trajectory proportion is calculated, and the upper quantile of this distribution is taken as the proportion threshold. After determining the proportion threshold, the proportion of abnormal trajectories in potential conflict areas is statistically analyzed to obtain the abnormal proportion. The abnormal proportion is compared with the proportion threshold. If the abnormal proportion exceeds the proportion threshold, an abnormal event is marked as occurring in the potential conflict area. Based on a time sliding window, the frequency, duration, and affected area of abnormal events in potential conflict areas are statistically analyzed. For example, the frequency, duration, and affected area of abnormal events in potential conflict areas within a 10-minute period are calculated to obtain the abnormal statistical results.
[0052] Step S5: Determine the potential hazard type of the abnormal event based on the abnormal statistical results and historical accidents, obtain the risk index based on the potential hazard type and risk level standard of the abnormal event, and generate an early warning notification if the risk index of the abnormal event exceeds the risk threshold.
[0053] Specifically, the step of determining the potential hazard type of the abnormal event based on the abnormal statistical results and historical accidents includes:
[0054] A historical accident database is pre-established. The historical accidents in the database include: abnormal event types and hazard types. In the historical accident database, the occurrence frequency of each hazard type is counted for each abnormal event type.
[0055] The similarity between each abnormal event is calculated based on the occurrence frequency, duration, and affected area. The abnormal events are then classified according to the similarity to obtain the abnormal event type.
[0056] Based on the abnormal event type of each abnormal event, the hazard type that occurs most frequently in historical accidents under each abnormal event type is taken as the potential hazard type of each abnormal event.
[0057] In practical applications, hierarchical clustering algorithms are used to classify abnormal events based on their frequency, duration, and affected area. Specifically, Euclidean distance is used to calculate the similarity between abnormal events. Initially, the two closest abnormal events are merged into one cluster. The average distance method is used to update the distance between clusters. This process is repeated until all abnormal events are clustered into three main clusters: a high-frequency short-duration cluster, a low-frequency long-duration cluster, and a medium-frequency wide-area cluster. This achieves the classification of abnormal events and determines their types. In the historical accident database, hazard types include: equipment failure, fire, leakage, collapse, poisoning, etc. For example, in historical accidents, the number of times equipment failure, fire, leakage, collapse, and poisoning occurred under the abnormal events of the high-frequency short-duration cluster can be statistically analyzed by region. Assuming the detected abnormal event is of the type of low-frequency long-duration cluster, then based on the region where the abnormal event occurred, the most frequent hazard type under the abnormal event of low-frequency long-duration cluster in the corresponding region is counted. If the region where the detected abnormal event occurred is region A, then the most frequent hazard type under the abnormal event of low-frequency long-duration cluster in region A in the historical accidents is counted. If the hazard type is leakage, then the potential hazard type of the detected abnormal event is leakage.
[0058] Specifically, the risk index obtained based on the potential hazard type and risk level criteria of the abnormal event includes:
[0059] Determine the severity index based on the type of potential hazard;
[0060] Based on the frequency, duration, and affected area of the aforementioned abnormal events, the risk level standard is queried to obtain the risk level;
[0061] The risk index is calculated based on the severity index and risk level.
[0062] In practical applications, a severity index is pre-defined by statistically analyzing the average consequences of various hazard types from historical accidents. For example, the average consequences of each hazard type are calculated based on factors such as casualties, property damage, and environmental impact. After obtaining the average consequences of all historical accidents, a five-level severity index is assigned. For instance, equipment failure corresponds to a severity index of level 1, leakage to level 2, collapse to level 3, poisoning to level 4, and fire to level 5. Risk level standards are set based on the frequency, duration, and affected area of abnormal events. For example, the risk level for high-frequency, short-duration clusters of abnormal events is set to level 1, for low-frequency, long-duration clusters to level 2, and for medium-frequency, wide-area clusters to level 3. After obtaining the severity index and risk level, the severity index and risk level are multiplied together to obtain the risk index.
[0063] Specifically, generating an early warning notification when the risk index of the abnormal event exceeds a risk threshold includes:
[0064] Anomalies whose risk index exceeds the risk threshold are identified as target anomalies, and the target locations and target intensities of multiple target sensors corresponding to the detected target anomalies are extracted.
[0065] A triangulation algorithm is used to calculate the locations and intensities of multiple targets to obtain the coordinates of the hazard source, and high-risk areas are determined based on the coordinates of the hazard source.
[0066] Extract the population gathering points in the high-risk area, calculate the initial paths from the population gathering points to all refuge points, eliminate the initial paths with congestion coefficients exceeding the congestion threshold to obtain candidate paths, and determine the evacuation path for each candidate path by weighted calculation based on travel time and safety factor.
[0067] Early warning notifications are generated based on the coordinates of the hazard source, high-risk areas, and evacuation routes.
[0068] In practical applications, the risk index of each historical accident is calculated, and historical accidents are marked as "caused actual damage" and "did not cause damage." An ROC curve is plotted, and the risk index corresponding to the point with the largest Youden index is selected as the risk threshold. If the risk threshold is 3, and the anomaly type of anomaly event a is a low-frequency long-duration cluster, the hazard type is leakage, the risk level corresponding to low-frequency long-duration clusters is level 2, and the severity index corresponding to leakage is level 2, then the risk index is 4. Therefore, the risk index of anomaly event a exceeds the risk threshold, and anomaly event a is designated as the target anomaly event. Multiple sensors that detect abnormal event 'a' are used as target sensors. The target position and target intensity of the target sensors are extracted. There are at least three target sensors, and the three target sensors are not collinear. The target position of the target sensor is used as a mass point, and the target intensity is used as the target weight. The initial estimated coordinates are obtained by weighted averaging the target position and the corresponding target weight. The current residual vector and loss function value are calculated based on the initial estimated coordinates. Then, the Jacobian matrix is calculated, and a corrected normal equation is constructed. The increment is calculated based on the corrected normal variance. The parameters are updated using the increment. The loss is calculated using the updated parameters. The loss before and after the update is compared. The iteration is repeated until the convergence condition is met to obtain the coordinates of the hazard source. For example, if the location result of the hazard source coordinates is 116.3 degrees east longitude and 39.9 degrees north latitude, the area within a preset radius centered on the hazard source coordinates is determined as a high-risk area. The preset radius can be set to 1 km, 1.5 km, or 2 km, which can be determined according to the actual population density and terrain data. Based on the population density characteristics of high-risk areas, crowd gathering points are extracted. Nodes with population density greater than a density threshold are designated as crowd gathering points. Initial paths from these gathering points to various evacuation points can be generated using Dijkstra's algorithm or A*. Then, candidate paths are selected by combining traffic data. Specifically, a congestion threshold can be set to 0.8. Initial paths with congestion coefficients exceeding the congestion threshold are first removed from all initial paths to obtain candidate paths. Then, the travel time and safety factor of each candidate path are calculated. The travel time is determined based on real-time speed and the distance of the candidate path, while the safety factor is determined based on road type and historical traffic accident rate. The travel time and safety factor are then normalized. Finally, a weighted calculation is performed on the normalized travel time and safety factor. The weighted calculation formula is as follows:
[0069] ,
[0070] Where P represents the weighted score, T represents the normalized travel time, S represents the normalized safety factor, w1 represents the weight of travel time, and w2 represents the weight of the safety factor. The weights of travel time and safety factor can be determined based on the risk index of the abnormal event. If the risk index is high, the weight of travel time is increased and the weight of safety factor is decreased; if the risk index is low, the weight of travel time is decreased and the weight of safety factor is increased. After calculating the weighted scores of each candidate path, the candidate path with the lowest weighted score is selected as the evacuation path. Then, a warning notification is automatically pushed through a designated communication channel, such as a 5G network interface. The warning notification is broadcast and includes the coordinates of the hazard source, the high-risk area, and the evacuation path.
[0071] Specifically, after generating the early warning notification, the following steps are included:
[0072] A continuous monitoring mechanism is used to iteratively analyze video data and emotional fluctuation data. After a new abnormal event is detected, the abnormal statistics are updated based on the new abnormal event, and the early warning notification is updated based on the updated abnormal statistics and historical incidents.
[0073] In practical applications, video data and emotional fluctuation data are acquired in real time to obtain updated video and emotional fluctuation data. The methods in steps S1 to S4 are used to iteratively analyze the updated video and emotional fluctuation data. Upon detecting a new abnormal event, the abnormal event statistics are updated accordingly. The method in step S5 is then used to update the early warning notification based on the updated abnormal statistics and historical incidents. Through continuous monitoring and iterative analysis, new abnormal events can be captured in real time, the abnormal statistics can be updated, the risk index can be dynamically reassessed by combining historical incidents, and the early warning notification can be automatically updated, thus achieving dynamic early warning.
[0074] This application presents a safety warning robot control method with intelligent voice broadcasting. It integrates video data and emotional fluctuation data synchronously using timestamps, then extracts object location features, person location features, and person behavior features from the video data to construct a standardized multi-dimensional information set. Based on this multi-dimensional information set and synchronized emotional data, it builds an environmental and behavioral feature map, accurately identifying potential conflict areas and abnormal events. By comparing with historical incidents, when the risk index of an abnormal event exceeds a risk threshold, it automatically generates a warning notification including the coordinates of the hazard source, high-risk areas, and evacuation routes. Through continuous monitoring and iterative analysis, it optimizes the state representation and dynamically adjusts the high-risk area division and evacuation route planning.
[0075] like Figure 2 As shown in the illustration, this application also provides a safety warning robot control system with intelligent voice broadcasting, including:
[0076] The synchronization module 10 is used to acquire video data and emotional fluctuation data, and to synchronize the video data and emotional fluctuation data with timestamps to obtain synchronized video data and synchronized emotional data.
[0077] The extraction and processing module 20 is used to extract object position features, person position features, and person behavior features from the synchronized video data, and to perform standardization processing and data correction on the object position features, person position features, and person behavior features to obtain a multi-dimensional information set.
[0078] The graph construction module 30 is used to construct an environmental and behavioral feature graph based on the multidimensional information set and synchronized sentiment data.
[0079] The identification and analysis module 40 is used to identify potential conflict areas based on the environmental and behavioral feature map, analyze abnormal events in the potential conflict areas, and obtain abnormal statistical results.
[0080] The judgment and early warning module 50 is used to judge the potential danger type of the abnormal event based on the abnormal statistical results and historical accidents, obtain a risk index based on the potential danger type and risk level standard of the abnormal event, and generate an early warning notification when the risk index of the abnormal event exceeds the risk threshold.
[0081] The various modules of the aforementioned safety warning robot control system with intelligent voice broadcasting can be implemented entirely or partially through software, hardware, or a combination thereof. These modules and units can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the corresponding operations of each module.
[0082] It should be noted that, in this document, the terms "comprising," "including," and any other variations are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Specific examples have been used in this document to illustrate the principles and implementation methods of the technical solutions of this application. The above examples are only for the purpose of helping to understand the methods and core ideas of this application. The above descriptions are merely preferred embodiments of this application. It should be pointed out that, due to the limitations of written expression and the objective existence of infinite specific structures, those skilled in the art can make several improvements, modifications, or changes without departing from the principles of this application, and can also combine the above technical features in an appropriate manner; these improvements, modifications, changes, or combinations, or the direct application of the concept and technical solutions of this application to other situations without modification, should all be considered within the scope of protection of this application.
Claims
1. A method for controlling a safety early warning robot with intelligent voice broadcasting, characterized in that, include: Acquire video data and emotional fluctuation data, and synchronize the video data and emotional fluctuation data with timestamps to obtain synchronized video data and synchronized emotional data; Object position features, person position features, and person behavior features are extracted from the synchronized video data. The object position features, person position features, and person behavior features are then standardized and corrected to obtain a multidimensional information set. Construct an environmental and behavioral feature map based on the multidimensional information set and synchronized sentiment data; Based on the environmental and behavioral feature maps, potential conflict areas are identified, and abnormal events in the potential conflict areas are analyzed to obtain abnormal statistical results. Based on the abnormal statistical results and historical accidents, the potential hazard type of the abnormal event is determined. A risk index is obtained based on the potential hazard type and risk level standard of the abnormal event. If the risk index of the abnormal event exceeds the risk threshold, an early warning notification is generated.
2. The safety early warning robot control method with intelligent voice broadcasting according to claim 1, characterized in that, The object position features, person position features, and person behavior features are extracted from the synchronized video data, including: A shared encoder is used to extract shared features from the synchronized video data; Convolution is performed on the shared features to obtain object position features and person position features; The location features of the person are mapped to shared features to obtain the region of interest. The region of interest is then pooled and convolved to obtain the behavioral features of the person.
3. The safety early warning robot control method with intelligent voice broadcasting according to claim 1, characterized in that, The standardization and data correction of the object location features, person location features, and person behavior features yield a multi-dimensional information set, including: The Z-score normalization algorithm was used to standardize the object position features, person position features, and person behavior features to obtain a standard dataset. Acquire light intensity change data, adjust the correction factor based on the light intensity change data, and correct the standard data set based on the adjusted correction factor to obtain a multidimensional information set.
4. The safety early warning robot control method with intelligent voice broadcasting according to claim 1, characterized in that, The construction of the environmental and behavioral feature map based on the multidimensional information set and synchronized sentiment data includes: Calculate the person density features based on the person location features in the multidimensional information set; The synchronous sentiment data is classified using a support vector machine to determine the correlation between the synchronous sentiment data and the density features of the individuals. Based on the aforementioned relationships, the location features of objects, locations of people, and behavioral features in the multidimensional information set are overlaid to obtain an environmental and behavioral feature map.
5. The safety early warning robot control method with intelligent voice broadcasting according to claim 1, characterized in that, The process of identifying potential conflict areas based on the environmental and behavioral feature map, analyzing abnormal events in the potential conflict areas, and obtaining abnormal statistical results includes: The environmental and behavioral feature maps were analyzed using a density clustering algorithm to identify potential conflict regions. Obtain the individual trajectory of the person, which is a coordinate sequence with a timestamp, and use the individual trajectory in the potential conflict area as the target trajectory; Extract the path features of the target trajectory, perform time series analysis on the path features to obtain feature trends, take the feature trends that exceed the preset behavior deviation threshold as target feature trends, take the target trajectory corresponding to the target feature trends as abnormal trajectories, count the abnormal proportion of abnormal trajectories in the potential conflict area, and mark the potential conflict area as having an abnormal event when the abnormal proportion in the potential conflict area exceeds the proportion threshold. A sliding window method was used to statistically analyze the frequency, duration, and affected area of abnormal events within potential conflict zones, yielding anomaly statistics.
6. The safety early warning robot control method with intelligent voice broadcasting according to claim 1, characterized in that, The anomaly statistics include the frequency of occurrence of the anomaly, the duration of the anomaly, and the area affected by the anomaly. The determination of the potential hazard type of the anomaly based on the anomaly statistics and historical incidents includes: A historical accident database is pre-established. The historical accidents in the database include: abnormal event types and hazard types. In the historical accident database, the occurrence frequency of each hazard type is counted for each abnormal event type. The similarity between each abnormal event is calculated based on the occurrence frequency, duration, and affected area. The abnormal events are then classified according to the similarity to obtain the abnormal event type. Based on the abnormal event type of each abnormal event, the hazard type that occurs most frequently in historical accidents under each abnormal event type is taken as the potential hazard type of each abnormal event.
7. The safety early warning robot control method with intelligent voice broadcasting according to any one of claims 1-6, characterized in that, The risk index, derived based on the potential hazard type and risk level criteria of the abnormal event, includes: Determine the severity index based on the type of potential hazard; Based on the frequency, duration, and affected area of the aforementioned abnormal events, the risk level standard is queried to obtain the risk level; The risk index is calculated based on the severity index and risk level.
8. The safety early warning robot control method with intelligent voice broadcasting according to any one of claims 1-6, characterized in that, When the risk index of the abnormal event exceeds the risk threshold, an early warning notification is generated, including: Anomalies whose risk index exceeds the risk threshold are identified as target anomalies, and the target locations and target intensities of multiple target sensors corresponding to the detected target anomalies are extracted. A triangulation algorithm is used to calculate the locations and intensities of multiple targets to obtain the coordinates of the hazard source, and high-risk areas are determined based on the coordinates of the hazard source. Extract the population gathering points in the high-risk area, calculate the initial paths from the population gathering points to all refuge points, eliminate the initial paths with congestion coefficients exceeding the congestion threshold to obtain candidate paths, and determine the evacuation path for each candidate path by weighted calculation based on travel time and safety factor. Early warning notifications are generated based on the coordinates of the hazard source, high-risk areas, and evacuation routes.
9. The safety early warning robot control method with intelligent voice broadcasting according to any one of claims 1-6, characterized in that, After generating the warning notification, the following is included: A continuous monitoring mechanism is used to iteratively analyze video data and emotional fluctuation data. After a new abnormal event is detected, the abnormal statistics are updated based on the new abnormal event, and the early warning notification is updated based on the updated abnormal statistics and historical incidents.
10. A safety early warning robot control system with intelligent voice broadcasting, characterized in that, include: The synchronization module is used to acquire video data and emotional fluctuation data, and to synchronize the video data and emotional fluctuation data with timestamps to obtain synchronized video data and synchronized emotional data. The extraction and processing module is used to extract object position features, person position features, and person behavior features from the synchronized video data, and to perform standardization processing and data correction on the object position features, person position features, and person behavior features to obtain a multi-dimensional information set. The graph construction module is used to construct an environmental and behavioral feature graph based on the multidimensional information set and synchronized sentiment data. The identification and analysis module is used to identify potential conflict areas based on the environmental and behavioral feature map, analyze abnormal events in the potential conflict areas, and obtain abnormal statistical results. The judgment and early warning module is used to determine the potential danger type of the abnormal event based on the abnormal statistical results and historical accidents, obtain a risk index based on the potential danger type and risk level standard of the abnormal event, and generate an early warning notification when the risk index of the abnormal event exceeds the risk threshold.