Methods, devices, computer equipment and software products for extracting personal risk characteristics

CN121210977BActive Publication Date: 2026-08-14ELECTRIC POWER RES INST CHINA SOUTHERN POWER GRID CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-15
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0003]在相关技术中,基于传统机器学习的特征提取法,对非结构化数据处理能力弱,在电力施工中,难以从大量的事故报告文本、现场作业图片和操作视频(如违规操作的监控录像)中提取有效风险特征,依赖于人工特征工程,若初始特征设计不合理,会严重影响提取效果,且泛化能力有限,在不同电压等级的电网作业、输变电线路施工等跨场景应用时,准确性显著下降,导致最终提取的人身风险特征不够准确

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121210977B_ABST
    Figure CN121210977B_ABST
Patent Text Reader

Abstract

This application relates to a method, apparatus, computer equipment, computer-readable storage medium, and computer program product for extracting personal risk features. The method includes: acquiring multi-source heterogeneous data on risks occurring in the power industry; extracting risk features based on a preset model to obtain four-level features from the multi-source heterogeneous data; determining the mapping relationship between a general risk ontology and a power industry ontology; adapting the four-level features to the power industry ontology based on the mapping relationship and performing feature enhancement processing to obtain processed four-level features; calculating mutual information values ​​and generating an association strength matrix of the processed four-level features; using the accuracy and coverage of the processed four-level features as reward functions to determine the importance ranking of the processed four-level features; generating a personal risk feature map; and calibrating it to obtain a calibrated personal risk feature map. This method can improve the accuracy of personal risk feature extraction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of feature extraction technology, and in particular to a method, apparatus, computer equipment, computer-readable storage medium, and computer program product for extracting personal risk features. Background Technology

[0002] In the field of power safety, accurately extracting personal risk characteristics is crucial for ensuring the safety of workers.

[0003] In related technologies, feature extraction methods based on traditional machine learning have weak processing capabilities for unstructured data. In power construction, it is difficult to extract effective risk features from a large number of accident report texts, on-site operation pictures and operation videos (such as surveillance videos of violations). It relies on manual feature engineering. If the initial feature design is unreasonable, it will seriously affect the extraction effect. Moreover, the generalization ability is limited. When applied across different scenarios such as power grid operations at different voltage levels and transmission and transformation line construction, the accuracy drops significantly, resulting in the final extracted personal risk features being inaccurate. Summary of the Invention

[0004] Therefore, it is necessary to provide a method, apparatus, computer equipment, computer-readable storage medium, and computer program product for personal risk feature extraction that can improve the accuracy of feature extraction, in response to the above-mentioned technical problems.

[0005] Firstly, this application provides a method for extracting personal risk characteristics, including:

[0006] Acquire multi-source heterogeneous data on risks in the power industry;

[0007] Based on a preset model, risk features are extracted from the multi-source heterogeneous data to obtain four-level features of the multi-source heterogeneous data.

[0008] The mapping relationship between the general risk ontology and the power industry ontology is determined. Based on the mapping relationship, the four-level features are adapted to the power industry, and the adapted four-level features are subjected to feature enhancement processing to obtain the processed four-level features.

[0009] Calculate the mutual information value of the processed four-level features, and generate the association strength matrix of the processed four-level features based on the mutual information value. Based on the association strength matrix, use the accuracy and coverage of the processed four-level features as reward functions to determine the importance ranking of the processed four-level features.

[0010] Based on the importance ranking, a personal risk feature map is generated, and the features in the feature map are calibrated with a preset rule base to obtain a calibrated personal risk feature map.

[0011] In one embodiment, the preset model includes a basic feature layer, an interaction feature layer, a latent feature layer, and an evolutionary feature layer; the four-level features include basic explicit features, interactive linkage features, latent latent features, and evolutionary trend features; the risk feature extraction based on the preset model to obtain the four-level features of the multi-source heterogeneous data includes:

[0012] The multi-source heterogeneous data is input into the basic feature layer to extract the basic explicit features of the multi-source heterogeneous data;

[0013] The basic explicit features are input into the interaction feature layer. Based on the graph neural network, an interaction model among the user, device and environment is constructed. The basic explicit features are processed through the interaction model to obtain the interaction linkage features of the multi-source heterogeneous data.

[0014] The interactive linkage features are input into the latent feature layer, and the interactive linkage features are compared with the interactive linkage features of the verification data in the power industry to obtain the latent potential features of the multi-source heterogeneous data.

[0015] The time-series data, basic explicit features, interactive linkage features, and latent features from the multi-source heterogeneous data are input into the evolutionary feature layer. Based on the time-series data, the basic explicit features, interactive linkage features, and latent features are fused to obtain the evolutionary trend features of the multi-source heterogeneous data.

[0016] In one embodiment, the evolutionary trend features include the risk change vector and time-varying sensitivity factor of the multi-source heterogeneous data.

[0017] In one embodiment, generating a personal risk feature map based on the importance ranking includes:

[0018] Determine the data source, extraction basis, and confidence value of the processed four-level features;

[0019] Based on the importance ranking, the causal probability among the processed four-level features is transformed into the directional direction between the level where the cause feature is located and the level where the effect feature is located.

[0020] Based on the data source, the extraction criteria, the confidence level, and the direction, a personal risk feature map is generated.

[0021] In one embodiment, the method further includes:

[0022] Based on the features in the calibrated personal risk feature map, the directional interference vector is calculated;

[0023] Based on the directional interference vector, Gaussian noise is injected into the basic feature layer of the preset model, and gradient masking attack is performed on the latent feature layer of the preset model.

[0024] Calculate the adversarial loss function of the preset model, and optimize the parameters in the preset model based on the adversarial loss function to obtain the optimized preset model.

[0025] In one embodiment, the multi-source heterogeneous data includes text data, image data, and sensor data; prior to acquiring the multi-source heterogeneous data indicating risks in the power industry, the process includes:

[0026] The text data is segmented and labeled to extract the risk subject, behavior, and object triples from the text data;

[0027] The image data is processed using a multimodal processing model to generate a spatial semantic description vector for the image data.

[0028] Extract the frequency domain energy features from the sensor data, and determine the time domain abrupt change points in the sensor data based on the frequency domain energy features.

[0029] Secondly, this application also provides a personal risk characteristic extraction device, comprising:

[0030] The acquisition module is used to acquire multi-source heterogeneous data on risks in the power industry.

[0031] The extraction module is used to extract risk features from the multi-source heterogeneous data based on a preset model, and obtain four-level features of the multi-source heterogeneous data.

[0032] The processing module is used to determine the mapping relationship between the general risk ontology and the power industry ontology. Based on the mapping relationship, the four-level features are adapted to the power industry, and the adapted four-level features are subjected to feature enhancement processing to obtain the processed four-level features.

[0033] The determination module is used to calculate the mutual information value of the processed four-level features, generate the association strength matrix of the processed four-level features based on the mutual information value, and determine the importance ranking of the processed four-level features based on the association strength matrix, using the accuracy and coverage of the processed four-level features as reward functions.

[0034] The calibration module is used to generate a personal risk feature map based on the importance ranking, and to calibrate the features in the feature map with a preset rule base to obtain a calibrated personal risk feature map.

[0035] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:

[0036] Acquire multi-source heterogeneous data on risks in the power industry;

[0037] Based on a preset model, risk features are extracted from the multi-source heterogeneous data to obtain four-level features of the multi-source heterogeneous data.

[0038] The mapping relationship between the general risk ontology and the power industry ontology is determined. Based on the mapping relationship, the four-level features are adapted to the power industry, and the adapted four-level features are subjected to feature enhancement processing to obtain the processed four-level features.

[0039] Calculate the mutual information value of the processed four-level features, and generate the association strength matrix of the processed four-level features based on the mutual information value. Based on the association strength matrix, use the accuracy and coverage of the processed four-level features as reward functions to determine the importance ranking of the processed four-level features.

[0040] Based on the importance ranking, a personal risk feature map is generated, and the features in the feature map are calibrated with a preset rule base to obtain a calibrated personal risk feature map.

[0041] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the following steps:

[0042] Acquire multi-source heterogeneous data on risks in the power industry;

[0043] Based on a preset model, risk features are extracted from the multi-source heterogeneous data to obtain four-level features of the multi-source heterogeneous data.

[0044] The mapping relationship between the general risk ontology and the power industry ontology is determined. Based on the mapping relationship, the four-level features are adapted to the power industry, and the adapted four-level features are subjected to feature enhancement processing to obtain the processed four-level features.

[0045] Calculate the mutual information value of the processed four-level features, and generate the association strength matrix of the processed four-level features based on the mutual information value. Based on the association strength matrix, use the accuracy and coverage of the processed four-level features as reward functions to determine the importance ranking of the processed four-level features.

[0046] Based on the importance ranking, a personal risk feature map is generated, and the features in the feature map are calibrated with a preset rule base to obtain a calibrated personal risk feature map.

[0047] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, performs the following steps:

[0048] Acquire multi-source heterogeneous data on risks in the power industry;

[0049] Based on a preset model, risk features are extracted from the multi-source heterogeneous data to obtain four-level features of the multi-source heterogeneous data.

[0050] The mapping relationship between the general risk ontology and the power industry ontology is determined. Based on the mapping relationship, the four-level features are adapted to the power industry, and the adapted four-level features are subjected to feature enhancement processing to obtain the processed four-level features.

[0051] Calculate the mutual information value of the processed four-level features, and generate the association strength matrix of the processed four-level features based on the mutual information value. Based on the association strength matrix, use the accuracy and coverage of the processed four-level features as reward functions to determine the importance ranking of the processed four-level features.

[0052] Based on the importance ranking, a personal risk feature map is generated, and the features in the feature map are calibrated with a preset rule base to obtain a calibrated personal risk feature map.

[0053] The aforementioned method, apparatus, computer equipment, computer-readable storage medium, and computer program product for extracting personal risk features first acquire multi-source heterogeneous data on risks occurring in the power industry; based on a preset model, risk features are extracted from the multi-source heterogeneous data to obtain four-level features; the mapping relationship between a general risk ontology and a power industry ontology is determined; based on the mapping relationship, the four-level features are adapted to the power industry, and feature enhancement processing is performed on the adapted four-level features to obtain processed four-level features; the mutual information value of the processed four-level features is calculated, and an association strength matrix of the processed four-level features is generated based on the mutual information value; based on the association strength matrix, the accuracy and coverage of the processed four-level features are used as reward functions to determine the importance ranking of the processed four-level features; based on the importance ranking, a personal risk feature map is generated, and the features in the feature map are calibrated with a preset rule base to obtain a calibrated personal risk feature map. Thus, by constructing a dynamic risk feature library that includes a basic feature layer, an interactive feature layer, a latent feature layer, and an evolutionary feature layer, cross-domain knowledge transfer and industry adaptation are carried out, dual-loop feature optimization is implemented, interpretable feature maps are generated, and feature robustness is enhanced, making the final extracted personal risk features more accurate. Attached Figure Description

[0054] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0055] Figure 1 This is an application environment diagram of the personal risk feature extraction method in one embodiment;

[0056] Figure 2 This is a flowchart illustrating a method for extracting personal risk characteristics in one embodiment;

[0057] Figure 3 This is a structural block diagram of a personal risk feature extraction device in one embodiment;

[0058] Figure 4 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0059] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0060] It should be noted that the terms "first," "second," etc., used in this application can be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from the second element. The terms "comprising" and "having," and any variations thereof, used in this application, are intended to cover non-exclusive inclusion. The term "multiple" used in this application refers to two or more. The term "and / or" used in this application refers to one of the embodiments, or any combination of multiple embodiments.

[0061] The personal risk feature extraction method provided in this application embodiment can be applied to, for example... Figure 1 In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104 or located on the cloud or other network servers. Terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, drones, low-altitude aircraft, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, projection devices, etc. Portable wearable devices can include smartwatches, smart bracelets, head-mounted devices, etc. Head-mounted devices can be virtual reality (VR) devices, augmented reality (AR) devices, smart glasses, etc. Server 104 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.

[0062] In one exemplary embodiment, such as Figure 2 As shown, a method for extracting personal risk characteristics is provided, and this method is applied to... Figure 1 Taking terminal 102 as an example, the explanation includes the following steps 202 to 210. Wherein:

[0063] Step 202: Obtain multi-source heterogeneous data on risks in the power industry.

[0064] Among them, multi-source heterogeneous data includes text data, image data, and sensor data.

[0065] For example, acquiring risk-related text data, image data, and sensor data in the power industry.

[0066] In one embodiment, real-time temperature parameters of equipment such as high-voltage switchgear and transformers can be collected by infrared temperature sensors; partial discharge sensors can monitor cable joints and the insulation status of GIS equipment; high-definition intelligent monitoring cameras (deployed at substations and transmission line operation sites to collect personnel operation videos and on-site environmental images, providing raw image data) can be used; drone inspection systems (equipped with visible light and thermal imaging cameras to acquire image data of transmission line towers and insulators) can be used; power work ticket management systems (providing text data containing operator information, operation steps, and risk prevention and control measures) can be used; and edge computing gateways (terminals) can be deployed locally at substations to achieve real-time aggregation and preprocessing of multi-source sensor data, reducing the latency of transmission to the cloud; and a power safety expert system (stores expert rules such as power industry accident cases and operating procedures, and interacts bidirectionally with the expert rule base of interpretable feature maps to support the calibration of feature extraction logic) can be used.

[0067] Step 204: Based on the preset model, risk features are extracted from the multi-source heterogeneous data to obtain four-level features of the multi-source heterogeneous data.

[0068] Optionally, a preset model can be used to analyze multi-source heterogeneous data, extract risk features, and obtain the basic explicit features, interactive linkage features, hidden potential features, and evolution trend features of multi-source heterogeneous data.

[0069] The preset model is a pre-trained large model, or it can be other models with feature extraction functions. This application does not limit this.

[0070] Step 206: Determine the mapping relationship between the general risk ontology and the power industry ontology. Based on the mapping relationship, adapt the four-level features to the power industry and perform feature enhancement processing on the adapted four-level features to obtain the processed four-level features.

[0071] For example, the mapping relationship between the general risk ontology and the power industry ontology is determined by the semantic understanding module in the preset model. Based on the mapping relationship, the four-level features are adapted to the power industry. The BERT-BiLSTM-CRF model can be used to achieve automatic alignment between terms. The adapted four-level features are then subjected to feature enhancement processing to obtain the processed four-level features.

[0072] Among them, the general risk ontology is a set of general risk concepts applicable to multiple industries, such as the general relationship of equipment failure caused by human error; the power industry ontology is a set of risk concepts specific to the power industry, such as the power industry concepts of "step voltage" and "arc burn", and the industry characteristic of electric shock caused by failure to ground during high-voltage maintenance.

[0073] In one embodiment, the feature risk contribution can be calibrated using an industry accident case library through a Bayesian optimizer, thereby further dynamically adjusting the weights.

[0074] In one embodiment, when a new accident case is added, the dynamic feature weights are adjusted using formula (1).

[0075]

[0076] in, The learning rate adjustment coefficients generated for the Bayesian optimizer. The adjusted weights, The weights before adjustment This is a loss function based on accident cases.

[0077] In another embodiment, industry-specific implicit relationships are mined using knowledge graph completion technology, triggering dynamic ontology expansion to enhance features.

[0078] Specifically, based on the existing industry risk knowledge graph, missing associations (such as the lack of association between new energy storage equipment and "thermal runaway") are identified. Knowledge graph completion technology (such as the TransE algorithm) is used to mine potential associations, triggering dynamic expansion of the ontology and adding new entities (new energy storage equipment) and new relationships (energy storage → thermal runaway → burns) to the ontology.

[0079] Step 208: Calculate the mutual information value of the processed four-level features, and generate the association strength matrix of the processed four-level features based on the mutual information value. Based on the association strength matrix, use the accuracy and coverage of the processed four-level features as reward functions to determine the importance ranking of the processed four-level features.

[0080] Optionally, the mutual information values ​​between different feature layers (such as the base layer and the interaction layer) and within the same layer are calculated through a multi-head self-attention mechanism. A feature association strength matrix (with elements representing the degree of association between features) is constructed based on the mutual information. Accuracy and coverage are used as reward functions. The feature extraction path is dynamically optimized through the DDPG reinforcement learning algorithm, and the updated list of feature importance ranking is output, i.e., the importance ranking.

[0081] Accuracy rate = (Number of correctly identified risk samples / Total number of risk samples) × 100%;

[0082] Coverage rate = (Number of extracted risk features / Total number of industry standard features) × 100%;

[0083] In one embodiment, the DDPG agent is initialized with the sum of accuracy and coverage as the reward function. The agent adjusts the extraction path according to the reward value, iteratively calculates the feature contribution, and generates an importance ranking based on the contribution.

[0084] Step 210: Based on importance ranking, generate a personal risk feature map, and calibrate the features in the feature map with a preset rule base to obtain a calibrated personal risk feature map.

[0085] For example, the data source, extraction basis, and confidence value of the processed four-level features are determined; based on the order of importance ranking, the causal probability between the processed four-level features is transformed into the directional direction between the level where the cause feature is located and the level where the effect feature is located; based on the data source, extraction basis, confidence value, and directional direction, a personal risk feature map is generated, and the personal risk feature map is calibrated with a preset rule base to obtain a calibrated personal risk feature map.

[0086] In one embodiment, the preset rule base can be an expert rule base. The reasoning chain in the personal risk feature map (such as "electric shock caused by not wearing insulating clothing") is compared with the reasoning chain in the expert rule base. If there is a deviation, the reason is analyzed through meta-learning (such as unreasonable feature weights), the extraction logic is dynamically adjusted (such as increasing the weight of "wet ground"), and the personal risk feature map is updated.

[0087] In one embodiment, after the personal risk feature map is generated, the results are synchronized to the power safety monitoring platform. The platform duty personnel can view the dynamic risk indicators through the display screen. If the risk value of a certain work area exceeds the preset threshold, the platform’s built-in instruction sending module sends a warning message to the smart safety helmet terminal of the on-site workers, and at the same time, the on-site sound and light alarm device is activated to issue a warning.

[0088] In the aforementioned method for extracting personal risk features, multi-source heterogeneous data on risks occurring in the power industry are acquired. Based on a preset model, risk features are extracted from the multi-source heterogeneous data to obtain four-level features. The mapping relationship between a general risk ontology and a power industry ontology is determined. Based on the mapping relationship, the four-level features are adapted to the power industry, and feature enhancement processing is performed on the adapted four-level features to obtain processed four-level features. The mutual information value of the processed four-level features is calculated, and an association strength matrix of the processed four-level features is generated based on the mutual information value. Based on the association strength matrix, the accuracy and coverage of the processed four-level features are used as reward functions to determine the importance ranking of the processed four-level features. Based on the importance ranking, a personal risk feature map is generated, and the features in the feature map are calibrated with a preset rule base to obtain a calibrated personal risk feature map. Thus, by constructing a dynamic risk feature library that includes a basic feature layer, an interactive feature layer, a latent feature layer, and an evolutionary feature layer, cross-domain knowledge transfer and industry adaptation are carried out, dual-loop feature optimization is implemented, interpretable feature maps are generated, and feature robustness is enhanced, making the final extracted personal risk features more accurate.

[0089] In an exemplary embodiment, the preset model includes a basic feature layer, an interaction feature layer, a latent feature layer, and an evolutionary feature layer; the four-level features include basic explicit features, interactive linkage features, latent latent features, and evolutionary trend features; based on the preset model, risk features are extracted from multi-source heterogeneous data to obtain the four-level features of the multi-source heterogeneous data, including: inputting the multi-source heterogeneous data into the basic feature layer to extract the basic explicit features of the multi-source heterogeneous data; inputting the basic explicit features into the interaction feature layer, and constructing an interaction model between users, devices, and the environment based on a graph neural network. The model processes the basic explicit features through an interaction model to obtain the interaction and linkage features of multi-source heterogeneous data. These interaction and linkage features are then input into the latent feature layer, and compared with the interaction and linkage features of verification data in the power industry to obtain the latent potential features of the multi-source heterogeneous data. Finally, the time-series data, basic explicit features, interaction and linkage features, and latent potential features from the multi-source heterogeneous data are input into the evolutionary feature layer. Based on the time-series data, the basic explicit features, interaction and linkage features, and latent potential features are fused to obtain the evolutionary trend features of the multi-source heterogeneous data.

[0090] In practice, the basic feature layer uses a rule engine-based primary feature extraction to process multi-source heterogeneous data and output basic explicit features such as device status and environmental indicators.

[0091] Among them, equipment status parameters are the core operating parameters of the equipment, such as the dielectric loss value of the transformer's insulating oil, the partial discharge of the cable, the real-time temperature of the high-voltage switchgear, and the contact resistance of the disconnector switch; environmental indicators are the environmental parameters of the power industry work site and the surrounding environment of the equipment, such as the temperature, humidity, and wind speed of the substation, and the data on thunderstorms and high temperatures of the transmission line; basic explicit characteristics are the characteristics that can be directly observed and quantified, namely the above-mentioned equipment status parameters and environmental indicators, as well as the certification information of the operators and the equipment model.

[0092] The interaction feature layer constructs an interaction model between users, devices, and the environment through a graph neural network (GNN). It processes basic explicit features with users, devices, interaction models (environment), and interaction relationships (such as "personnel operating devices" and "environment affecting devices") as edges, and outputs interactive linkage features by learning the association patterns between nodes through GNN.

[0093] Among them, the interactive linkage feature includes a behavior linkage matrix, where the matrix elements represent the linkage strength between different subjects (such as the correlation value of "personnel misoperation causing equipment temperature to rise").

[0094] The latent feature layer employs a contrastive learning strategy, extracting latent features through training by comparing positive and negative samples.

[0095] Among them, both positive and negative samples come from interactive linkage features. Positive samples contain risk features (such as monitoring features of "operation without wearing insulating gloves" and sensor features of "partial discharge exceeding the standard"); negative samples contain no risk and compliance features (such as video features of "standard operation" and equipment data of "normal parameters").

[0096] In one embodiment, positive and negative sample pairs are trained by comparative learning, and a weighted loss function is constructed to enable the model to learn the similarity differences between positive and negative samples, thereby mining latent features such as "abnormal mental state of personnel" and "potential aging of equipment" from interactive linkage features.

[0097] In one embodiment, the formula for constructing the weighted loss function is shown in formula (2).

[0098]

[0099] in, The weighting coefficients for industry risk level labels are sim(・), and the similarity function is f. i For the current sample; f j + For positive samples; f k - For negative samples; τ is the temperature coefficient.

[0100] The evolutionary feature layer uses an improved LSTM-Transformer hybrid architecture to jointly model time-series data, basic explicit features, interactive features, and latent features, and outputs evolutionary trend features.

[0101] Specifically, the time-series data can be data related to time changes, such as equipment temperature sequences continuously collected by sensors, personnel operation time sequences, and environmental parameter change data. This application embodiment does not limit this.

[0102] In one embodiment, a gated temporal convolutional network (GTCN) is used instead of a standard LSTM unit to fuse the temporal correlations of the features from the first three layers to obtain the final evolutionary trend features.

[0103] In the above embodiments, the four-layer feature extraction structure uses a progressive mining approach of "basic-interaction-implicit-evolution" to transform the originally scattered and fragmented risk information into a complete evidence chain of "data → behavior → potential hidden dangers → future trends", which significantly improves the comprehensiveness, foresight and interpretability of power personal risk identification.

[0104] In one exemplary embodiment, the evolutionary trend features include risk change vectors and time-varying sensitivity factors from multi-source heterogeneous data.

[0105] In practice, evolutionary trend characteristics include risk change vectors and time-varying sensitive factors from multi-source heterogeneous data.

[0106] Among them, the risk change vector can be "the trend of partial discharge increasing by 10% in 3 hours", and the time-varying sensitive factor can be "ambient temperature" with a higher influence weight during high-temperature periods.

[0107] In the above embodiments, the "rate of change" and "who is most sensitive" are quantified, making the evolution trend of personal risk characteristics more intuitive.

[0108] In an exemplary embodiment, a personal risk feature map is generated based on importance ranking, including: determining the data source, extraction basis, and confidence value of the processed four-level features; converting the causal probability between the processed four-level features into a directional orientation between the level where the cause feature is located and the level where the effect feature is located, based on the order of importance ranking; and generating a personal risk feature map based on the data source, extraction basis, confidence value, and directional orientation.

[0109] In practice, knowledge distillation technology is used to transform the hidden logic into a visual rule chain, and the data source, extraction basis, and confidence value of the processed four-level features are determined. Based on the order of importance ranking, the causal probability between the processed four-level features is transformed into the directional direction between the level where the cause feature is located and the level where the effect feature is located. Based on the data source, extraction basis, confidence value, and directional direction, a personal risk feature map is generated.

[0110] In one embodiment, feature data and extraction criteria are labeled in the personal risk feature map, and the causal probability between features is quantified by a directed graph to generate a risk inference chain containing confidence values.

[0111] In the above embodiments, by using a three-level node system of "data source → correlation probability → inference chain" to break down each risk assessment into a retrievable chain of evidence, the tracing process becomes simpler and faster when a risk occurs.

[0112] In an exemplary embodiment, the method further includes: calculating a directional interference vector based on features in the calibrated personal risk feature map; injecting Gaussian noise into the basic feature layer of the preset model based on the directional interference vector, and performing gradient masking attacks on the latent feature layer of the preset model; calculating the adversarial loss function of the preset model, and optimizing the parameters in the preset model based on the adversarial loss function to obtain an optimized preset model.

[0113] In practice, based on the features in the calibrated personal risk feature map, a directional interference vector is calculated using the FGSM algorithm. Based on the directional interference vector, Gaussian noise is injected into the basic feature layer of the preset model, and gradient masking attack is performed on the latent feature layer of the preset model. The adversarial loss function of the preset model is calculated, and the parameters in the preset model are optimized based on the adversarial loss function to obtain the optimized preset model.

[0114] The formula for calculating the adversarial loss function is shown in formula (3).

[0115]

[0116] in, Let cross-entropy be the loss function. For balance coefficient, This represents the gradient of the loss function with respect to the input features.

[0117] In the above embodiments, the problem of high noise sensitivity is solved by adversarial training and hierarchical defense, and high accuracy can still be maintained under data noise interference.

[0118] In an exemplary embodiment, the multi-source heterogeneous data includes text data, image data, and sensor data. Before acquiring multi-source heterogeneous data on risks in the power industry, the process includes: performing word segmentation and annotation on the text data to extract risk subject, behavior, and object triples from the text data; processing the image data through a multimodal processing model to generate a spatial semantic description vector of the image data; extracting frequency domain energy features from the sensor data, and determining temporal abrupt change points in the sensor data based on the frequency domain energy features.

[0119] In practice, risk subject, behavior and object triples are extracted based on dynamic word segmentation and semantic role labeling (SRL) of industry dictionary; image data is processed by ViT-BERT model (multimodal processing model) to obtain spatial semantic description vector; wavelet packet transform is used to extract frequency domain energy features in sensor data, and a one-dimensional convolutional autoencoder is combined to capture temporal abrupt change points.

[0120] In the above embodiments, by preprocessing multi-source heterogeneous data in advance, the subsequent processing flow is faster, and the efficiency of personal risk feature extraction is improved.

[0121] To illustrate the personal risk feature extraction method in this application in detail, an embodiment is described below. For example, this application describes a personal risk feature extraction method in a specific scenario.

[0122] First, acquire textual, image, and sensor data related to risks in the power industry.

[0123] By using a pre-set model to analyze multi-source heterogeneous data, risk features are extracted, resulting in the basic explicit features, interactive linkage features, hidden potential features, and evolutionary trend features of the multi-source heterogeneous data.

[0124] By using the semantic understanding module in the preset model, the mapping relationship between the general risk ontology and the power industry ontology is determined. Based on the mapping relationship, the four-level features are adapted to the power industry. The BERT-BiLSTM-CRF model can be used to achieve automatic alignment between terms. The adapted four-level features are then subjected to feature enhancement processing to obtain the processed four-level features.

[0125] The mutual information values ​​between different feature layers (such as the base layer and the interaction layer) and within the same layer are calculated through a multi-head self-attention mechanism. Based on the mutual information, a feature association strength matrix is ​​constructed (the elements represent the degree of association between features). Accuracy and coverage are used as reward functions. The feature extraction path is dynamically optimized through the DDPG reinforcement learning algorithm, and the updated list of feature importance ranking is output, which is the importance ranking.

[0126] The data sources, extraction criteria, and confidence values ​​of the processed four-level features are determined. Based on the order of importance ranking, the causal probabilities among the processed four-level features are transformed into directional directions between the levels of the cause features and the effect features. Based on the data sources, extraction criteria, confidence values, and directional directions, a personal risk feature map is generated. The personal risk feature map is then calibrated against a preset rule base to obtain a calibrated personal risk feature map.

[0127] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps. It is understood that the steps in different embodiments can be freely combined as needed, and all non-contradictory solutions formed by such combinations are within the scope of protection of this application.

[0128] Based on the same inventive concept, this application also provides a personal risk feature extraction device for implementing the aforementioned personal risk feature extraction method. The solution provided by this device is similar to the solution described in the above method; therefore, the specific limitations in one or more embodiments of the personal risk feature extraction device provided below can be found in the limitations of the personal risk feature extraction method described above, and will not be repeated here.

[0129] In one exemplary embodiment, such as Figure 3 As shown, a personal risk feature extraction device is provided, comprising: an acquisition module 301, an extraction module 302, a processing module 303, a determination module 304, and a calibration module 305, wherein:

[0130] The acquisition module is used to acquire multi-source heterogeneous data that may indicate risks in the power industry.

[0131] The extraction module is used to extract risk features from the multi-source heterogeneous data based on a preset model, thereby obtaining four-level features of the multi-source heterogeneous data.

[0132] The processing module is used to determine the mapping relationship between the general risk ontology and the power industry ontology. Based on the mapping relationship, the four-level features are adapted to the power industry, and the adapted four-level features are subjected to feature enhancement processing to obtain the processed four-level features.

[0133] The determination module is used to calculate the mutual information value of the processed four-level features, generate the association strength matrix of the processed four-level features based on the mutual information value, and determine the importance ranking of the processed four-level features based on the association strength matrix, using the accuracy and coverage of the processed four-level features as reward functions.

[0134] The calibration module is used to generate a personal risk feature map based on the importance ranking, and to calibrate the features in the feature map with a preset rule base to obtain a calibrated personal risk feature map.

[0135] In one exemplary embodiment, the extraction module is further configured to input the multi-source heterogeneous data into the basic feature layer to extract the basic explicit features of the multi-source heterogeneous data;

[0136] The basic explicit features are input into the interaction feature layer. Based on the graph neural network, an interaction model among the user, device and environment is constructed. The basic explicit features are processed through the interaction model to obtain the interaction linkage features of the multi-source heterogeneous data.

[0137] The interactive linkage features are input into the latent feature layer, and the interactive linkage features are compared with the interactive linkage features of the verification data in the power industry to obtain the latent potential features of the multi-source heterogeneous data.

[0138] The time-series data, basic explicit features, interactive linkage features, and latent features from the multi-source heterogeneous data are input into the evolutionary feature layer. Based on the time-series data, the basic explicit features, interactive linkage features, and latent features are fused to obtain the evolutionary trend features of the multi-source heterogeneous data.

[0139] In one exemplary embodiment, the evolutionary trend features include the risk change vector and time-varying sensitivity factor of the multi-source heterogeneous data.

[0140] In one exemplary embodiment, the above-described apparatus further includes a generation module for determining the data source, extraction basis, and confidence value of the processed four-level features;

[0141] Based on the importance ranking, the causal probability among the processed four-level features is transformed into the directional direction between the level where the cause feature is located and the level where the effect feature is located.

[0142] Based on the data source, the extraction criteria, the confidence level, and the direction, a personal risk feature map is generated.

[0143] In one exemplary embodiment, the above-described apparatus further includes an optimization module for calculating a directional interference vector based on features in the calibrated personal risk feature map;

[0144] Based on the directional interference vector, Gaussian noise is injected into the basic feature layer of the preset model, and gradient masking attack is performed on the latent feature layer of the preset model.

[0145] Calculate the adversarial loss function of the preset model, and optimize the parameters in the preset model based on the adversarial loss function to obtain the optimized preset model.

[0146] In one exemplary embodiment, the above-mentioned processing module is further configured to perform word segmentation and annotation processing on the text data, and extract the risk subject, behavior and object triplet from the text data;

[0147] The image data is processed using a multimodal processing model to generate a spatial semantic description vector for the image data.

[0148] Extract the frequency domain energy features from the sensor data, and determine the time domain abrupt change points in the sensor data based on the frequency domain energy features.

[0149] Each module in the aforementioned personal risk feature extraction device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.

[0150] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4As shown, the computer device includes a processor, memory, input / output interfaces (I / O), a communication interface, a display unit, and input devices. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface, display unit, and input devices are also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores risk characteristic data. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communication with external terminals via a network connection. When executed by the processor, the computer program implements a method for extracting personal risk characteristics.

[0151] The display unit of this computer device is used to form a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of this computer device can be a touch layer covering the display screen, or buttons, a trackball, or a touchpad set on the casing of the computer device, or an external keyboard, touchpad, or mouse, etc.

[0152] Those skilled in the art will understand that Figure 4 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0153] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:

[0154] Acquire multi-source heterogeneous data on risks in the power industry;

[0155] Based on a preset model, risk features are extracted from the multi-source heterogeneous data to obtain four-level features of the multi-source heterogeneous data.

[0156] The mapping relationship between the general risk ontology and the power industry ontology is determined. Based on the mapping relationship, the four-level features are adapted to the power industry, and the adapted four-level features are subjected to feature enhancement processing to obtain the processed four-level features.

[0157] Calculate the mutual information value of the processed four-level features, and generate the association strength matrix of the processed four-level features based on the mutual information value. Based on the association strength matrix, use the accuracy and coverage of the processed four-level features as reward functions to determine the importance ranking of the processed four-level features.

[0158] Based on the importance ranking, a personal risk feature map is generated, and the features in the feature map are calibrated with a preset rule base to obtain a calibrated personal risk feature map.

[0159] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:

[0160] Acquire multi-source heterogeneous data on risks in the power industry;

[0161] Based on a preset model, risk features are extracted from the multi-source heterogeneous data to obtain four-level features of the multi-source heterogeneous data.

[0162] The mapping relationship between the general risk ontology and the power industry ontology is determined. Based on the mapping relationship, the four-level features are adapted to the power industry, and the adapted four-level features are subjected to feature enhancement processing to obtain the processed four-level features.

[0163] Calculate the mutual information value of the processed four-level features, and generate the association strength matrix of the processed four-level features based on the mutual information value. Based on the association strength matrix, use the accuracy and coverage of the processed four-level features as reward functions to determine the importance ranking of the processed four-level features.

[0164] Based on the importance ranking, a personal risk feature map is generated, and the features in the feature map are calibrated with a preset rule base to obtain a calibrated personal risk feature map.

[0165] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, performs the following steps:

[0166] Acquire multi-source heterogeneous data on risks in the power industry;

[0167] Based on a preset model, risk features are extracted from the multi-source heterogeneous data to obtain four-level features of the multi-source heterogeneous data.

[0168] The mapping relationship between the general risk ontology and the power industry ontology is determined. Based on the mapping relationship, the four-level features are adapted to the power industry, and the adapted four-level features are subjected to feature enhancement processing to obtain the processed four-level features.

[0169] Calculate the mutual information value of the processed four-level features, and generate the association strength matrix of the processed four-level features based on the mutual information value. Based on the association strength matrix, use the accuracy and coverage of the processed four-level features as reward functions to determine the importance ranking of the processed four-level features.

[0170] Based on the importance ranking, a personal risk feature map is generated, and the features in the feature map are calibrated with a preset rule base to obtain a calibrated personal risk feature map.

[0171] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0172] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.

[0173] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0174] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A method for extracting personal risk characteristics, characterized in that, The method includes: Acquire multi-source heterogeneous data on risks in the power industry; the preset model includes a basic feature layer, an interaction feature layer, a hidden feature layer, and an evolutionary feature layer; the four-level features include basic explicit features, interactive linkage features, hidden potential features, and evolutionary trend features; the multi-source heterogeneous data includes text data, image data, and sensor data; The multi-source heterogeneous data is input into the basic feature layer to extract the basic explicit features of the multi-source heterogeneous data; The basic explicit features are input into the interaction feature layer. Based on the graph neural network, an interaction model among the user, device and environment is constructed. The basic explicit features are processed through the interaction model to obtain the interaction linkage features of the multi-source heterogeneous data. The interactive linkage features are input into the latent feature layer, and the interactive linkage features are compared with the interactive linkage features of the verification data in the power industry to obtain the latent potential features of the multi-source heterogeneous data. The time-series data, basic explicit features, interactive linkage features, and latent features from the multi-source heterogeneous data are input into the evolutionary feature layer. Based on the time-series data, the basic explicit features, interactive linkage features, and latent features are fused to obtain the evolutionary trend features of the multi-source heterogeneous data. The mapping relationship between the general risk ontology and the power industry ontology is determined. Based on the mapping relationship, the four-level features are adapted to the power industry, and the adapted four-level features are subjected to feature enhancement processing to obtain the processed four-level features. Calculate the mutual information value of the processed four-level features, and generate the association strength matrix of the processed four-level features based on the mutual information value. Based on the association strength matrix, use the accuracy and coverage of the processed four-level features as reward functions to determine the importance ranking of the processed four-level features. Based on the importance ranking, a personal risk feature map is generated, and the features in the feature map are calibrated with a preset rule base to obtain a calibrated personal risk feature map.

2. The method according to claim 1, characterized in that, The evolutionary trend features include the risk change vector and time-varying sensitivity factor of the multi-source heterogeneous data.

3. The method according to claim 1, characterized in that, The process of generating a personal risk feature map based on the importance ranking includes: Determine the data source, extraction basis, and confidence value of the processed four-level features; Based on the importance ranking, the causal probability among the processed four-level features is transformed into the directional direction between the level where the cause feature is located and the level where the effect feature is located. Based on the data source, the extraction criteria, the confidence level, and the direction, a personal risk feature map is generated.

4. The method according to claim 1, characterized in that, The method further includes: Based on the features in the calibrated personal risk feature map, the directional interference vector is calculated; Based on the directional interference vector, Gaussian noise is injected into the basic feature layer of the preset model, and gradient masking attack is performed on the latent feature layer of the preset model. Calculate the adversarial loss function of the preset model, and optimize the parameters in the preset model based on the adversarial loss function to obtain the optimized preset model.

5. The method according to claim 1, characterized in that, Before acquiring multi-source heterogeneous data indicating risks in the power industry, the following steps are included: The text data is segmented and labeled to extract the risk subject, behavior, and object triples from the text data; The image data is processed using a multimodal processing model to generate a spatial semantic description vector for the image data. Extract the frequency domain energy features from the sensor data, and determine the time domain abrupt change points in the sensor data based on the frequency domain energy features.

6. A personal risk characteristic extraction device, characterized in that, The device includes: The acquisition module is used to acquire multi-source heterogeneous data on risks in the power industry; the preset model includes a basic feature layer, an interaction feature layer, a hidden feature layer, and an evolutionary feature layer; the four-level features include basic explicit features, interactive linkage features, hidden potential features, and evolutionary trend features; the multi-source heterogeneous data includes text data, image data, and sensor data; An extraction module is used to input the multi-source heterogeneous data into the basic feature layer to extract the basic explicit features of the multi-source heterogeneous data; input the basic explicit features into the interaction feature layer, construct an interaction model between users, devices, and the environment based on a graph neural network, process the basic explicit features through the interaction model to obtain the interaction linkage features of the multi-source heterogeneous data; input the interaction linkage features into the latent feature layer, compare the interaction linkage features with the interaction linkage features of verification data in the power industry to obtain the latent potential features of the multi-source heterogeneous data; input the time series data in the multi-source heterogeneous data, the basic explicit features, the interaction linkage features, and the latent potential features into the evolution feature layer, and fuse the basic explicit features, the interaction linkage features, and the latent potential features based on the time series data to obtain the evolution trend features of the multi-source heterogeneous data; The processing module is used to determine the mapping relationship between the general risk ontology and the power industry ontology. Based on the mapping relationship, the four-level features are adapted to the power industry, and the adapted four-level features are subjected to feature enhancement processing to obtain the processed four-level features. The determination module is used to calculate the mutual information value of the processed four-level features, generate the association strength matrix of the processed four-level features based on the mutual information value, and determine the importance ranking of the processed four-level features based on the association strength matrix, using the accuracy and coverage of the processed four-level features as reward functions. The calibration module is used to generate a personal risk feature map based on the importance ranking, and to calibrate the features in the feature map with a preset rule base to obtain a calibrated personal risk feature map.

7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 5.

9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Risk prediction method and device, equipment and storage medium

    CN113822494A

  • Internet marketing platform risk early warning management method and system

    CN119919144A

  • Power distribution network state sensing and control method and system based on industrial Internet of Things

    CN120474194A