Insurance business data processing system for reinforcement learning
By integrating IoT devices and employee behavior data and utilizing cross-validation and reinforcement learning models, the problem of unreliable data in traditional insurance business has been resolved, dynamic risk assessment and intelligent decision-making have been achieved, and the refined management capabilities of insurance business have been enhanced.
Patent Information
- Application Number
- CN202510742887.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-09-19
AI Technical Summary
Traditional insurance businesses rely on static historical data in the risk assessment and pricing process, making it difficult to capture and respond to dynamic changes in corporate operations in real time. In addition, employee behavioral compliance data is subjective and unreliable, affecting the reliability of model judgment and decision-making.
Integrate real-time data from IoT devices and employee behavior compliance data, ensure data quality through cross-validation, and input verified multi-dimensional data streams into the reinforcement learning model to dynamically optimize decision-making strategies and form a complete intelligent closed loop.
It realizes the intelligent management of corporate insurance business, outputs targeted dynamic pricing, risk warning and fraud identification, and improves the accuracy and effectiveness of insurance business processing.
Smart Images

Figure CN120670798A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing, and more specifically, to an insurance business data processing system based on reinforcement learning. Background Art
[0002] Currently, traditional corporate insurance businesses face challenges in risk assessment, pricing, and claims settlement, primarily due to their reliance on static historical data and periodic audits, which makes it difficult to capture and respond to dynamically changing risks in business operations in real time. The widespread adoption of the Internet of Things (IoT) enables businesses to collect massive amounts of real-time data on device status, operating environment, and personnel activities through various sensors and devices. Combined with employee behavior and compliance data generated during business operations, this provides an unprecedented data foundation for refined and dynamic insurance management. In this context, the introduction of reinforcement learning (RL) technology has become a key path to intelligent insurance operations.
[0003] During the research, we discovered that ensuring the reliability of reinforcement learning model outputs requires quality control of input data. This is especially true given the differences between real-time data from IoT devices and employee behavioral compliance data. IoT data is typically automatically generated by devices, is highly objective, and reflects hard facts about the physical world. In contrast, employee behavioral compliance data may involve manual recording or reporting, which carries the risk of subjectivity, incompleteness, and even errors or fraud. The performance of reinforcement learning models is highly dependent on the accuracy and authenticity of input data. If learning is based on unreliable employee behavioral data, the model will make erroneous judgments and decisions, impacting the fairness and effectiveness of insurance business processing.
[0004] Therefore, an optimized insurance business data processing system based on reinforcement learning is expected. Summary of the Invention
[0005] In order to solve the above technical problems, the present application is proposed. The embodiment of the present application provides a reinforcement learning insurance business data processing system, which integrates objective IoT real-time data representing the physical environment and equipment status and employee behavior compliance data reflecting human factors and process compliance, and innovatively uses IoT data to cross-validate employee behavior data to ensure data quality. Subsequently, the verified high-quality multi-dimensional data stream is input into the reinforcement learning model, which dynamically optimizes its decision-making strategy through continuous learning environment feedback (such as enterprise operating status and risk changes), and finally outputs targeted and intelligent enterprise insurance business processing results (such as dynamic pricing, risk warning, fraud identification), thereby forming a complete intelligent closed loop from data collection, verification, learning to decision-making.
[0006] According to one aspect of the present application, a reinforcement learning insurance business data processing system is provided, comprising:
[0007] The data acquisition module is used to obtain real-time data from IoT devices and employee behavior compliance data of target enterprise objects;
[0008] The data verification module is used to verify employee behavior compliance data based on real-time data from IoT devices to obtain verification results;
[0009] The enterprise insurance business processing result generation module is used to input the IoT device real-time data and employee behavior compliance data into the reinforcement learning model to obtain the enterprise insurance business processing result in response to the verification result that the employee behavior compliance data verification is qualified.
[0010] Compared with the existing technology, the present application provides a reinforcement learning insurance business data processing system, which integrates objective IoT real-time data representing the physical environment and equipment status and employee behavior compliance data reflecting human factors and process compliance, and innovatively uses IoT data to cross-validate employee behavior data to ensure data quality. Subsequently, the verified high-quality multi-dimensional data stream is input into the reinforcement learning model, which dynamically optimizes its decision-making strategy through continuous learning environment feedback (such as enterprise operating status and risk changes), and ultimately outputs targeted, intelligent enterprise insurance business processing results (such as dynamic pricing, risk warning, fraud identification), thereby forming a complete intelligent closed loop from data collection, verification, learning to decision-making. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The above and other purposes, features, and advantages of the present application will become more apparent through a more detailed description of the embodiments of the present application in conjunction with the accompanying drawings. The accompanying drawings are intended to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation of the present application. In the drawings, the same reference numerals generally represent the same components or steps.
[0012] Figure 1 1 is a block diagram of an insurance business data processing system based on reinforcement learning according to an embodiment of the present application;
[0013] Figure 2 Schematic diagram of data flow in an insurance business data processing system using reinforcement learning according to an embodiment of the present application;
[0014] Figure 3 This is a block diagram of a data verification module in an insurance business data processing system based on reinforcement learning according to an embodiment of the present application. DETAILED DESCRIPTION
[0015] Below, the exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application, and it should be understood that the present application is not limited to the exemplary embodiments described herein.
[0016] As used in this application and the claims, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" are not intended to refer to the singular but may include the plural. Generally speaking, the terms "comprises" and "include" only indicate the inclusion of the steps and elements specifically identified, and these steps and elements do not constitute an exclusive list. A method or apparatus may also include other steps or elements.
[0017] Although the present application makes various references to certain modules in the system according to embodiments of the present application, any number of different modules can be used and run on the user terminal and / or server. The modules are illustrative only, and different aspects of the system and method can use different modules.
[0018] Flowcharts are used in this application to illustrate the operations performed by the systems according to the embodiments of the present application. It should be understood that the preceding or following operations are not necessarily performed in exact order. Instead, the various steps may be processed in reverse order or simultaneously, as needed. Furthermore, other operations may be added to these processes, or one or more operations may be removed from these processes.
[0019] Below, the exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application, and it should be understood that the present application is not limited to the exemplary embodiments described herein.
[0020] In the technical solution of this application, a reinforcement learning insurance business data processing system is proposed. Figure 1 1 is a block diagram of an insurance business data processing system based on reinforcement learning according to an embodiment of the present application;
[0021] Figure 2 Schematic diagram of data flow of insurance business data processing system based on reinforcement learning according to the embodiment of the present application. Figure 1 and Figure 2As shown, the insurance business data processing system based on reinforcement learning according to the embodiment of the present application includes: a data acquisition module 310, which is used to obtain the real-time data of the IoT device and the employee behavior compliance data of the target enterprise object; a data verification module 320, which is used to perform data verification on the employee behavior compliance data based on the real-time data of the IoT device to obtain a verification result; an enterprise insurance business processing result generation module 330, which is used to input the IoT device real-time data and the employee behavior compliance data into the reinforcement learning model in response to the verification result that the employee behavior compliance data verification is qualified to obtain the enterprise insurance business processing result.
[0022] In particular, the data acquisition module 310 is used to acquire real-time IoT device data and employee behavior compliance data of the target enterprise object. The real-time IoT device data of the target enterprise object includes device vibration data, gas concentration, environmental coefficient, wearable device data, and location data; the employee behavior compliance data includes safety operation frequency, device operation log, safety procedure execution timestamp, safety training completeness, and employee self-report. It should be understood that the spatiotemporal correlation between equipment operating status and personnel operations in industrial scenarios directly affects the triggering logic of risk events. For example, when a gas concentration sensor detects an anomaly, the system needs to combine whether employees perform gas leak emergency operations in accordance with safety regulations during the same period (verified by the timestamp in the equipment operation log) to accurately determine whether the anomaly is an equipment failure or a human error. Therefore, in order to establish a causal chain between equipment state evolution and employee operation behavior, in the technical solution of this application, the real-time IoT device data and employee behavior compliance data of the target enterprise object are acquired. The system can construct a multi-dimensional information field covering the physical state of the equipment (such as vibration data reflecting the degree of mechanical wear) and the trajectory of personnel behavior (such as wearable device data locating the range of employee activities), providing a data basis for subsequent spatiotemporal correlation reasoning. This data acquisition strategy provides input features that combine physical authenticity and behavioral logic for end-to-end risk assessment and dynamic premium calculation, achieving a leap from isolated event detection to systemic risk management.
[0023] In particular, the data verification module 320 is used to perform data verification on employee behavior compliance data based on real-time data from IoT devices to obtain verification results. In a specific example of this application, Figure 3As shown, the data verification module 320 includes: an employee behavior log record extraction unit 321, which is used to extract employee behavior log records for a preset time period from employee behavior compliance data; an IoT device time series data extraction unit 322, which is used to extract a time queue of IoT device real-time data for a preset time period from IoT device real-time data; a time series semantic interaction encoding unit 323, which is used to perform time series semantic interaction encoding on the employee behavior log records and the time queue of IoT device real-time data to obtain a device state-employee operation event interaction response reasoning encoding vector; and a verification unit 324, which is used to generate a verification result based on the device state-employee operation event interaction response reasoning encoding vector.
[0024] Specifically, the employee behavior log record extraction unit 321 is used to extract employee behavior log records for a preset time period from employee behavior compliance data. That is, extracting employee behavior log records (such as power on / off timestamps and safety procedure execution records in the equipment operation log) through a preset time window (such as 30 minutes before the equipment anomaly) is essentially to construct a spatiotemporal correlation framework between equipment state mutations and operational behavior antecedents. This time slicing mechanism enables the system to focus on key operation cycles, such as aligning the alarm period of the gas concentration sensor with the time series of emergency operation actions recorded by the employee's wearable device, thereby verifying the timeliness of the execution of safety regulations. In addition, this time-constrained data extraction not only avoids the noise interference brought by the full amount of data, but more importantly, establishes a time anchor point for causal inference - for example, whether employees with low safety training completeness show operational delay characteristics before equipment overload. The extraction of such time-sensitive features directly affects the effectiveness of subsequent feature-level interaction coding.
[0025] Specifically, the IoT device time series data extraction unit 322 is used to extract a time queue of IoT device real-time data for a preset time period from the IoT device real-time data. Here, by extracting the time series data queue (including environmental factors, wearable device location, etc.) of a specific time window (such as 1 hour before the device alarm), the system can reconstruct the complete trajectory of device state changes and provide a chain of evidence in the time dimension for tracing the cause of the risk. This time queue-based feature extraction provides the reinforcement learning model with a complete view of the evolution of the device state, ultimately achieving a paradigm shift from passive abnormality alarms to active risk intervention.
[0026] Specifically, the temporal semantic interaction encoding unit 323 is used to perform temporal semantic interaction encoding on the time queue of employee behavior log records and IoT device real-time data to obtain a device status-employee operation event interaction response inference encoding vector. That is, in an embodiment of the present application, first, the employee behavior log records are passed through a semantic embedding encoder based on the BERT model to obtain a semantic embedding encoding vector of employee operation device events. It should be understood that employee device operation logs often contain specialized log entries. These textual information not only involves the accurate understanding of industry terms, but also requires capturing the logical associations between operation steps. Traditional text processing methods are insufficient in capturing the contextual semantics of operation events. The BERT model, through the multi-layer Transformer structure obtained in its pre-training stage, can parse the sequential dependencies in employee behavior log records. Therefore, in order to construct a high-dimensional semantic representation space for operation behavior, in the technical solution of the present application, the employee behavior log records are passed through a semantic embedding encoder based on the BERT model to obtain a semantic embedding encoding vector of employee operation device events. During this process, BERT's multi-head attention mechanism automatically focuses on key fields (such as "failed to meet standards") during encoding while maintaining an understanding of the entire context of the operation event. This enables the generated semantic embedding encoding vector of employee operation equipment events to reflect the multi-dimensional characteristics of operation quality, laying the foundation for subsequent interactive encoding with time-series equipment data.
[0027] Next, the time queue of the IoT device's real-time data is passed through a time series encoder based on the LSTM model to obtain a device state time series pattern feature encoding vector. It should be understood that during the operation of the device, the time queue of parameters such as temperature and pressure does not fluctuate independently, but is a complex system output formed by the coupling of multiple factors such as device health and operating load. When processing time series data such as vibration sensors and environmental parameters, traditional methods often use sliding window statistics or discretized feature extraction methods, resulting in the implicit gradual change laws and mutation signs in the device operation state being separated. The LSTM model, with its unique gating mechanism (forget gate, input gate, output gate) and cell state transmission path, can effectively model the nonlinear evolution pattern of device parameters in the time dimension. Therefore, in the technical solution of the present application, the time queue of the IoT device's real-time data is passed through a time series encoder based on the LSTM model to obtain a device state time series pattern feature encoding vector. During this process, the LSTM time series encoder maps the time series of device parameters into a high-dimensional vector containing state transition laws through multi-level nonlinear transformations. This vector not only contains the instantaneous characteristics of the current moment, but also encodes deep physical laws such as the device performance degradation rate and abnormal event triggering thresholds, providing an accurate time series baseline for building a causal reasoning model of device status and operating behavior.
[0028] Furthermore, the semantic embedding encoding vector of the employee operation device event and the device state temporal pattern feature encoding vector are encoded using a device-event semantic mapping based on time alignment to obtain a device state-employee operation event interactive response reasoning encoding vector. It should be understood that there is an implicit causal or accompanying relationship between device state changes and employee operation events in the temporal dimension. However, traditional methods lack a fine-grained time alignment mechanism, resulting in the separation of the interactive response patterns between the two. For example, when a gas concentration sensor detects an anomaly, it is difficult to trace whether the employee has followed the complete process of safety regulations during that period through isolated timestamp comparisons alone. The temporal misalignment between device vibration data and operation logs may mask the delayed impact of illegal operations on device status. Therefore, it is necessary to eliminate the time drift between the device data stream and the behavioral event stream through time alignment to construct a cross-modal spatiotemporal correlation benchmark. Therefore, in the technical solution of this application, the semantic embedding encoding vector of the employee operation device event and the device state temporal pattern feature encoding vector are encoded using a device-event semantic mapping based on time alignment to obtain a device state-employee operation event interactive response reasoning encoding vector. That is, by capturing the cross-modal temporal dependency between dynamic signals such as equipment vibration and gas concentration and operational events (such as the timestamp of safety procedure execution), the limitations of one-way causal inference in traditional threshold verification are broken through. Specifically, in this process, first, through the interaction of local implicit features and the chain reasoning mechanism, the temporal evolution of equipment status and the semantic logic of operational events are integrated into an interpretable interaction pattern, such as identifying the compound risk characteristics of employees not wearing protective equipment according to regulations during the period of rising gas concentration; then, the attention weight correction mechanism is adopted to strengthen the interactive response of key timing nodes in the abnormal event chain (such as the overlapping period of equipment failure precursor period and high-frequency illegal operation) and suppress the noise interference of non-critical time windows; in addition, the global risk evolution path is formed through the LSTM chain reasoning engine, so that the interactive impact of equipment status and operational behavior can be traced back to continuous time periods, providing high-order feature support for the risk accumulation effect modeling of the reinforcement learning model, and ultimately achieving consistent optimization of risk assessment in the spatiotemporal dimension. The generated equipment state-employee operation event interaction response reasoning encoding vector essentially constructs a joint state space of equipment and behavior, enabling the reinforcement learning model to quantify the associated potential energy of risk units on the time axis, and realize the paradigm upgrade of premium calculation from single-point risk assessment to dynamic risk situation awareness.
[0029] Specifically, first, local latent feature extraction based on one-dimensional convolutional coding is performed on the semantic embedding encoding vector of employee operation equipment events and the feature encoding vector of equipment status temporal pattern, respectively, to obtain a set of local latent feature encoding vectors of employee operation equipment events and a set of local latent feature encoding vectors of equipment status temporal pattern. It should be understood that the equipment status temporal feature vector carries the continuous evolution law of physical parameters (such as the energy distribution trend of the vibration spectrum), while the operation semantic vector contains the logical association between discrete events (such as the execution order of procedure steps). There are significant differences in the expression dimension and abstraction level of the two in the feature space. Traditional global interaction methods directly fully connect and fuse high-dimensional feature vectors, which easily leads to fine-grained temporal patterns (such as the inflection point characteristics of the equipment temperature curve) and key semantic fragments (such as the exception description field in the operation log) being overwhelmed by macro features. Therefore, in order to construct a differential association framework between the physical state of the equipment and the logic of the operation behavior, in the technical solution of the present application, the semantic embedding coding vector of the employee operation equipment event and the equipment status temporal pattern feature coding vector are respectively subjected to local implicit feature extraction based on one-dimensional convolutional coding to obtain a set of local implicit feature coding vectors of the employee operation equipment event and a set of local implicit feature coding vectors of the equipment status temporal pattern.
[0030] That is, through parameter-sharing convolution kernels, local contextual relationships of different scales are dynamically perceived, and multi-level feature primitives are constructed for subsequent interactive reasoning. In this process, the set of local implicit vectors extracted from the device's timing pattern features through one-dimensional convolution essentially decomposes the continuous signal into feature primitives with clear physical meaning (such as pressure mutation intervals and vibration cycle phases); the local implicit set of operation semantic vectors is the semantic deconstruction of the behavior log (such as the logical units of operation steps and the dependencies of procedural clauses). By designing convolution kernels of different scales, the system establishes multi-level and multi-scale local association anchors in the feature extraction stage, providing structured spatiotemporal analysis units for chain reasoning.
[0031] In a specific example of the present application, the semantic embedding coding vector of the employee operation device event and the feature coding vector of the device status time series pattern are respectively subjected to local implicit feature extraction based on one-dimensional convolution coding to obtain a set of local implicit feature coding vectors of the employee operation device event and a set of local implicit feature coding vectors of the device status time series pattern; wherein the one-dimensional convolution formula is:
[0032] Conv l×1 (X)={x1,x2,...,x i ,...,x n}
[0033] Conv l×1 (Y)={y1,y2,...,yi ,...,y n}
[0034] Among them, X is the semantic embedding encoding vector of the employee operation device event, Conv l×1 (·) represents a one-dimensional convolution operation, x1,x2,...,x i ,...,x n are the first, second, i-th and n-th local implicit feature coding vectors of the employee operation device event in the set of local implicit feature coding vectors of the employee operation device event, Y is the device state time series pattern feature coding vector, y1, y2, ..., y i ,...,y n They are respectively the 1st, 2nd, i-th and n-th local implicit feature coding vectors of the device state timing pattern in the set of the local implicit feature coding vectors of the device state timing pattern.
[0035] Next, each set of corresponding local implicit feature coding vectors for employee-operated device events and local implicit feature coding vectors for device state temporal patterns are input into a device state-employee operation event single feature interaction engine to obtain a set of device state-employee operation event local implicit feature interaction response coding vectors. Because traditional risk verification methods often use linear weighting or simple splicing strategies when dealing with the correlation between device state and operation logs, it is difficult to capture complex causal relationships such as device state and operation logs. Therefore, in order to establish a set of implicit differential equations between the device physical space and the operation semantic space, in the technical solution of the present application, each set of corresponding local implicit feature coding vectors for employee-operated device events and local implicit feature coding vectors for device state temporal patterns are input into a device state-employee operation event single feature interaction engine to obtain a set of device state-employee operation event local implicit feature interaction response coding vectors. Here, by constructing a differential action unit of local feature pairs, it is possible to decouple the asymmetric association between local patterns of device status (such as abnormal fluctuation ranges in temperature sensor data) and semantic fragments of operation events (such as missing key steps in maintenance records), overcoming the defect of traditional methods that key signals are obliterated due to global feature mixing.
[0036] In a specific example of the present application, the following feature interaction formula is used to input each corresponding set of employee operation device event local implicit feature coding vectors and device state timing pattern local implicit feature coding vectors in the set of employee operation device event local implicit feature coding vectors and the set of device state timing pattern local implicit feature coding vectors into the device state-employee operation event monomer feature interaction engine to obtain a set of device state-employee operation event local implicit feature interaction response coding vectors; wherein, the feature interaction formula is:
[0037]
[0038] Among them, concat(·) represents cascade, ⊙ represents point multiplication by position, It means adding by position. Indicates difference by position, W i and b i Represent the weight matrix and bias vector respectively, v i Represents x i and y i The corresponding equipment status-employee operation event local implicit feature interaction response encoding vector.
[0039] Then, based on the characteristic distribution characteristics of each local latent feature interaction response encoding vector within the set of local latent feature interaction response encoding vectors for equipment status and employee operation events, the chained inference attention weights for each local latent feature interaction response encoding vector are determined to obtain the initial set of chained inference attention weights for equipment status and employee operation events. In other words, by establishing a dynamic weight allocation mechanism, the model can autonomously identify key indicators such as the coupling strength between sudden equipment status changes and abnormal employee operations, as well as the risk transmission speed. During periods of overlap between resonance phenomena in specific frequency bands of the equipment vibration spectrum and frequent illegal employee operations, sudden entropy changes or covariance differences in the feature vectors often trigger adaptive increases in attention weights. Conversely, interaction features corresponding to standard operations under stable operating conditions receive lower weights due to their convergent distribution. This mechanism enables the system to transcend the traditional risk model's reliance on explicit thresholds. For example, when gas concentration has not reached the alarm threshold but is showing an exponential upward trend, potential leak risks can be identified in advance by analyzing the gradient distribution characteristics of the interaction feature vectors. By generating initial attention weights, the system achieves dual optimization of both spatiotemporal focusing and noise suppression of risk signals.
[0040] In a specific example of the present application, based on the characteristic distribution characteristics of each device state-employee operation event local implicit feature interaction response encoding vector in the set of device state-employee operation event local implicit feature interaction response encoding vectors, the following attention formula is used to determine the chain reasoning attention weight of each device state-employee operation event local implicit feature interaction response encoding vector to obtain a set of initial device state-employee operation event chain reasoning attention weights; wherein, the attention formula is:
[0041]
[0042] Among them, softmax(·) represents the softmax function, v i,j Represents the eigenvalue of the jth position of the i-th device state-employee operation event local latent feature interaction response encoding vector in the set of device state-employee operation event local latent feature interaction response encoding vectors, ||·|| 2 represents the square of the norm, L represents the length of the equipment status-employee operation event local implicit feature interaction response encoding vector, a i Indicates v i Corresponding initial equipment state-employee operation event chain reasoning attention weight.
[0043] Then, the set of initial equipment state-employee operation event chain reasoning attention weights is spatially corrected based on the interaction specification to obtain the set of equipment state-employee operation event chain reasoning attention weights. In particular, it should be understood that the traditional attention weight calculation method regards the interaction between equipment state and operation behavior as a homogenization process, ignoring the differential action mechanism of different interaction specifications (such as the gradual migration of the physical state of the equipment and the discrete transition of the operation behavior) in the feature space. In one example, when the equipment vibration mode presents a steady fluctuation (translation-dominated), its interaction weight distribution with the conventional operation record should be different from the interaction mode of sudden gas leakage (fluctuation-dominated) and the lack of emergency operation. Therefore, in order to construct an interaction specification differential geometry space with physical interpretability, in a preferred example of the present application, the set of initial equipment state-employee operation event chain reasoning attention weights is spatially corrected based on the interaction specification to obtain the set of equipment state-employee operation event chain reasoning attention weights.
[0044] Specifically, by mathematically decoupling translational effects (with gradients aligned with the interaction manifold) from fluctuation effects (orthogonal to the primary interaction direction), the system overcomes the attention weight bias problem caused by normative confusion in traditional methods. Specifically, by calculating the periodic compensation of fluctuation effects by the local logarithmic size of the translation (e.g., the periodic variation in operational response delay caused by accumulated equipment wear), as well as the auxiliary quantities of fluctuation-induced translational phase transitions (e.g., the perturbation effect of sudden failures on routine maintenance cycles), the system reconstructs the normative manifold of the equipment-operator interaction in Hilbert space. This mathematical correction mechanism enables attention weights to reflect not only the strength of the interaction but also the topological characteristics of the risk transmission path. For example, operational delays during periods of high-frequency vibration are identified as high-risk nodes dominated by fluctuations, while slight deviations from routine maintenance cycles are classified as low-risk units dominated by translations, thus achieving normative field quantitative modeling of risk energy. This norm-sensitive attention allocation mechanism enables reinforcement learning models to accurately quantify the differentiated pricing factors of "translation norm cumulative risk" and "fluctuation norm sudden risk" when optimizing premium strategies, promoting the evolution of insurance business from single risk event response to risk normative field situation awareness.
[0045] In this preferred example, the set of attention weights of the initial device state-employee operation event chain reasoning is spatially corrected based on the interaction specification using the following spatial correction formula to obtain the set of attention weights of the device state-employee operation event chain reasoning; wherein, the spatial correction formula is:
[0046]
[0047] ζ=e γ ×(α+β)
[0048] a′ i =(ω1×ξ+ω2×ζ)a i
[0049] Among them, α and β represent the statistics of the translation action vector, γ represents the statistics of the fluctuation action vector, ξ represents the periodic local regular compensation, ζ represents the periodic local auxiliary phase transition, ω1 and ω2 represent the modulation weights, and a′ i Indicates v i The corresponding equipment status-employee operation event chain reasoning attention weight.
[0050] Specifically, in this process, when calculating the local implicit feature interaction response coding vector of each device state-employee operation event, the interaction features between the corresponding local implicit feature coding vector of the employee operation device event and the local implicit feature coding vector of the device state time series pattern are introduced, such as x i ⊙y i , Etc., that is, they essentially correspond to different interaction norms, and thus have different spatial auxiliary constraint associations in the interaction space.
[0051] Therefore, in order to enhance the norm variability assistance in chain reasoning attention weights, it is preferred to modify the set of initial equipment state-employee operation event chain reasoning attention weights based on the spatial action decomposition of the interaction norm. is regarded as a translation effect, that is, the gradient direction is consistent with the spatial interaction direction, and x i ⊙y i It is regarded as a wave effect, that is, the gradient direction is orthogonal to the spatial interaction direction. In this way, for vector statistics of different effects, for example γ=||x i ⊙y i || 2 , it is believed that the wave action leads to a localized periodic effect in the translation direction, and thus the periodic local regular compensation is calculated as:
[0052]
[0053] That is, if α+β is expressed as translational localization, the wave action γ increases as the logarithmic size of the translational localization increases.
[0054] On the other hand, the wave action will also lead to a localized translational phase transition, thus obtaining a periodic local auxiliary phase transition:
[0055] ζ=e γ ×(α+β)
[0056] Therefore, the initial equipment state-employee operation event chain reasoning attention weight a is modified based on the weighted sum of the above two items i :
[0057] a′ i =(ω1×ξ+ω2×ζ)a i
[0058] By distinguishing the normative roles of different interaction norms corresponding to the local implicit feature interaction responses of each device status-employee operation event in the interaction space, the correlation between different spatial auxiliary constraint norms in the interaction space can be improved, thereby improving the calculation accuracy of the attention weight of the device status-employee operation event chain reasoning.
[0059] In particular, based on the set of attention weights of the equipment state-employee operation event chain reasoning, the set of equipment state-employee operation event local implicit feature interaction response coding vectors is weighted modulated to obtain a set of modulated equipment state-employee operation event local implicit feature interaction response coding vectors. It should be understood that when traditional methods deal with the correlation between equipment parameters (such as gas concentration, environmental coefficient) and operation records (such as safety training completeness, procedure execution timing), the signal strength of key risk conduction nodes is often diluted by background noise due to the global average weighting mechanism. Therefore, in the technical solution of the present application, based on the set of attention weights of the equipment state-employee operation event chain reasoning, the set of equipment state-employee operation event local implicit feature interaction response coding vectors is weighted modulated to obtain a set of modulated equipment state-employee operation event local implicit feature interaction response coding vectors.
[0060] Specifically, by chaining the spatial mapping of attention weights, a risk-sensitivity-driven feature modulation mechanism is constructed to establish a dynamic importance scale for interactive features, enabling the system to autonomously adjust the strength of feature expression based on risk transmission rules. During this process, when local features of device state (such as anomalous clustering of wearable device location data) are spatiotemporally coupled with local features of operational behavior (such as discrete offsets in safety procedure execution timestamps), the system calculates the spatial effects of their interaction norms using attention weights (e.g., translational effects characterize the gradual impact of environmental parameters, while fluctuation effects reflect the impact of sudden operational errors). This norm-sensitive modulation strategy significantly increases the information density of high-weight interaction units (e.g., records of operational slack during slowly varying environmental coefficients) in the feature space, while suppressing low-weight units (e.g., compliant operations during periods of regular parameter fluctuations) as background noise. In this way, the system possesses the ability to topologically reconstruct risk transmission pathways.
[0061] In a specific example of the present application, based on a set of attention weights of the device state-employee operation event chain reasoning, a set of device state-employee operation event local implicit feature interaction response coding vectors is weighted modulated using the following modulation formula to obtain a set of modulated device state-employee operation event local implicit feature interaction response coding vectors; wherein, the modulation formula is:
[0062] I={I1,I2,...,I i ,...,I n}
[0063] I i =a i ·v i
[0064] Where I is the set of the modulation device state-employee operation event local implicit feature interaction response coding vectors, I1, I2, ..., I i ,...,I n They are respectively the 1st, 2nd, i-th and n-th modulation device state-employee operation event local implicit feature interaction response coding vectors in the set of the modulation device state-employee operation event local implicit feature interaction response coding vectors.
[0065] Furthermore, the set of modulated equipment state-employee operation event local implicit feature interaction response encoding vectors is input into a chained inference engine based on a forward LSTM model to obtain an equipment state-employee operation event interaction response inference encoding vector. Considering that abnormal equipment operating parameters and employee operational errors form a "risk transmission chain" within a continuous production cycle, for example, a single unperformed equipment maintenance operation may trigger a cascading state degradation in subsequent processes, such implicit risk dependencies across time windows cannot be captured through isolated event detection. Traditional methods lack the ability to retain historical interaction states, making it difficult to construct a temporal causal graph between equipment state evolution and employee behavior patterns, resulting in risk assessment lagging behind the actual risk evolution process. To establish a dynamic memory model of equipment-behavior interaction states, enabling the system to automatically track the propagation path of risk signals, the technical solution of this application uses the sequential inference mechanism of a forward LSTM to reconstruct discrete local interaction events in time and space into a continuous risk evolution trajectory. In one example, for long-term risk patterns such as abnormal employee operation frequency accompanying the gradual change of equipment pressure curves, the LSTM gating mechanism selectively memorizes key nodes (such as the operation record when the parameter first deviates from the baseline) and forgets non-critical background noise. Furthermore, for the instantaneous correlation between sudden equipment alarms and emergency operation responses, precise timing alignment is achieved through cell state updates. This approach overcomes the causal inversion problem caused by traditional models that ignore the direction of the time arrow.
[0066] In a specific example of the present application, a set of modulated device state-employee operation event local implicit feature interaction response encoding vectors is input into a chain inference engine based on a forward LSTM model using the following inference formula to obtain a device state-employee operation event interaction response inference encoding vector; wherein, the inference formula is:
[0067] v f =LSTM(I)
[0068] Among them, LSTM(·) represents the forward LSTM, v f A coding vector is inferred for the device status-employee operation event interaction response.
[0069] Specifically, the verification unit 324 is configured to generate a verification result based on the device state-employee operation event interaction response inference encoding vector. Specifically, in the technical solution of the present application, the device state-employee operation event interaction response inference encoding vector is passed through a decoder-based verification result generator to obtain a verification result. The verification result is used to represent the difference between employee behavior compliance data and real-time data from IoT devices. It should be understood that although the device state-employee operation event interaction response inference encoding vector carries spatiotemporal correlation features, its high-dimensional abstract form cannot be directly mapped into compliance indicators that are understandable to the business. In the technical solution of the present application, the decoder-based verification result generator interprets implicit interaction patterns in the feature space into explicit difference values, for example, converting the "temporal mismatch between the environmental deterioration rate and the protection operation response efficiency" into a quantifiable risk deviation. In this process, after the device state temporal pattern (such as the abnormal clustering of employee locations monitored by wearable devices) and the semantic features of operational behavior (such as the discreteness of safety procedure execution timestamps) are deeply fused through the encoder, the decoder reconstructs the compliance difference between the two through inverse feature mapping. The generated difference value reflects the compliance status at the current moment, giving insurance risk assessment dynamic traceability capabilities.
[0070] 11. In particular, the enterprise insurance business processing result generation module 330 is configured to, in response to a verification result indicating that the employee behavior compliance data has passed verification, input the IoT device real-time data and employee behavior compliance data into a reinforcement learning model to obtain the enterprise insurance business processing result. The reward function of the reinforcement learning model includes a continuous cycle reward and a risk event penalty. It should be understood that when the system verifies that there is a compliance correlation between operational behavior and device status (such as a match between the safety procedure execution timestamp and the environmental parameter fluctuation trend), it is still necessary to evaluate the sustainability of this compliance status and its impact on long-term risks. The reinforcement learning model, through a continuous cycle reward mechanism (such as exponential reward accumulation for continued compliance) and a risk event penalty function (such as a gradient penalty triggered by a sudden leak event), can model the nonlinear relationship between the "time decay rate of operational compliance" and the "dynamic contraction of the device status safety boundary," overcoming the drawback of traditional models that simplifies premium calculation as a superposition of discrete events. This enables insurance companies to identify companies that appear to be compliant but have hidden risk transmission paths (such as factories that rely on high-frequency safety operations to compensate for aging equipment), and to design differentiated risk-sharing plans, promoting the transformation and upgrading of industrial insurance from a loss compensation tool to a collaborative safety production governance platform.
[0071] As described above, the insurance business data processing system 300 using reinforcement learning according to an embodiment of the present application can be implemented in various wireless terminals, such as a server equipped with an insurance business data processing algorithm using reinforcement learning. In one possible implementation, the insurance business data processing system 300 using reinforcement learning according to an embodiment of the present application can be integrated into a wireless terminal as a software module and / or a hardware module. For example, the insurance business data processing system 300 using reinforcement learning can be a software module in the operating system of the wireless terminal, or can be an application developed for the wireless terminal; of course, the insurance business data processing system 300 using reinforcement learning can also be one of the many hardware modules of the wireless terminal.
[0072] Alternatively, in another example, the insurance business data processing system 300 based on reinforcement learning and the wireless terminal may also be separate devices, and the insurance business data processing system 300 based on reinforcement learning may be connected to the wireless terminal via a wired and / or wireless network and transmit interactive information in accordance with an agreed data format.
[0073] While various embodiments of the present disclosure have been described above, the above descriptions are illustrative, non-exhaustive, and not intended to be limiting of the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A reinforcement learning insurance business data processing system, characterized by: include: The data acquisition module is used to obtain real-time data from IoT devices and employee behavior compliance data of target enterprise objects; A data verification module is used to perform data verification on employee behavior compliance data based on real-time data from IoT devices to obtain verification results. The data verification on employee behavior compliance data based on real-time data from IoT devices is used to: analyze the temporal mapping relationship between employee behavior compliance data and real-time data from IoT devices through feature-level interactive coding to generate verification results; The enterprise insurance business processing result generation module is used to input the IoT device real-time data and employee behavior compliance data into the reinforcement learning model to obtain the enterprise insurance business processing result in response to the verification result that the employee behavior compliance data verification is qualified.
2. The insurance business data processing system based on reinforcement learning according to claim 1, characterized in that: The real-time data of IoT devices of target enterprises includes equipment vibration data, gas concentration, environmental coefficients, wearable device data and location data; employee behavior compliance data includes safety operation frequency, equipment operation logs, safety procedure execution timestamps, safety training completeness and employee self-reports.
3. The insurance business data processing system based on reinforcement learning according to claim 1, characterized in that: Data verification module, including: An employee behavior log record extraction unit, configured to extract employee behavior log records for a preset time period from employee behavior compliance data; An IoT device time series data extraction unit, configured to extract a time queue of IoT device real-time data in a preset time period from IoT device real-time data; The temporal semantic interaction encoding unit is used to perform temporal semantic interaction encoding on the time queues of employee behavior log records and IoT device real-time data to obtain the device status-employee operation event interaction response reasoning encoding vector; The verification unit is used to generate a verification result based on the inference coding vector of the equipment status-employee operation event interaction response.
4. The insurance business data processing system based on reinforcement learning according to claim 3, characterized in that: Temporal semantic interaction encoding unit, including: The semantic embedding encoding subunit is used to pass the employee behavior log records through the semantic embedding encoder based on the BERT model to obtain the semantic embedding encoding vector of the employee operation device event; The time series encoding subunit is used to pass the time queue of IoT device real-time data through the time series encoder based on the LSTM model to obtain the device state time series pattern feature encoding vector; The device-event semantic mapping encoding subunit is used to perform device-event semantic mapping encoding based on time alignment on the semantic embedding encoding vector of the employee operation device event and the device status temporal pattern feature encoding vector to obtain the device status-employee operation event interaction response inference encoding vector.
5. The insurance business data processing system based on reinforcement learning according to claim 4, characterized in that: The device-event semantic mapping encoding subunit includes: The local latent feature interaction secondary sub-unit is used to calculate the set of equipment state-employee operation event local latent feature interaction response coding vectors based on the local latent coding features of the employee operation equipment event semantic embedding coding vector and the equipment state temporal pattern feature coding vector; a feature modulation secondary subunit, configured to perform feature modulation based on a chained inference attention mechanism on the set of device state-employee operation event local implicit feature interaction response encoding vectors based on the feature distribution characteristics of each device state-employee operation event local implicit feature interaction response encoding vector in the set of device state-employee operation event local implicit feature interaction response encoding vectors, so as to obtain a set of modulated device state-employee operation event local implicit feature interaction response encoding vectors; The chain reasoning secondary subunit is used to perform device state-employee operation event state chain reasoning on the set of modulated device state-employee operation event local implicit feature interaction response coding vectors to obtain the device state-employee operation event interaction response reasoning coding vector.
6. The insurance business data processing system based on reinforcement learning according to claim 5, characterized in that: The local hidden feature interaction secondary subunit is used to: Performing local latent feature extraction based on one-dimensional convolutional coding on the semantic embedding coding vector of employee operation equipment events and the feature coding vector of equipment status temporal pattern, respectively, to obtain a set of local latent feature coding vectors of employee operation equipment events and a set of local latent feature coding vectors of equipment status temporal pattern; Each corresponding group of local implicit feature coding vectors of employee operation device events and local implicit feature coding vectors of device status timing patterns in the set of local implicit feature coding vectors of employee operation device events and the set of local implicit feature coding vectors of device status timing patterns are respectively input into the device status-employee operation event monomer feature interaction engine to obtain a set of device status-employee operation event local implicit feature interaction response coding vectors.
7. The insurance business data processing system based on reinforcement learning according to claim 6, characterized in that: Characteristic modulation secondary subunit, used for: Based on the characteristic distribution characteristics of each device state-employee operation event local latent feature interaction response encoding vector in the set of device state-employee operation event local latent feature interaction response encoding vectors, determine the chain reasoning attention weight of each device state-employee operation event local latent feature interaction response encoding vector to obtain an initial device state-employee operation event chain reasoning attention weight set; Performing spatial correction on the initial set of attention weights of the chain reasoning of device state-employee operation events based on interaction norms to obtain the set of attention weights of the chain reasoning of device state-employee operation events; Based on the set of attention weights of the device state-employee operation event chain reasoning, the set of device state-employee operation event local implicit feature interaction response encoding vectors is weighted modulated to obtain a set of modulated device state-employee operation event local implicit feature interaction response encoding vectors.
8. The insurance business data processing system based on reinforcement learning according to claim 7, characterized in that: Chain reasoning secondary subunit, used for: The set of modulated device state-employee operation event local implicit feature interaction response encoding vectors is input into the chain inference engine based on the forward LSTM model to obtain the device state-employee operation event interaction response inference encoding vector.
9. The insurance business data processing system based on reinforcement learning according to claim 8, characterized in that: Verification unit, used for: The device status-employee operation event interaction response inference encoding vector is passed through a decoder-based verification result generator to obtain a verification result. The verification result is used to represent the difference value between employee behavior compliance data and real-time data of IoT devices.
10. The insurance business data processing system based on reinforcement learning according to claim 1, characterized in that: The reward function of the reinforcement learning model includes continuous period rewards and risk event penalties.
Citation Information
Cited By
Sewage quality prediction and dynamic treatment method and system
CN121502153A