Software vulnerability detection and protection method oriented to biological monitoring system
By simulating attack scenario feature datasets using multi-source sensor data and historical attack logs, and adopting abnormal state clustering and reinforcement learning algorithms, the prediction model is dynamically updated, and the collection frequency and reinforcement strategy are adjusted. This solves the dynamic adaptability problem of the biological monitoring system in complex attack scenarios and improves the security and reliability of the system.
Patent Information
- Application Number
- CN202510888771.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-10-03
AI Technical Summary
Existing biological monitoring systems lack dynamic adaptability and real-time performance when facing diverse and complex attack scenarios, which makes the system prone to failure under unknown threats, and reinforcement measures are difficult to balance system performance and security.
By simulating multi-source sensor data and historical attack logs, we generate attack scenario feature datasets, adopt abnormal state clustering and reinforcement learning algorithms, dynamically update the prediction model, adjust the collection frequency and reinforcement strategy, and combine adaptive encryption and resource optimization to achieve resilience strategy optimization against complex attacks.
Effectively respond to dynamic and complex attack environments, improve the security and reliability of biological monitoring systems, possess the ability of continuous learning and self-improvement, and enhance anti-attack capabilities and system stability.
Smart Images

Figure CN120744933A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of biological monitoring, and in particular to a software vulnerability detection and protection method for biological monitoring systems. Background Art
[0002] Biomonitoring systems play an irreplaceable role in healthcare, environmental monitoring, and public safety, and their reliability is directly related to human health and social stability. With technological advances, biomonitoring systems are widely used in complex scenarios, but they face increasingly severe attack threats, and system failures can have catastrophic consequences.
[0003] Existing solutions mostly focus on defending against a single attack scenario, such as data encryption or sensor protection, but lack the ability to comprehensively respond to diverse attack scenarios, and have obvious shortcomings in terms of system resilience and anti-attack capabilities. These methods are often unable to dynamically adapt to complex attack patterns, resulting in the system being prone to failure when facing unknown threats. The core challenges facing the current field stem from how to accurately predict failure modes under different attack scenarios and how to design reinforcement measures for these vulnerabilities. Predicting failure modes requires a comprehensive analysis of the system's behavior under multiple attacks, but the diversity and complexity of attack scenarios make it difficult for a single model to cover all possible situations. The resulting technical difficulty is that the system needs to quickly identify and adapt to unknown attacks in a dynamic environment, and existing technologies are insufficient in real-time and adaptability. This deficiency further makes it difficult to design reinforcement measures that balance system performance and security, and over-emphasizing one aspect may weaken overall resilience. Summary of the Invention
[0004] The purpose of the present invention is to provide a software vulnerability detection and protection method for biological monitoring systems to solve the problems existing in the prior art.
[0005] To achieve the above objectives, the present invention provides the following technical solution: a software vulnerability detection and protection method for a biological monitoring system, the method comprising: Multi-source sensor data streams and historical attack logs are obtained from the biomonitoring system. A feature dataset containing a state deviation matrix and attack scenario identifiers is generated through attack scenario simulation. An abnormal state clustering algorithm is used to classify the attack sequences and obtain a classified attack scenario set. Based on the classified attack scenario set, equilibrium state parameters are obtained. A reinforcement learning algorithm is used to optimize the resilience strategy based on these equilibrium state parameters. The acquisition frequency and real-time feedback channel strategy parameters are adjusted through attack scenario simulation. The strategy weights are optimized by combining state data caching to obtain an optimized resilience strategy configuration. According to the optimized resilience strategy configuration, the operating parameters of the biological monitoring system are updated. System status data is continuously collected through multi-source sensor fusion and time series data sharding. Combined with dynamic monitoring frequency adjustment, reinforcement strategy switching and adaptive encryption switching, the improved anti-attack capability parameters are obtained.
[0006] Preferably, obtaining the equilibrium state parameters based on the classified attack scenario set includes extracting the sensor data stream and state feature vector generated by multi-source sensor fusion through state space definition for the classified attack scenario set, training a single failure mode prediction model, and using a reward function to evaluate the prediction accuracy to obtain a prediction result for a complex failure mode.
[0007] Preferably, the method of obtaining the equilibrium state parameters according to the classified attack scenario set includes extracting the vulnerability features of the system under each failure mode according to the prediction results of the complex failure mode, constructing a vulnerability feature vector including state deviation weight and deviation trend analysis, determining the key vulnerability points through feature importance analysis, and obtaining a key vulnerability set.
[0008] Preferably, the equilibrium state parameters are obtained based on the classified attack scenario set, including updating the failure mode prediction model through an online learning algorithm if the vulnerability points in the key vulnerability point set match the state feature vectors extracted by real-time state monitoring and the anomaly detection results, adjusting the acquisition frequency adjustment parameters to adapt to the dynamic attack scenario, and obtaining an updated prediction model.
[0009] Preferably, the equilibrium state parameters are obtained according to the classified attack scenario set, including obtaining real-time sensor data streams for the updated prediction model, calculating the system state deviation including the state deviation matrix and deviation trend analysis through time series data sharding and data stream preprocessing, and if the state deviation exceeds the deviation threshold set, the reinforcement strategy switch is triggered to obtain the reinforcement strategy configuration.
[0010] Preferably, obtaining the equilibrium state parameters based on the classified attack scenario set includes encrypting the sensor data stream after data stream preprocessing using adaptive encryption switching according to the reinforcement strategy configuration, optimizing the state data cache and the calculation process of the real-time feedback channel through dynamic resource allocation, and obtaining the encrypted data stream and the optimized processing flow.
[0011] Preferably, the equilibrium state parameters are obtained based on the classified attack scenario set, including calculating the system performance indicators and security indicators including state deviation weights and outlier detection through encrypted data streams and optimized processing procedures. If the deviation between the performance indicator and the security indicator is less than a preset performance-safety balance threshold, it is determined that the system has reached an equilibrium state and the equilibrium state parameters are obtained.
[0012] Preferably, the abnormal state clustering algorithm adopts a density-based spatial clustering algorithm or a Gaussian mixture model.
[0013] Preferably, the reinforcement learning algorithm comprises a deep Q-network or proximal policy optimization.
[0014] Preferably, the state data cache optimization strategy dynamically adjusts the weight coefficient of historical data to improve learning efficiency through a sliding window mechanism and a priority cache strategy.
[0015] It can be seen from the above technical solution that the present invention has the following beneficial effects: This software vulnerability detection and protection method for biomonitoring systems simulates and generates attack scenario feature datasets using multi-source sensor data and historical attack logs, employs anomaly clustering and state-space modeling to predict failure modes and extract key vulnerabilities. Combined with real-time state monitoring, the present invention dynamically updates the prediction model, triggers reinforcement strategies based on state deviations, and employs adaptive encryption and resource optimization to protect data streams. By balancing performance and security indicators, the present invention utilizes reinforcement learning to optimize resilience strategies, adjusts acquisition frequency and feedback channels, and achieves a continuous improvement in the ability to counter attacks. The present invention can effectively respond to dynamic and complex attack environments and improve the security and reliability of biomonitoring systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 Flow chart of the method of the present invention. DETAILED DESCRIPTION
[0017] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0018] like Figure 1 As shown, the present invention provides a technical solution: a software vulnerability detection and protection method for a biological monitoring system, the method comprising: Multi-source sensor data streams and historical attack logs are obtained from the biomonitoring system. A feature dataset containing a state deviation matrix and attack scenario identifiers is generated through attack scenario simulation. An abnormal state clustering algorithm is used to classify the attack sequences and obtain a classified attack scenario set. Based on the classified attack scenario set, equilibrium state parameters are obtained. A reinforcement learning algorithm is used to optimize the resilience strategy based on these equilibrium state parameters. The acquisition frequency and real-time feedback channel strategy parameters are adjusted through attack scenario simulation. The strategy weights are optimized by combining state data caching to obtain an optimized resilience strategy configuration. According to the optimized resilience strategy configuration, the operating parameters of the biological monitoring system are updated. System status data is continuously collected through multi-source sensor fusion and time series data sharding. Combined with dynamic monitoring frequency adjustment, reinforcement strategy switching and adaptive encryption switching, the improved anti-attack capability parameters are obtained.
[0019] This implementation uses the biomonitoring system's multi-source sensor data streams and historical attack logs as input to simulate various potential attack scenarios. This generates a feature dataset consisting of a state deviation matrix (which measures the difference between the current system state and the normal state) and attack scenario identifiers (which characterize different types of attacks). An abnormal state clustering algorithm is then used to cluster the collected attack sequences, identifying sets of attack scenarios of different categories. Based on this, the system further extracts equilibrium state parameters for each attack scenario, which serve as indicators for evaluating the system's response performance under different attack conditions. Subsequently, a reinforcement learning algorithm is applied to train and optimize a resilience strategy based on the equilibrium state parameters. This strategy guides the system in adjusting data collection frequency and feedback channel parameters when facing an attack. Furthermore, a state data caching mechanism is incorporated to optimize strategy weights, ensuring the strategy remains effective and efficient under varying attack conditions. After verification and optimization through simulation experiments, the final resilience strategy configuration is obtained and applied to a real-world system to dynamically update operating parameters. During system operation, multi-source sensor fusion and time-series data slicing are continuously performed to reflect the system status in real time. At the same time, the monitoring frequency is dynamically adjusted, strategy switching is implemented, and the appropriate encryption mechanism is adaptively selected according to the threat level, ultimately obtaining system performance indicators with enhanced anti-attack capabilities.
[0020] This implementation has the following advantages: Comprehensive threat identification: Through multi-source data fusion and attack scenario simulation, it can identify complex and diverse attack patterns. Dynamic response and optimization: Combined with reinforcement learning algorithms, it automatically optimizes resilience strategies to achieve adaptive adjustments in the face of unknown threats. Enhanced system stability: Through state data caching and policy weight optimization, the robustness and stability of protection strategies are improved. Improved anti-attack capabilities: Dynamically adjust monitoring frequency and encryption strategies to improve the system's defense against persistent and advanced persistent threats (APTs). Continuous evolution capability: The system has the ability to continuously learn and self-improve, and can adjust protection strategies in a timely manner according to emerging attack methods.
[0021] According to the classified attack scenario set, the equilibrium state parameters are obtained, including extracting the sensor data stream and state feature vector generated by multi-source sensor fusion for the classified attack scenario set through state space definition, training a single failure mode prediction model, and using a reward function to evaluate the prediction accuracy to obtain the prediction results of complex failure modes.
[0022] Based on claim 1, this embodiment further refines the process for obtaining equilibrium state parameters. First, for a set of classified attack scenarios, by defining a state space suitable for the biomonitoring system, the data streams from the multi-source sensor fusion during system operation are extracted and converted into feature vectors representing the system state. These feature vectors contain key indicators reflecting the current system health and potential failure signals. Next, a single failure mode prediction model is trained using these state feature vectors. This model focuses on identifying and predicting a single type of failure that may occur in the system under specific attack scenarios. For example, single-point failures such as abnormal sensor response, data delay, or abnormal fluctuations. To improve model training, a reward function is established to measure the model's prediction accuracy over different training cycles. The reward function provides positive incentives for accurate predictions and penalties for incorrect predictions, thereby guiding the model to continuously optimize parameters and improve its ability to identify failure modes. Finally, by analyzing the predicted outputs of single failure modes, complex failure modes (i.e., system anomalies caused by the combination or interaction of multiple single failures) are further combined and identified to form a comprehensive prediction result. The prediction information of these complex failure modes will serve as an important basis for subsequent optimization of resilience strategies and equilibrium state parameters.
[0023] This implementation achieves comprehensive extraction of system states through state space definition and fusion of multi-source sensor data, ensuring the representativeness and completeness of feature vectors. Leveraging a single failure mode prediction model, it significantly improves the ability to identify potential single-point failures in the system. Furthermore, a reward function mechanism guides adaptive optimization of the model, continuously improving prediction accuracy. Furthermore, by combining the output of a single failure mode, it further identifies and predicts complex failure modes, enabling the system to handle multiple failures and complex attacks. This provides strong support for the scientific formulation of resilience strategies and ultimately significantly enhances the biomonitoring system's responsiveness and resilience to various threats.
[0024] According to the classified attack scenario set, the equilibrium state parameters are obtained, including the prediction results of complex failure modes, the extraction of the system's vulnerability features under each failure mode, the construction of a vulnerability feature vector including state deviation weight and deviation trend analysis, and the determination of key vulnerabilities through feature importance analysis to obtain a key vulnerability set.
[0025] This implementation further identifies potential vulnerable areas of the system based on prediction results based on complex failure modes. First, system vulnerability characteristics corresponding to each identified complex failure mode are extracted. These characteristics include the impact and sensitivity of each system component or parameter on overall functionality under a specific failure scenario. Subsequently, a state deviation weight is assigned to each vulnerability to quantify the severity of its deviation from normal operating conditions. Deviation trend analysis is also performed to identify whether the deviation exhibits a trend of continuous deterioration, fluctuation, or self-recovery. The state deviation weights are integrated with the deviation trend analysis results to form a vulnerability feature vector, which comprehensively characterizes the dynamic behavior and importance of each vulnerability. Next, a feature importance analysis algorithm (such as one based on Shapley value, information gain, or recursive feature elimination) is applied to evaluate the vulnerability feature vector to identify critical vulnerabilities that have a decisive impact on system stability and security. Ultimately, a set of critical vulnerabilities is generated. This set will be prioritized for reinforcement during subsequent resilience strategy optimization and system operating parameter adjustment.
[0026] This implementation not only identifies potential vulnerabilities in system operation through in-depth analysis of complex failure modes, but also comprehensively quantifies and dynamically tracks the importance of each vulnerability, combining state deviation weights and deviation trends. This improves the accuracy and real-time nature of identifying system weaknesses. Feature importance analysis further ensures that the identified key vulnerabilities are highly representative and relevant, providing a clear and scientific basis for prioritizing subsequent defense strategy optimization. This significantly enhances the resilience, stability, and overall risk mitigation capabilities of the biomonitoring system in the face of complex attacks or abnormal scenarios.
[0027] According to the classified attack scenario set, the equilibrium state parameters are obtained. If the vulnerable points in the key vulnerable point set match the state feature vectors extracted by real-time state monitoring and the anomaly detection results, the failure mode prediction model is updated through the online learning algorithm, and the acquisition frequency adjustment parameters are adjusted to adapt to the dynamic attack scenario to obtain the updated prediction model.
[0028] Based on the analysis results of the key vulnerability set, this embodiment further realizes the dynamic adaptation and continuous optimization of the failure mode prediction model. Specifically, the system continuously monitors real-time status data and identifies abnormal patterns during operation through the established state feature vector and anomaly detection mechanism. When the detection results show that the current state characteristics match a vulnerability in the key vulnerability set, it is determined that the vulnerability has actual performance in the current operating environment. Once the matching vulnerability is identified, the system starts the online learning algorithm and incrementally updates the original failure mode prediction model so that it can learn and adapt to the latest system behavior and threat characteristics. The online learning process avoids retraining of the entire data, reduces computing resource consumption and achieves rapid response. At the same time, according to the detected dynamic attack scenario, the frequency adjustment parameters of data collection are automatically adjusted to optimize the timeliness and accuracy of the data, ensuring that the prediction model maintains high prediction performance and responsiveness under new environmental conditions. Ultimately, an updated prediction model is formed and applied to subsequent status monitoring and failure prediction tasks to achieve continuous adaptive evolution of the model.
[0029] This implementation achieves on-demand online updates of the prediction model through real-time monitoring and dynamic matching of key vulnerability sets, improving the system's ability to quickly respond to emerging threats and abnormal conditions. The introduction of online learning algorithms avoids the frequent retraining of traditional models, significantly reducing update latency and computational costs, while maintaining the model's continuous learning capabilities. At the same time, dynamic adjustment of acquisition frequency parameters makes data collection and processing more efficient while still meeting prediction accuracy, further improving the system's adaptability and resource utilization efficiency, thereby significantly enhancing the resilience and continuous protection capabilities of the biomonitoring system in the face of rapidly evolving attack environments.
[0030] Based on the classified attack scenario set, the equilibrium state parameters are obtained, including obtaining real-time sensor data streams for the updated prediction model, and calculating the system state deviation including the state deviation matrix and deviation trend analysis through time series data sharding and data stream preprocessing. If the state deviation exceeds the deviation threshold set, the reinforcement strategy switch is triggered to obtain the reinforcement strategy configuration.
[0031] This embodiment is based on the updated prediction model to further realize the dynamic monitoring of the real-time status of the system and the switching of adaptive protection strategies. Specifically, the system continuously receives real-time data streams from various sensors. For the received data, the continuous data stream is first divided into time segments of fixed length through time series data slicing technology, which facilitates subsequent batch processing and feature extraction. Subsequently, data stream preprocessing operations are performed, including data cleaning, missing value filling, outlier smoothing and standardization, to ensure that the data quality meets the requirements of deviation analysis. Based on the preprocessed data, the state deviation matrix is calculated according to the current system state characteristics. The matrix quantifies the difference between the system operation state and the normal state defined by the prediction model, and combines the deviation trend analysis to identify the development direction and rate of the deviation. When the detected state deviation exceeds the pre-set deviation threshold set, it is determined that the system operation state may have entered the abnormal or attack impact range. At this time, the system automatically triggers the reinforcement strategy switching mechanism, and selects or adjusts the optimal reinforcement strategy configuration based on the current deviation characteristics, historical protection effects and feedback from the prediction model. This configuration may involve adjusting data acquisition parameters, enhancing encryption algorithms, changing the security level of feedback channels, or enabling specific isolation / redundancy mechanisms to enhance the system's adaptability and resistance to anomalies or attacks.
[0032] This implementation method achieves high-precision monitoring of system status and dynamic deviation analysis by closely integrating real-time sensor data with updated prediction models, enabling timely identification of potential operational anomalies or signs of attack. Utilizing a deviation threshold trigger mechanism, the system proactively switches to the reinforcement strategy that best suits the current environment before the deviation trend reaches a dangerous level, significantly improving the speed and accuracy of the protection response. Furthermore, through time-series data sharding and data preprocessing, the efficiency and consistency of data processing are improved, ensuring that the switching of reinforcement strategies is based on high-quality, reliable data, thereby effectively enhancing the continued resilience and defense capabilities of the biomonitoring system in complex and changing attack environments.
[0033] According to the classified attack scenario set, the equilibrium state parameters are obtained, including encrypting the sensor data stream after data stream preprocessing according to the reinforcement strategy configuration, using adaptive encryption switching, and optimizing the calculation process of the state data cache and real-time feedback channel through dynamic resource allocation to obtain the encrypted data stream and optimized processing flow.
[0034] Guided by the hardening policy configuration, this implementation further implements dynamic security hardening of system data streams and optimization of computational processes. Specifically, based on the selected hardening policy configuration, the system first performs adaptive encryption switching on the pre-processed sensor data in the data stream. This process dynamically selects the optimal encryption algorithm from a variety of pre-set encryption algorithms (such as AES, RSA, ECC, or quantum encryption schemes) based on the current threat level, system load, and real-time performance requirements, balancing security and processing efficiency. The system then dynamically allocates resources for the processed data stream and its associated computational processes. Through an intelligent scheduling mechanism, the system optimizes the allocation of state data caches and real-time feedback channels based on the availability and priority of computing resources (such as CPU, memory, and network bandwidth), ensuring maximum resource utilization for data processing and communication without compromising system responsiveness. Ultimately, the system obtains an encrypted data stream and establishes an optimized processing flow with enhanced security, processing efficiency, and scalability, ensuring continuous data confidentiality and efficient system operation in a dynamic attack environment.
[0035] This implementation combines a hardening strategy with adaptive encryption switching to achieve flexible encryption of data streams based on varying security requirements, enhancing the system's protection against dynamic threats. Furthermore, through a dynamic resource allocation mechanism, it intelligently optimizes the computational processes of state data caching and real-time feedback channels, significantly improving resource utilization and system response speed, and avoiding the impact of resource bottlenecks on system performance. This solution not only ensures data transmission security but also improves the processing efficiency and scalability of the biomonitoring system in complex operating environments, thereby comprehensively enhancing the system's resilience, stability, and anti-attack capabilities.
[0036] According to the classified attack scenario set, the equilibrium state parameters are obtained, including the encrypted data stream and optimized processing flow, and the system performance indicators and security indicators including state deviation weights and outlier detection are calculated. If the deviation between the performance indicator and the security indicator is less than the preset performance-security balance threshold, the system is judged to have reached a balanced state and the equilibrium state parameters are obtained.
[0037] This implementation, based on encrypted data streams and optimized processing flows, further enables dynamic assessment and balance of overall system performance and security. Specifically, during continuous system operation, the system uses encrypted data streams and state data from optimized processing flows to calculate key system performance metrics (such as data processing speed, response time, and resource utilization) and security metrics (such as encryption strength, outlier detection frequency, and intrusion detection trigger rate). During this calculation, the system combines state deviation weights to measure the deviation between current performance and historical baselines. It also uses outlier detection to identify potential security threats or abnormal events, thereby comprehensively reflecting the health and security of the current system operation. The system then calculates the deviation between the performance and security metrics and compares it with a preset performance-security balance threshold. If the deviation is below the threshold, the system indicates that the system has achieved an ideal balance between current resource allocation, processing power, and security protection. At this point, the system determines that equilibrium has been achieved and derives the equilibrium parameters based on these parameters, providing a basis for subsequent resilience policy adjustments and resource allocation. If the deviation exceeds the threshold, the system triggers a policy adjustment process to automatically optimize the relevant parameters to achieve equilibrium.
[0038] This implementation establishes a multi-dimensional dynamic assessment mechanism for system operating status by integrating state deviation weights, outlier detection, performance indicators, and security metrics, enabling real-time assessment of the balance between system performance and security. This approach not only promptly identifies potential performance bottlenecks or security risks but also dynamically adjusts policies to maintain optimal operating conditions, thereby improving the system's adaptability and stability in complex environments. Furthermore, by setting a performance-security balance threshold, the system maintains a high level of security without sacrificing critical performance, effectively enhancing the overall resilience and sustained protection capabilities of the biomonitoring system.
[0039] The abnormal state clustering algorithm adopts a density-based spatial clustering algorithm or a Gaussian mixture model. This embodiment is directed to the abnormal state clustering algorithm described in claim 1, and preferably adopts a density-based spatial clustering algorithm (such as DBSCAN, OPTICS) or a Gaussian mixture model (GMM) to classify the attack sequence. When a density-based spatial clustering algorithm is adopted, the method can divide the data into different clusters according to the density relationship of the data points, and effectively identify noise or abnormal points, which is suitable for processing irregularly shaped attack sequence data. In particular, when facing complex or non-spherical data distributions, density clustering does not require a preset number of clusters and has a high degree of flexibility and adaptability. When a Gaussian mixture model is adopted, the system assumes that the attack sequence data conforms to a mixture model composed of multiple Gaussian distributions, and iteratively estimates the parameters through the expectation maximization (EM) algorithm, and finally divides the data into multiple probabilistic clusters. This method is particularly suitable for situations where the data has clear statistical characteristics or there may be overlapping clusters. It can output the probability that each data point belongs to each cluster, thereby improving the flexibility and accuracy of classification. In practical applications, based on the characteristics, distribution patterns, and data volume of the attack scenario data, we select or dynamically switch between density clustering and Gaussian mixture models to obtain the best classification effect, forming a set of classified attack scenarios and providing a solid data foundation for subsequent equilibrium state parameter extraction and resilience strategy optimization.
[0040] This implementation significantly improves the classification capability and robustness of the abnormal state clustering algorithm in complex data environments by adopting a density-based spatial clustering algorithm or a Gaussian mixture model. Density clustering does not require a preset number of clusters and can adaptively identify complex and irregular data structures while effectively handling outliers. The Gaussian mixture model has powerful statistical modeling capabilities and is suitable for data sets with overlapping or fuzzy boundaries, and provides probabilistic classification results for data points. The combination or optimal application of the two methods not only improves the accuracy and flexibility of attack scenario identification, but also enhances the system's adaptability when facing different data features, thereby laying a solid foundation for the scientific formulation and dynamic adjustment of subsequent resilience strategies, and further enhancing the comprehensive response capabilities of the biomonitoring system to security threats.
[0041] Reinforcement learning algorithms include deep Q-networks or proximal policy optimization. Based on the method described in claim 1, this embodiment preferably employs deep Q-networks (DQNs) or proximal policy optimization (PPO) as reinforcement learning algorithms for the resilience strategy optimization process. When using a deep Q-network, the algorithm combines traditional Q-learning methods with deep neural networks, approximating the state-action value function (Q-value) through the neural network, thereby solving complex problems with high-dimensional state spaces that are difficult for traditional Q-learning to handle. During system operation, the DQN continuously learns the optimal resilience strategy based on the current equilibrium state parameters and historical experience, dynamically optimizing the data acquisition frequency, feedback channel parameters, and strategy weights. When proximal policy optimization is used, the algorithm belongs to the policy gradient method, which ensures the stability and efficiency of the training process by limiting the step size of each policy update. PPO exhibits excellent convergence performance in a continuous action space and is suitable for resilience strategy optimization in complex and dynamically changing attack scenarios. Based on real-time feedback and simulated attack scenario results, the system uses PPO to continuously adjust strategy parameters to enhance its adaptive capabilities in response to new threats. Based on data complexity, real-time requirements, and system resource conditions, you can flexibly select DQN or PPO, or dynamically switch based on training results to ensure that the resilience strategy remains in the optimal or suboptimal state.
[0042] This implementation significantly enhances the learning and adaptability of resilience strategy optimization by employing a deep Q-network or proximal policy optimization algorithm. DQN can effectively address biomonitoring systems with high-dimensional state spaces and complex state changes, improving the accuracy and robustness of policy decisions. PPO, with its excellent stability and convergence speed, is suitable for handling continuous action decisions and frequently changing attack environments. The application of these two algorithms not only improves the resilience strategy's response efficiency to unknown threats, but also optimizes data collection, feedback adjustment, and the dynamic management of policy weights, ultimately enhancing the biomonitoring system's intelligent protection level and sustained defense capabilities.
[0043] The state data cache optimization strategy dynamically adjusts the weight coefficients of historical data to improve learning efficiency through a sliding window mechanism and a priority caching strategy. Based on the method described in claim 1, this embodiment specifically designs a state data cache optimization strategy, combining a sliding window mechanism with a priority caching strategy to enhance the intelligence of data management and learning efficiency. First, the system introduces a sliding window mechanism into cache management, windowing historical state data in chronological order. The sliding window automatically moves forward, retaining only the latest time-segment data that is most relevant to the current learning task. This prevents excessive invalid or outdated data from interfering with model learning and improves cache timeliness. Second, the system applies a priority caching strategy, assigning different priorities to historical data in the cache based on its importance, access frequency, and relevance to the current learning objective. High-priority data is retained for a long time, while low-priority data is eliminated first when cache space is limited. The combined effects of the sliding window and priority caching allow the system to dynamically adjust the weight coefficients of each historical data item in the cache. The weight coefficients not only consider the temporal properties of the data, but also comprehensively consider the data's impact in the recent learning process and its access frequency. This mechanism ensures that more important and representative data are given higher learning weights during model training, thereby accelerating convergence and improving the accuracy and generalization ability of the model.
[0044] This implementation method combines a sliding window mechanism with a priority caching strategy to dynamically adjust the weight of historical data, achieving efficient management of cache resources and significantly improving model learning efficiency. The sliding window ensures the timeliness of cached data, preventing old data from affecting learning results; the priority caching strategy rationally allocates resources based on the importance of the data, improving the utilization of key data. At the same time, the dynamic adjustment of the weight coefficient enables the learning process to focus on the most representative data, improving the training speed and accuracy of the model. In addition, this strategy reduces unnecessary data storage and processing, optimizes the utilization of system resources, and provides strong support for the continuous learning and self-adaptation capabilities of the biomonitoring system.
[0045] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A software vulnerability detection and protection method for a biological monitoring system, characterized in that: The method comprises: Multi-source sensor data streams and historical attack logs are obtained from the biomonitoring system. A feature dataset containing a state deviation matrix and attack scenario identifiers is generated through attack scenario simulation. An abnormal state clustering algorithm is used to classify the attack sequences and obtain a classified attack scenario set. Based on the classified attack scenario set, equilibrium state parameters are obtained. A reinforcement learning algorithm is used to optimize the resilience strategy based on these equilibrium state parameters. The acquisition frequency and real-time feedback channel strategy parameters are adjusted through attack scenario simulation. The strategy weights are optimized by combining state data caching to obtain an optimized resilience strategy configuration. According to the optimized resilience strategy configuration, the operating parameters of the biological monitoring system are updated. System status data is continuously collected through multi-source sensor fusion and time series data sharding. Combined with dynamic monitoring frequency adjustment, reinforcement strategy switching and adaptive encryption switching, the improved anti-attack capability parameters are obtained.
2. The method for detecting and protecting software vulnerabilities in a biological monitoring system according to claim 1, wherein: Obtaining equilibrium state parameters based on the classified attack scenario set includes extracting sensor data streams and state feature vectors generated by multi-source sensor fusion through state space definition for the classified attack scenario set, training a single failure mode prediction model, and using a reward function to evaluate prediction accuracy to obtain prediction results for complex failure modes.
3. The method for detecting and protecting software vulnerabilities in a biological monitoring system according to claim 2, wherein: The method of obtaining the equilibrium state parameters based on the classified attack scenario set includes extracting the vulnerability features of the system under each failure mode based on the prediction results of the complex failure mode, constructing a vulnerability feature vector including state deviation weight and deviation trend analysis, determining the key vulnerability points through feature importance analysis, and obtaining a key vulnerability point set.
4. The method for detecting and protecting software vulnerabilities in a biological monitoring system according to claim 3, wherein: The equilibrium state parameters are obtained based on the classified attack scenario set, including if the vulnerable points in the key vulnerable point set match the state feature vectors and anomaly detection results extracted by real-time state monitoring, then the failure mode prediction model is updated through an online learning algorithm, and the acquisition frequency adjustment parameters are adjusted to adapt to the dynamic attack scenario to obtain an updated prediction model.
5. The method for detecting and protecting software vulnerabilities in a biological monitoring system according to claim 4, wherein: The equilibrium state parameters are obtained according to the classified attack scenario set, including obtaining real-time sensor data streams for the updated prediction model, calculating the system state deviation including the state deviation matrix and deviation trend analysis through time series data sharding and data stream preprocessing, and if the state deviation exceeds the deviation threshold set, triggering the reinforcement strategy switch to obtain the reinforcement strategy configuration.
6. The method for detecting and protecting software vulnerabilities in a biological monitoring system according to claim 5, wherein: The method of obtaining the equilibrium state parameters based on the classified attack scenario set includes encrypting the sensor data stream after data stream preprocessing using adaptive encryption switching according to the reinforcement strategy configuration, optimizing the calculation process of the state data cache and the real-time feedback channel through dynamic resource allocation, and obtaining the encrypted data stream and the optimized processing process.
7. The method for detecting and protecting software vulnerabilities in a biological monitoring system according to claim 6, wherein: The equilibrium state parameters are obtained based on the classified attack scenario set, including calculating the system performance indicators and security indicators including state deviation weights and outlier detection through encrypted data streams and optimized processing flows. If the deviation between the performance indicator and the security indicator is less than a preset performance-security balance threshold, it is determined that the system has reached an equilibrium state and the equilibrium state parameters are obtained.
8. The method for detecting and protecting software vulnerabilities in a biological monitoring system according to claim 1, wherein: The abnormal state clustering algorithm adopts a density-based spatial clustering algorithm or a Gaussian mixture model.
9. The method for detecting and protecting software vulnerabilities in a biological monitoring system according to claim 1, wherein: The reinforcement learning algorithms include deep Q-networks or proximal policy optimization.
10. The method for detecting and protecting software vulnerabilities in a biological monitoring system according to claim 1, wherein: The state data cache optimization strategy dynamically adjusts the weight coefficient of historical data to improve learning efficiency through a sliding window mechanism and a priority cache strategy.