A method and system for loading and presenting a persisted Flink job file
By using a blockchain consensus mechanism inspired by bio-intelligence and a dynamic trust scoring model, abnormal file operations are identified and dynamic encryption keys are generated, solving the single point of failure and security threat problems of the traditional Flink job file management system and achieving high security and efficient management.
Patent Information
- Application Number
- CN202510907607.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-02
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2045-07-02
AI Technical Summary
Traditional Flink job file management systems suffer from single point of failure risks, lack of intelligent anomaly detection capabilities in file version management, and difficulty in dealing with complex security threats due to static key encryption, thus failing to meet the requirements of high concurrency, high availability, and high security.
Employing a blockchain consensus mechanism inspired by biological intelligence, it identifies abnormal file operations through immune memory cells, establishes a dynamic trust scoring model, generates dynamic encryption keys, realizes encrypted transmission and storage of file fragments, and triggers an immune response mechanism to isolate suspicious nodes when suspicious operations are detected.
It achieves secure storage and efficient management of Flink job files, effectively prevents malicious attacks, improves system security and reliability, solves the problems of data synchronization and version conflict in traditional methods, and optimizes system performance and user experience.
Smart Images

Figure CN120408687B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to a file loading and display technology, in particular to a persistent Flink job file loading and display method and system. BACKGROUND
[0002] With the rapid development of distributed computing technology, Apache Flink as an open source computing framework for stream processing and batch processing has been widely applied in the field of big data processing. As a core component of big data processing, the persistent storage, version management and secure access of Flink job files have become an important requirement in enterprise-level applications. Traditional Flink job file management usually adopts a centralized storage method, mainly relying on file systems or databases for storage and management, and using conventional encryption and permission control mechanisms to ensure data security.
[0003] In the traditional Flink job file management system, the persistent storage of files often depends on a single storage node or a centralized storage architecture, and the tracking of file versions is mainly realized through simple timestamps or version numbers, while the secure access control relies on static username and password authentication and fixed permission allocation mechanisms. With the growth of big data processing needs and the complexity of security threats, this traditional file management method has been difficult to meet the requirements of high concurrency, high availability and high security.
[0004] The existing Flink job file management technology has the following main defects: first, the traditional centralized file storage method has a single point of failure risk, when the storage node fails, the files in the entire system become inaccessible, affecting the normal operation of Flink jobs and business continuity. Secondly, the existing file version management mechanism lacks intelligent anomaly detection capability, making it difficult to effectively identify and prevent malicious file operations, increasing the risk of system attacks. Finally, the traditional static key encryption and fixed permission model is difficult to cope with complex and changing security threats, especially in the face of advanced persistent threats, it lacks adaptive security protection capabilities and cannot achieve dynamic identification and response to suspicious behavior. SUMMARY
[0005] The embodiment of the application provides a persistent Flink job file loading and display method and system, which can solve the problems in the prior art.
[0006] In a first aspect, the embodiment of the application provides a persistent Flink job file loading and display method, comprising:
[0007] Obtaining a file identifier, a version mark and metadata information of a Flink job file, performing sharded storage on the Flink job file based on the file identifier, the version mark and the metadata information, and generating a file shard set and file fingerprint information;
[0008] Based on the file shard set and the file fingerprint information, a bio-intelligent inspired blockchain consensus mechanism is constructed, which identifies abnormal file operations through immune memory cells and establishes a dynamic trust score model of file operation behavior based on the principle of antigen-antibody reaction;
[0009] Based on the abnormal file operation and the dynamic trust score model, a blockchain node network is deployed and a dynamic encryption key is generated, which is used to establish a secure communication channel in the blockchain node network for encrypted transmission and storage of file shards;
[0010] When receiving a file access request, identity authentication and permission verification are performed based on the dynamic trust score model and the secure communication channel, and when a suspicious operation is detected, an immune response mechanism is triggered to isolate the suspicious node;
[0011] When a file version is updated, the legality of the update operation is verified using the file fingerprint information and the dynamic trust score model, and after verification, the file shard information of the new version is stored in the blockchain node network through the secure communication channel.
[0012] Based on the file shard set and the file fingerprint information, a bio-intelligent inspired blockchain consensus mechanism is constructed, which identifies abnormal file operations through immune memory cells and establishes a dynamic trust score model of file operation behavior based on the principle of antigen-antibody reaction, which includes:
[0013] Based on the file shard set and the file fingerprint information, the file shard set is divided into a plurality of feature subsets, and an immune memory cell model is constructed according to the feature hash value and the correlation degree in the feature subset, the immune memory cell model including a feature recognition layer, a pattern matching layer and a response output layer, wherein the feature recognition layer extracts the behavior characteristics of file operation based on the feature hash value, the pattern matching layer establishes a feature similarity calculation matrix based on the correlation degree, and the response output layer generates an immune response result according to the feature similarity calculation matrix;
[0014] The file operation information is fused with the behavior features to generate an antigen feature vector, and a feature mode of the pattern matching layer is represented as an antibody feature vector; a binding strength of the antigen feature vector and the antibody feature vector is calculated based on the feature similarity calculation matrix, and the binding strength is obtained by calculating a cosine similarity between the feature vectors; the immune response result is combined with the binding strength by weighting to obtain an immune response coefficient, and the immune response coefficient is mapped to a dynamic trust score model.
[0015] The feature recognition layer extracts behavior features of file operations based on the feature hash values, the pattern matching layer establishes a feature similarity calculation matrix based on the correlation degree, and the response output layer generates an immune response result according to the feature similarity calculation matrix, including:
[0016] In the feature recognition layer, the feature hash values are input to a prefrontal pattern calculation unit, the prefrontal pattern calculation unit extracts behavior features of file operations based on a neuron spontaneous firing mechanism to generate an initial feature vector;
[0017] The initial feature vector is input to a dynamic adaptive network, the dynamic adaptive network enhances the initial feature vector based on a synaptic plasticity mechanism to generate an enhanced feature vector, the enhanced feature vector inherits the time sequence features and spatial features of the initial feature vector and has a dynamic adaptive capability;
[0018] A feature similarity calculation matrix is constructed based on the enhanced feature vector, an antigen-antibody dynamic balance mechanism is introduced into the feature similarity calculation matrix, the feature similarity calculation matrix is input to an immune memory unit, the immune memory unit optimizes the feature similarity calculation matrix based on an immune tolerance mechanism to generate an optimized feature mode;
[0019] A bee swarm collaborative network is constructed based on the optimized feature mode, the bee swarm collaborative network includes a plurality of collaborative calculation nodes, and adjacent collaborative calculation nodes exchange feature information through a pheromone transmission mechanism; a group decision result is generated based on the feature information exchange result in the bee swarm collaborative network, and a final immune response result is generated according to a comparison result of the group decision result and a preset immune tolerance threshold.
[0020] Based on the abnormal file operation and the dynamic trust score model, a blockchain node network is deployed and a dynamic encryption key is generated, including:
[0021] A feature matrix of the abnormal file operation is obtained, the feature matrix includes operation type, abnormal score, timestamp and operation location information, a dynamic trust score is calculated according to the feature matrix, and the dynamic trust score is constructed as a multi-dimensional feature tensor;
[0022] According to the abnormal score corresponding to each operation position and the component value of the multi-dimensional feature tensor, a node deployment density of the operation position is calculated, and the node deployment density is obtained by weighted combination of an abnormal score weight coefficient and a feature tensor weight coefficient;
[0023] The multi-dimensional feature tensor is dimensionally reconstructed to obtain a key seed matrix, and the multi-dimensional feature tensor is shape-transformed in a preset dimension, and the preset dimension is determined by the node deployment density; the key seed matrix is input into a chaotic mapping model for nonlinear transformation, control parameters of the chaotic mapping model are dynamically adjusted according to the dynamic trust score, and an initial key sequence is obtained;
[0024] The information entropy value of the initial key sequence is calculated, the information entropy value is obtained by calculating the logarithmic product of the probability of occurrence of each character in the initial key sequence, the iteration number of the chaotic mapping model is adjusted according to the information entropy value, and a dynamic encryption key that meets a preset entropy value threshold is generated.
[0025] The key seed matrix is input into a chaotic mapping model for nonlinear transformation, control parameters of the chaotic mapping model are dynamically adjusted according to the dynamic trust score, and an initial key sequence is obtained, including:
[0026] The key seed matrix is converted into neuron input potential values, each neuron input potential value corresponds to a neuron position, a neuromorphic chaotic computing unit is constructed based on the neuron input potential values, and the neuromorphic chaotic computing unit includes a membrane potential and a synaptic weight;
[0027] The firing threshold of the neuromorphic chaotic computing unit and the synaptic weight are adjusted according to the dynamic trust score, a spiking neural network is constructed based on the adjusted firing threshold and synaptic weight, and the spiking neural network has a spatiotemporal dynamic characteristic;
[0028] State information of the spiking neural network is input into the neuromorphic chaotic computing unit, a time sequence evolution process of the spiking neural network is calculated through the neuromorphic chaotic computing unit, an output signal of the spiking neural network is decoded to obtain a chaotic sequence, and the chaotic sequence inherits the dynamic characteristic of the spiking neural network;
[0029] Environmental noise information generated by micro-particle motion is obtained, the environmental noise information is input into the neuromorphic chaotic computing unit for processing, and the processed noise information is fused with the chaotic sequence to obtain a chaotic sequence with enhanced randomness;
[0030] Detect phase transition information in the micro-particle motion, trigger parameter adjustment of the chaotic sequence based on the phase transition information, update the chaotic sequence with enhanced randomness according to the adjusted parameters, and generate an initial key sequence.
[0031] Based on the dynamic trust score model and the secure communication channel, perform identity authentication and permission verification, and when detecting suspicious operation, trigger an immune response mechanism to isolate the suspicious node, including:
[0032] Input the basic communication channel information and identity authentication information into the life encoder, which converts the basic communication channel information into entangled state information; input the entangled state information and trust score information into the gene entanglement unit based on the secure communication channel, biologically encode the entangled state information based on the gene entanglement unit, and generate a gene authentication result;
[0033] Input the gene authentication result and the secure communication channel into the life key defense unit, which generates defense channel information based on a biological key distribution mechanism, and the defense channel information inherits the biological features of the gene authentication result;
[0034] Input the trust score information, the gene authentication result, and the defense channel information into the permission evaluation unit, which calculates a comprehensive permission score according to a preset weight coefficient, compares the comprehensive permission score with a dynamic threshold, and generates a permission verification result;
[0035] Detect the deviation value between the current operation and the normal operation baseline, and when the deviation value exceeds a preset immune threshold, trigger an immune response mechanism, which biologically isolates suspicious nodes based on the permission verification result, and the biological isolation realizes immune shielding of the suspicious nodes through the defense channel information.
[0036] Biological encoding of the entangled state information based on the gene entanglement unit to generate a gene authentication result includes:
[0037] The entangled state information contains the association features between communication nodes, the entangled state information is input into the gene entanglement unit, the gene entanglement unit converts the association features into gene sequences based on a DNA sequence mapping mechanism, and the gene sequences contain the biological features of the entangled state information;
[0038] Based on the gene sequence, construct an authentication template, dynamically encode the authentication template through a gene recombination mechanism to generate a gene authentication result, and the gene authentication result inherits the biological features of the gene sequence.
[0039] In a second aspect, the embodiment of the present application provides a persistent Flink job file loading and display system, comprising:
[0040] A first unit is configured to acquire a file identifier, a version mark and metadata information of a Flink job file, perform sharded storage on the Flink job file based on the file identifier, the version mark and the metadata information, and generate a file shard set and file fingerprint information;
[0041] A second unit is configured to construct a bio-intelligent inspired blockchain consensus mechanism based on the file shard set and the file fingerprint information, identify abnormal file operations through immune memory cells, and establish a dynamic trust score model of file operation behaviors based on an antigen-antibody reaction principle;
[0042] A third unit is configured to deploy a blockchain node network and generate a dynamic encryption key based on the abnormal file operations and the dynamic trust score model, the dynamic encryption key is used to establish a secure communication channel in the blockchain node network, and the secure communication channel is used to realize encrypted transmission and storage of file shards;
[0043] A fourth unit is configured to perform identity authentication and permission verification based on the dynamic trust score model and the secure communication channel when a file access request is received, and trigger an immune response mechanism to isolate a suspicious node when a suspicious operation is detected.
[0044] A fifth unit is configured to perform legality verification on an update operation by using the file fingerprint information and the dynamic trust score model when a file version is updated, and store file shard information of a new version into the blockchain node network through the secure communication channel after verification.
[0045] In a third aspect, the embodiment of the present application provides an electronic device, comprising:
[0046] A processor;
[0047] A memory for storing processor-executable instructions;
[0048] The processor is configured to invoke the instructions stored in the memory to execute the method described above.
[0049] In a fourth aspect, the embodiment of the present application provides a computer readable storage medium having computer program instructions stored thereon, the computer program instructions being executed by a processor to implement the method described above.
[0050] The present application has the following advantages:
[0051] The application provides a kind of persistent Flink job file loading and display method, by biological intelligent inspired blockchain consensus mechanism and dynamic trust score model, realize the safe storage and efficient management of Flink job file, effectively prevent malicious attack and illegal access, significantly improve the security and reliability of system.
[0052] The method adopts file fragment storage technology combined with dynamic encryption key, establishes a multi-level protection system, not only ensures the security of data transmission and storage process, but also can actively identify and isolate suspicious nodes through immune mechanism, greatly improves the resistance and self-protection ability of system to external threats.
[0053] The file fingerprint information and version control mechanism of the application realize effective tracking and efficient management of Flink job file, especially in the file version update process, through legality verification to ensure the safety and consistency of operation, solve the problem of data synchronization and version conflict in traditional method, optimize the overall performance and user experience of system. BRIEF DESCRIPTION OF DRAWINGS
[0054] Figure 1 It is a flowchart of the persistent Flink job file loading and display method of the embodiment of the application;
[0055] Figure 2 It is a performance comparison analysis column chart of biological intelligent immune blockchain technology of the embodiment of the application;
[0056] Figure 3 It is a comparison radar chart of immune response efficiency under different file operation types of the embodiment of the application;
[0057] Figure 4 It is a logic block diagram of chaotic neural network key generation of the embodiment of the application;
[0058] Figure 5 It is a performance comparison analysis column chart of biological intelligent immune authentication and traditional scheme of the embodiment of the application. DETAILED DESCRIPTION
[0059] In order to make the purpose, technical scheme and advantages of the embodiment of the application clearer, the technical scheme in the embodiment of the application will be described clearly and completely in combination with the drawings in the embodiment of the application, obviously, the described embodiment is only a part of the embodiment of the application, not all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor belong to the scope of protection of the application.
[0060] The technical solutions of the present application will be described in detail below with specific examples. The following specific examples can be combined with each other, and the same or similar concepts or processes may not be described in detail in some examples.
[0061] Figure 1 A flowchart of a persistent Flink job file loading and display method according to an embodiment of the present application is shown in FIG. 1, which includes the following steps: Figure 1
[0062] Obtaining the file identifier, version marker and metadata information of the Flink job file, and performing sharded storage on the Flink job file based on the file identifier, the version marker and the metadata information to generate a file shard set and file fingerprint information;
[0063] Based on the file shard set and the file fingerprint information, a bio-intelligent inspired blockchain consensus mechanism is constructed, which identifies abnormal file operations through immune memory cells and establishes a dynamic trust score model of file operation behavior based on the principle of antigen-antibody reaction;
[0064] Based on the abnormal file operation and the dynamic trust score model, a blockchain node network is deployed and a dynamic encryption key is generated, which is used to establish a secure communication channel in the blockchain node network for encrypted transmission and storage of file shards;
[0065] When a file access request is received, identity authentication and permission verification are performed based on the dynamic trust score model and the secure communication channel, and when a suspicious operation is detected, an immune response mechanism is triggered to isolate the suspicious node;
[0066] When a file version is updated, the legality of the update operation is verified using the file fingerprint information and the dynamic trust score model, and after verification, the file shard information of the new version is stored in the blockchain node network through the secure communication channel.
[0067] In an optional implementation, based on the file shard set and the file fingerprint information, a bio-intelligent inspired blockchain consensus mechanism is constructed, which identifies abnormal file operations through immune memory cells and establishes a dynamic trust score model of file operation behavior based on the principle of antigen-antibody reaction, which includes:
[0068] Based on the file fragment set and the file fingerprint information, the file fragment set is divided into multiple feature subsets, and an immune memory cell model is constructed according to feature hash values and correlation degrees in the feature subsets, the immune memory cell model including a feature recognition layer, a pattern matching layer and a response output layer, wherein the feature recognition layer extracts behavior features of file operations based on the feature hash values, the pattern matching layer establishes a feature similarity calculation matrix based on the correlation degrees, and the response output layer generates an immune response result according to the feature similarity calculation matrix;
[0069] Fusion is performed on file operation information and the behavior features to generate an antigen feature vector, and a feature pattern of the pattern matching layer is represented as an antibody feature vector; the binding strength of the antigen feature vector and the antibody feature vector is calculated based on the feature similarity calculation matrix, and the binding strength is obtained by calculating the cosine similarity between feature vectors; the immune response result and the binding strength are combined by weighting to obtain an immune response coefficient, and the immune response coefficient is mapped to a dynamic trust score model.
[0070] After obtaining the Flink job file, the file is processed by fragmentation according to the file size, and the size of each fragment is usually 4MB. For example, a 120MB Flink job file is divided into 30 fragments to form a file fragment set. SHA-256 hash values are calculated for each fragment to generate 256-bit feature fingerprints. At the same time, file metadata information is extracted, including creation time, modification time, owner, permission and other attributes. The hash values of all fragments are connected and hashed again to obtain the overall file fingerprint information, such as "7a8b9c0d1e2f3g4h...". These fingerprint information and metadata serve as the unique identifier of the file, which is used for subsequent operation recognition and verification.
[0071] Based on the file fragment set and the file fingerprint information, a feature clustering method is used to divide the file fragment set into multiple feature subsets. In a specific implementation, the K-means clustering algorithm is used, and K=5 is set, i.e. the 30 fragments are divided into 5 feature subsets, each containing semantically similar fragments. For example, subset 1 contains fragments of file headers and configuration information, subset 2 contains fragments of main code logic, and subset 3 contains fragments of data processing flow, etc. For each feature subset, the hash similarity between the internal fragments and the hash difference between the external and other subset fragments are calculated to obtain the cohesion and separation degree indicators of the feature subset. For example, the cohesion of subset 1 is 0.82, and the average separation degree with other subsets is 0.67.
[0072] The immune memory cell model is constructed based on the feature hash values and the correlation degree in the feature subset. The immune memory cell model consists of three layers: feature recognition layer, pattern matching layer, and response output layer. The feature recognition layer contains 128 neurons, each responsible for identifying a specific type of file operation feature. For example, neuron #42 is specifically designed to identify the content replacement pattern in file modification operations, and neuron #87 is responsible for identifying the sequential access pattern in file read operations. The feature hash values are organized through a Merkle tree structure, with the leaf nodes corresponding to the hash values of the file fragments and the root node being the overall file fingerprint. When a file operation occurs, the affected fragment hash values change, and the change location can be quickly located through the Merkle tree, and the relevant hash information is input into the feature recognition layer. The feature recognition layer uses a convolutional neural network to extract high-level features from the hash values, such as an 84-dimensional vector representing the operation type, impact range, and time characteristics from the hash sequence "7a8b9c0d...".
[0073] The pattern matching layer contains 256 neurons and adopts a self-organizing mapping network structure to map the feature vectors output by the feature recognition layer to a two-dimensional feature space. Similar operation patterns are closer in the feature space, and different operation patterns are farther apart. A feature similarity calculation matrix is constructed based on the correlation degree between file fragments, with a size of 256x256, and each element represents the similarity between two features. The correlation degree is calculated by analyzing the co-occurrence frequency, modification order, and dependency relationship of the file fragments. For example, the modification of fragment 3 and fragment 7 often occurs simultaneously, and their correlation degree is 0.86; while fragment 2 and fragment 19 are almost unrelated, with a correlation degree of only 0.12. The feature similarity calculation matrix is smoothed using a Gaussian radial basis function to ensure that similar features produce similar activation patterns. The pattern matching layer outputs a 256x256 feature similarity matrix, reflecting the matching degree of the current operation with known patterns.
[0074] The response output layer contains 64 neurons, which receive the feature similarity matrix from the pattern matching layer and generate the final immune response result through a weighted voting mechanism. The response result includes operation type identification (such as read, write, delete), anomaly degree evaluation (range 0-1), and recommended response strategy (such as allow, record, verify, or block). For example, for a normal file read operation, the response result is {operation type: "read", anomaly degree: 0.05, response strategy: "allow"}; while for a suspicious file overwrite operation, the response result is {operation type: "overwrite", anomaly degree: 0.78, response strategy: "verify"}.
[0075] When the file operation information is fused with the behavior features, the file operation information includes operation type, timestamp, execution user, access permission, and other metadata. For example, the user "admin" performs a "write" operation at "2025-06-19 14:30:25" with a permission of "rw-r--r--". These information is fused with the behavior features (such as operation frequency, access pattern, content change amplitude) extracted by the feature identification layer to generate a 128-dimensional antigen feature vector. In the antigen feature vector, the first 64 dimensions represent operation metadata features, and the last 64 dimensions represent behavior features. For example, the value 0.75 in the 10th dimension of the antigen feature vector represents the regularity score of the operation time, and the value 0.92 in the 75th dimension represents the locality score of the content change.
[0076] When the feature patterns of the pattern matching layer are represented as antibody feature vectors, the principal feature directions are extracted from the feature similarity matrix of the pattern matching layer to obtain antibody feature vectors with the same dimension as the antigen feature vectors. The antibody feature vectors represent the normal operation patterns learned by the system. For example, the antibody feature vector corresponding to a normal code update operation has higher values (0.8-0.95) in the "modification time regularity" and "content change locality" dimensions, and lower values (0.1-0.3) in the "access sensitive area" dimension.
[0077] The binding strength between the antigen feature vector and the antibody feature vector is calculated based on the feature similarity calculation matrix. In specific implementation, the cosine similarity between the two vectors is calculated by dividing the dot product of the two vectors by the product of their respective module lengths. For example, the cosine similarity between the antigen feature vector of the current file modification operation and the antibody feature vector of the normal modification pattern is 0.92, indicating a high match; while the cosine similarity between the antigen feature vector and the antibody feature vector of the abnormal modification pattern is only 0.31, indicating a low match. The binding strength is in the range of [0, 1], and the larger the value, the more consistent the current operation with the known normal pattern.
[0078] The immune response result and the binding strength are combined by weighting to obtain the immune response coefficient. The weighting method is: immune response coefficient = abnormal degree of immune response result x 0.6 + (1 - binding strength) x 0.4. For example, the abnormal degree of immune response result is 0.25, and the binding strength is 0.85, then the immune response coefficient is 0.25 x 0.6 + (1 - 0.85) x 0.4 = 0.15 + 0.06 = 0.21. The immune response coefficient is in the range of [0, 1], and the larger the value, the more abnormal the operation.
[0079] When mapping the immune response coefficient to the dynamic trust score model, a nonlinear mapping function is used to transform the immune response coefficient into a trust score of 0-100. The specific mapping method is: trust score = 100 x (1- immune response coefficient squared). For example, when the immune response coefficient is 0.21, the trust score is 100 x (1-0.21 2 )=100 x (1-0.044)=95.6. The dynamic trust score model also considers operation history, user reputation and environmental factors to dynamically adjust the basic trust score. For example, when it is detected that the user has performed similar abnormal operations multiple times in a short period of time, the trust score of each operation will decrease by 5-10 points; when the operation occurs at an irregular time (such as 3am), the trust score decreases by 15-20 points; when the operation comes from an infrequently used IP address or device, the trust score decreases by 10-15 points.
[0080] The dynamic trust score model is integrated with the blockchain consensus mechanism as an important basis for decision-making by the verification node. For example, operations with a trust score higher than 80 can be directly verified by consensus; operations with a trust score between 50 and 80 need to be confirmed by additional verification nodes; operations with a trust score lower than 50 will trigger a risk alert and be rejected. When the file operation is confirmed by the verification node, the relevant information is packaged into a block and propagated and confirmed in the network through the consensus algorithm. The verified operation record and trust score are permanently stored on the blockchain, forming an unalterable operation audit log.
[0081] The bio-intelligent heuristic-based blockchain consensus mechanism realizes accurate identification of abnormal file operations through an immune memory cell model and establishes a dynamic trust score model through the principle of antigen-antibody reaction, providing reliable protection for the secure access and management of Flink job files. This method combines the advantages of artificial neural networks, immune systems and blockchain technology, with adaptive learning, distributed verification and tamper resistance, and is suitable for file security protection in large-scale distributed computing environments.
[0082] Figure 2 For the performance comparison and analysis of the bio-intelligent immune blockchain technology of the embodiments of the present application, the column chart is as follows:
[0083] The bar chart shows the comparison data of three different blockchain security solutions in three key performance indicators. Among them, the "traditional blockchain method" (white column) represents the blockchain security solution based on the conventional consensus mechanism, the "AI enhanced blockchain method" (slash column) represents the blockchain solution combined with artificial intelligence technology, and the "bio-intelligent immune blockchain" (grid column) represents the innovative blockchain technology integrating bio-intelligence and immune mechanism. In terms of abnormal operation identification accuracy, the bio-intelligent immune blockchain solution achieved the highest level of 95.6%, significantly exceeding the AI-enhanced solution of 78.2% and the traditional solution of 61.5%; in the key indicator of consensus establishment speed, the bio-intelligent immune solution also performed outstandingly, reaching 92.4%, while the AI-enhanced solution was 74.1% and the traditional solution was only 57.6%; in the comprehensive evaluation of system security protection level, the bio-intelligent immune solution achieved a remarkable result of 98.8%, which had a significant advantage over the AI-enhanced solution of 82.3% and the traditional solution of 65.3%. The overall data shows that the innovative blockchain technology integrating bio-intelligence and immune mechanism has achieved a qualitative leap in various performance indicators, with an average improvement of more than 30 percentage points over the traditional solution and 15 percentage points over the AI-enhanced solution, fully proving the innovative value and application potential of the technology solution.
[0084] In an optional implementation, the feature recognition layer extracts the behavior features of the file operation based on the feature hash value, the pattern matching layer establishes a feature similarity calculation matrix based on the correlation degree, and the response output layer generates an immune response result according to the feature similarity calculation matrix, including:
[0085] The feature recognition layer inputs the feature hash value to a prefrontal pattern calculation unit, the prefrontal pattern calculation unit extracts the behavior features of the file operation based on the neuron spontaneous firing mechanism, and generates an initial feature vector;
[0086] The initial feature vector is input to a dynamic adaptive network, the dynamic adaptive network performs enhancement processing on the initial feature vector based on the synaptic plasticity mechanism, and generates an enhanced feature vector, the enhanced feature vector inherits the time sequence features and spatial features of the initial feature vector and has dynamic adaptive capability;
[0087] A feature similarity calculation matrix is constructed based on the enhanced feature vector, an antigen-antibody dynamic balance mechanism is introduced into the feature similarity calculation matrix, the feature similarity calculation matrix is input to an immune memory unit, the immune memory unit optimizes the feature similarity calculation matrix based on the immune tolerance mechanism, and generates an optimized feature pattern;
[0088] Based on the optimized feature mode, a bee colony coordination network is constructed, the bee colony coordination network including a plurality of coordination computing nodes, and adjacent coordination computing nodes exchanging feature information through a pheromone transmission mechanism; a group decision result is generated based on the feature information exchange result in the bee colony coordination network, and a final immune response result is generated according to a comparison result of the group decision result and a preset immune tolerance threshold.
[0089] When the feature hash value is input to the prefrontal pattern calculation unit in the feature recognition layer, the feature hash value is generated by using an SHA-256 algorithm, has a length of 256 bits, and is represented by 64 hexadecimal characters. For example, the hash value of the file reading operation is "8f4e7d6c5b4a3210...". The prefrontal pattern calculation unit is composed of 128 artificial neurons and arranged as a two-dimensional array of 8x16. Each neuron has a resting potential of -70 mV, a threshold potential of -55 mV, and a refractory period of 2 ms. The neuron spontaneous firing mechanism is realized by a random background current, and the background current obeys a Gaussian distribution with a mean of 4 pA and a standard deviation of 1 pA. The feature hash value is divided into 32 8-bit groups, and each group is mapped to the input current of the corresponding neuron, and the mapping rule is: input current (pA) = basic current 5 pA + hash group value x 0.1. For example, when the hash group value is 143, the corresponding neuron receives an input current of 19.3 pA. The neuron generates an action potential according to the input current and its own state to form a spatiotemporal pulse pattern. The discharge sequence of the neuron array within 500 ms is recorded as a pulse matrix P, where P[i,j]=1 indicates that neuron i fires at time j, and otherwise 0. By analyzing the spatiotemporal pattern in the pulse matrix P, features including firing rate (20-50 times per second), firing pattern (regular, clustered or burst type), interneuron synchrony (0.2-0.7) are extracted to form an initial feature vector V of 84 dimensions init .
[0090] When the initial feature vector is input to the dynamic adaptation network for enhancement processing, the dynamic adaptation network adopts a three-layer structure: 84 neurons in the input layer, 256 neurons in the hidden layer, and 128 neurons in the output layer. The network implements synaptic plasticity mechanisms, including short-term synaptic plasticity (STP) and long-term synaptic plasticity (LTP / LTD). STP simulates short-term changes in synaptic transmission efficiency, with a time constant of 50-200 ms; LTP / LTD simulates long-term synaptic enhancement or inhibition, with a duration of more than 30 minutes. In the implementation, when frequent reading of the same file (more than 10 times within 5 seconds) is detected, the relevant synaptic weight is enhanced by 20%; when an abnormally large file write operation (more than 3 times the standard deviation) is detected, the relevant synaptic weight is inhibited by 30%. Synaptic weight W[i,j] represents the connection strength from neuron j to neuron i, with an initial value randomly generated in the [0.2, 0.8] interval. The weight update adopts the time difference rule: when the presynaptic neuron j fires t milliseconds before the postsynaptic neuron i, the weight changes positively (enhancement); otherwise, it changes negatively (inhibition). For example, when the time difference t = 10 ms, the weight is enhanced by 8%; when t = -15 ms, the weight is inhibited by 12%. After synaptic plasticity adjustment, the network shows enhanced response to normal operation modes (such as sequential reading during working hours) and inhibited response to abnormal modes (such as random access at night). The output layer generates a 128-dimensional enhanced feature vector V enh , which inherits the timing features (operation sequence, interval) and spatial features (access location, range) of the initial feature vector and has the ability to dynamically adapt to environmental changes.
[0091] When constructing the feature similarity calculation matrix based on the enhanced feature vectors, a 128x128 similarity matrix S is used to represent the correlation strength between features. The matrix element S[i,j] represents the similarity between feature i and feature j, with a value range of [0, 1]. When calculating the similarity, the cosine distance, Euclidean distance, and Hamming distance of the feature vectors are considered, with weights of 0.5, 0.3, and 0.2, respectively. The antigen-antibody dynamic balance mechanism is introduced into the similarity matrix, with normal operation features considered as "self" and abnormal operation features considered as "non-self". A dynamic threshold θ (initial value 0.75) is set, and when S[i,j] > θ, it is determined as "self", otherwise as "non-self". θ is dynamically adjusted with environmental changes: for every 10% increase in system load, θ decreases by 0.02; for every 5 additional active users, θ decreases by 0.01. When the feature similarity calculation matrix is input into the immune memory unit for optimization, the immune memory unit distinguishes between "danger" and "safe" signals based on the immune tolerance mechanism. Danger signals include: a sharp increase in file access frequency (>200%) within a short period of time (1 minute), sensitive file area access, and abnormal time period operation (e.g., 3 a.m.). Safe signals include: operation during regular working hours, operation consistent with historical access patterns, and authorized user's regular behavior. The immune memory unit maintains a memory bank M, which stores historical feature patterns and their danger scores. For newly input features, their matching degrees with each pattern in the memory bank are calculated, and the danger scores are adjusted based on the matching results. For example, a certain operation feature has a highest matching degree of 0.87 with a dangerous pattern in the memory bank, and an initial danger score of 0.82; after 10 similar operations, the danger score is reduced to 0.45 as no actual harm is found, achieving immune tolerance. The immune memory unit outputs the optimized feature pattern P opt , which contains 128 features and their adjusted weights and danger scores.
[0092] When constructing the swarm coordination network based on the optimized feature patterns, the network contains 16 coordination computing nodes, each of which is responsible for processing a sub-region of the feature space. The nodes are arranged in a 4x4 grid, and adjacent nodes exchange feature information through the pheromone transmission mechanism. The pheromone intensity I[i,j] represents the information transmission intensity from node i to node j, and the initial value is 0.5. When node i detects abnormal features, it releases pheromones to adjacent node j, and I[i,j] increases by 0.2; when no abnormalities are detected for 3 consecutive times, the pheromone intensity decays, and I[i,j] decreases by 0.1. After node j receives the pheromones, it increases the sensitivity to related features, and the detection threshold decreases by 15%. For example, node 2 detects the abnormal frequency of file deletion operations (50 files deleted in 10 minutes), and releases pheromones to adjacent nodes 3, 6, and 7. These nodes subsequently increase the monitoring sensitivity to file deletion operations. Each node in the swarm network makes independent decisions and votes, and the final group decision result is generated through majority voting or weighted averaging. When a certain feature is determined to be abnormal by more than 60% of the nodes, a global alarm is triggered; when it is determined to be abnormal by 30%-60% of the nodes, it enters an observation state; and when it is below 30%, it is considered to be normal fluctuations. The group decision result is compared with the preset immune tolerance threshold (default 0.65) to generate the final immune response result. The immune response is divided into four levels: ignore (score <0.4), observe (0.4≤score<0.65), warn (0.65≤score<0.85), and block (score≥0.85). For example, an operation with a score of 0.72 triggers a warning response, and the system records detailed logs and sends a notification to the administrator; an operation with a score of 0.93 triggers a blocking response, and immediately terminates the related process and locks the involved account.
[0093] The above technical implementation provides a biologically inspired feature recognition and immune response mechanism, which can effectively identify abnormal behavior in file operations and take appropriate measures according to the threat level. This method combines the advantages of neural computing, immune systems, and swarm intelligence, with characteristics of adaptive learning, distributed decision-making, and strong fault tolerance, making it suitable for file security protection in complex network environments.
[0094] Figure 3 The radar chart for comparing the immune response efficiency of different file operation types in the embodiments of the present application is as follows:
[0095] The radar chart shows the security protection efficiency comparison of three different technical solutions in six file operation scenarios. Among them, "the technical solution" (solid dot) represents the immune defense algorithm based on the fusion of biological intelligence and machine consciousness, "AI enhanced immune method" (block solid line) represents the immune algorithm introducing deep learning and neural network, and "traditional immune method" (dashed line plus) represents the conventional immune defense method based on rule matching. In terms of file reading operation, the biological intelligence fusion scheme achieves the highest efficiency of 92.7%, the AI enhanced scheme is 80.4%, and the traditional scheme is only 60.2%; in file writing operation, the three schemes achieve 88.5%, 76.3% and 55.7% respectively; the protection efficiency of file deletion operation is 85.6%, 73.1% and 52.3% respectively; the performance of file copy operation is 83.2%, 69.8% and 49.6% respectively; the efficiency of file moving operation is 86.3%, 71.2% and 53.5% respectively; and the protection efficiency of file permission modification operation is 89.2%, 77.9% and 58.1% respectively. From the overall data, the technical solution based on the fusion of biological intelligence and machine consciousness is significantly better than the other two schemes in all six dimensions, with a protection efficiency generally 10-15 percentage points higher than the AI enhanced scheme and 30-35 percentage points higher than the traditional scheme, fully proving the technical innovation and practical value of the biological intelligence fusion scheme in the field of file operation security protection. The shape of the radar chart also intuitively shows that the technical solution has a larger and more balanced coverage range, reflecting its all-round protection capability.
[0096] In an optional implementation, based on the abnormal file operation and the dynamic trust scoring model, deploying a blockchain node network and generating a dynamic encryption key comprises:
[0097] Obtaining a feature matrix of the abnormal file operation, the feature matrix including operation type, abnormal score, timestamp and operation location information, calculating a dynamic trust score according to the feature matrix, and constructing the dynamic trust score into a multi-dimensional feature tensor;
[0098] Based on the operation location information in the feature matrix and the multi-dimensional feature tensor, calculating the node deployment density of each operation location according to the abnormal score corresponding to each operation location and the component value of the multi-dimensional feature tensor, the node deployment density being obtained by weighted combination of abnormal score weight coefficient and feature tensor weight coefficient;
[0099] The multi-dimensional feature tensor is dimensionally reconstructed to obtain a key seed matrix, and the multi-dimensional feature tensor is shape-transformed in a preset dimension, which is determined by the node deployment density; the key seed matrix is input into a chaotic mapping model for nonlinear transformation, and the control parameters of the chaotic mapping model are dynamically adjusted according to the dynamic trust score, to obtain an initial key sequence;
[0100] The information entropy value of the initial key sequence is calculated, the information entropy value is obtained by calculating the logarithmic product of the probability of occurrence of each character in the initial key sequence, the iteration number of the chaotic mapping model is adjusted according to the information entropy value, and a dynamic encryption key satisfying a preset entropy value threshold is generated.
[0101] The feature matrix of the abnormal file operation is obtained to perform blockchain node network deployment and dynamic encryption key generation. The feature matrix includes the following four key dimensions: operation type (such as deletion, modification, copying, etc.), abnormal score (a numerical value of 0-100 representing the degree of abnormality), timestamp (accurate to the millisecond level), and operation location information (represented by device IP and file path). An example of the feature matrix collected by the system is as follows: the feature of a certain file deletion operation is [deletion operation, 85-point abnormal score, 20230615143022-millisecond timestamp, 192.168.1.100 / documents / financial / ]. Based on the feature matrix, the system calculates a dynamic trust score by a designed algorithm, and the score is obtained by comprehensively considering the operation type risk weight (deletion operation weight 0.9), the abnormal score standardized value (85 / 100=0.85), and the time correlation factor (calculated as 0.78 according to the time interval) to obtain a score value between 0 and 1, such as 0.81. These score values are constructed into a multi-dimensional feature tensor, for example, a tensor with a shape of 3×4×5, wherein each dimension corresponds to different feature categories, time periods, and location groups.
[0102] The operation location information in the feature matrix and the multi-dimensional feature tensor are used to calculate the node deployment density. For each operation location, such as 192.168.1.100 / documents / financial / , the system extracts the corresponding abnormal score of 85 points and the corresponding component value of 0.81 in the feature tensor. By setting the abnormal score weight coefficient α=0.6 and the feature tensor weight coefficient β=0.4, the system calculates the node deployment density as 0.6×(85 / 100)+0.4×0.81≈0.83. This density determines the number of blockchain nodes that need to be deployed in this location, for example, in the area with a density of 0.83, the system will deploy 8 blockchain nodes to form a highly concentrated security monitoring network.
[0103] To generate a high-security dynamic encryption key, the system first reconstructs the multi-dimensional feature tensor in dimensions to obtain a key seed matrix. Specifically, if the original feature tensor has a shape of 3x4x5 and the node deployment density is 0.83, the system will determine a preset dimension transformation scheme according to the density value, such as expanding the tensor along the second dimension into a 12x5 matrix form. In this way, the value 0.75 at the (1,2,3) position in the original tensor will be mapped to the (5,3) position of the new matrix. This key seed matrix contains the key feature information of the original abnormal file operation, but has been reconstructed to improve security.
[0104] The system inputs the obtained key seed matrix into a chaotic mapping model for nonlinear transformation. Using Logistic chaotic mapping, the control parameter r is adjusted according to the dynamic trust score 0.81, calculated as r = 3.7 + 0.3 x 0.81 ≈ 3.94. The initial value is set to the normalized average value 0.67 of the key seed matrix, and the chaotic sequence is calculated by iteration, such as [0.67, 0.87, 0.45, 0.98,...]. This sequence is quantized and mapped to a 256-bit binary array, and the initial key sequence "A5F83C7D..." is finally obtained.
[0105] The system calculates the information entropy value of the initial key sequence to evaluate its randomness and security strength. For the generated initial key sequence "A5F83C7D...", the system calculates the probability of each character appearing, such as the probability of 'A' appearing is 0.05 and the probability of '5' appearing is 0.04, etc. By calculating the negative value of the logarithmic product of each character probability, the information entropy value is about 4.85 bits / character. The system's preset entropy threshold is 4.9 bits / character, and since the current entropy value does not reach the threshold, the system automatically adjusts the number of iterations of the chaotic mapping model from the initial 100 to 150, and regenerates the key sequence. After adjustment, the information entropy value of the newly generated key sequence "F7D2B9E3..." is 4.93 bits / character, meeting the preset threshold requirement, and the system adopts this sequence as the final dynamic encryption key.
[0106] This dynamic encryption key is distributed to the deployed blockchain node network for encrypting transaction records of sensitive file operations. Each node obtains different segments of the key according to its own trust score, ensuring that even if a single node is compromised, the complete key will not be leaked. The system further implements a time window-based key update mechanism on each node, which re-triggers the key generation process when a new abnormal file operation is detected, ensuring the dynamic nature and continuous security of the encryption system. Through this method, the system realizes the deployment of a blockchain node network based on abnormal file operations and a dynamic trust score model, and generates a dynamic encryption key, providing a highly secure protection mechanism for the file system.
[0107] In an alternative embodiment, the key seed matrix is input into a chaotic map model for nonlinear transformation, control parameters of the chaotic map model are dynamically adjusted according to the dynamic trust score, and an initial key sequence is obtained, including:
[0108] The key seed matrix is converted into neuron input potential values, each of which corresponds to a neuron position, a neuromorphic chaotic computing unit is constructed based on the neuron input potential values, and the neuromorphic chaotic computing unit includes membrane potential and synaptic weight;
[0109] The firing threshold and the synaptic weight of the neuromorphic chaotic computing unit are adjusted according to the dynamic trust score, and a spiking neural network is constructed based on the adjusted firing threshold and synaptic weight, the spiking neural network having spatiotemporal dynamics;
[0110] The state information of the spiking neural network is input into the neuromorphic chaotic computing unit, the spatiotemporal evolution process of the spiking neural network is calculated through the neuromorphic chaotic computing unit, and the output signal of the spiking neural network is decoded to obtain a chaotic sequence, the chaotic sequence inheriting the dynamics of the spiking neural network;
[0111] The environmental noise information generated by the motion of micro-particles is obtained, the environmental noise information is input into the neuromorphic chaotic computing unit for processing, the processed noise information is fused with the chaotic sequence, and a chaotic sequence with enhanced randomness is obtained;
[0112] Phase transition information in the motion of micro-particles is detected, parameter adjustment of the chaotic sequence is triggered based on the phase transition information, the chaotic sequence with enhanced randomness is updated according to the adjusted parameters, and an initial key sequence is generated.
[0113] As shown in Figure 4 the method further comprises:
[0114] The key seed matrix is a two-dimensional matrix of 256x256, with element values ranging from 0 to 255. The conversion uses a linear mapping function to map the element value v to the neuron potential space [-70, -50] millivolts: potential value = -70 + (v / 255) x 20. For example, when the element value at position (35, 127) in the matrix is 128, the converted neuron potential value is -60 millivolts; when the element value at position (142, 89) is 200, the potential value is -54.12 millivolts. The matrix is divided into 16x16 blocks, and the average value of each block is calculated as the input potential of the corresponding neuron. For example, the average value of block (2, 5) is 175, so the input potential of the neuron corresponding to this block is -55.9 millivolts. After conversion, a 16x16 neuron array is formed, each neuron has an initial input potential value between [-70, -50] millivolts. During the potential conversion process, floating-point precision is used, with two decimal places, to ensure the continuity and smoothness of the potential distribution.
[0115] Each neuron simulates the electrophysiological characteristics of biological neurons, including membrane potential, threshold potential, refractory period and ion channel dynamics. The initial value of the membrane potential is set to the corresponding input potential value, such as the initial membrane potential of neuron (3, 7) being -58.2 millivolts. Ion channels include sodium channels, potassium channels and leakage current channels, with the following dynamic parameters: sodium channel activation time constant 3 milliseconds, inactivation time constant 1 millisecond; potassium channel activation time constant 10 milliseconds; leakage current conductance 0.02 microsiemens. The synaptic weight matrix W represents the connection strength between neurons, with a size of 256x256, and is initially generated using a truncated Gaussian distribution: mean 0.5, standard deviation 0.1, value range [0.2, 0.8]. For example, W
[42]
[153] =0.67 represents the synaptic weight from neuron #153 to neuron #42 as 0.67. The connection probability between neurons is 0.3, and a Bernoulli distribution is used to generate a connection mask matrix M. Combined with W and M, the actual synaptic connection matrix S = W o M is obtained, where "o" represents element-wise multiplication. The synaptic transmission delay matrix D between neurons has element values between 1-5 milliseconds, uniformly randomly distributed. For example, D
[42]
[153] =2.7 represents that the signal from neuron #153 reaches neuron #42 after 2.7 milliseconds.
[0116] The dynamic trust score ranges from [0, 100], and the neuron firing threshold and synaptic weights are dynamically adjusted based on the score. The firing threshold adjustment formula is: Threshold = -55 + (Trust Score - 50) × 0.1. When the trust score is 85, the firing threshold is -51.5 mV; when the trust score is 35, the firing threshold is -56.5 mV. Synaptic weight adjustment uses a scaling method: New weight matrix = Original weight matrix × (0.5 + Trust Score / 200). For example, when the trust score drops from 85 to 65, the synaptic weights are adjusted from the [0.2, 0.8] interval to the [0.16, 0.66] interval. Neuron excitability also adjusts with the trust score: when the trust score is below 50, the neuron excitability parameter is increased by 0.2, causing the network to fire more frequently; when the trust score is above 80, the excitability parameter is decreased by 0.1, making network activity more regular. Synaptic conduction delays were also fine-tuned: for every 20-point change in trust score, the delay matrix shifted by ±0.5 milliseconds overall. After parameter adjustments, all neurons were reset to a resting state, ready to enter the spiking neural network construction phase.
[0117] Based on the adjusted firing threshold and synaptic weights, a spiking neural network with 256 neurons was constructed. The network topology adopted a small-world network structure, with each neuron establishing connections with an average of 15 other neurons. The nearest neighbor connection probability was 0.8, and the far neighbor connection probability was 0.1. The neurons were arranged in a 16×16 two-dimensional grid, with neuron (i,j) corresponding to the (i×16+j)th neuron in the one-dimensional index. 85% of the synapses in the network were excitatory (positive weights), and 15% were inhibitory (negative weights). The initial membrane potential of each neuron was set to the resting potential plus ±5 mV of random perturbation, introducing diversity in the initial state. The network refractory period was set to 2 milliseconds, meaning that no new action potentials would be generated within 2 milliseconds after a neuron fires. After network construction, a 50-millisecond warm-up phase was initially run to allow the network to reach a dynamic equilibrium state and eliminate the influence of the initial conditions. The output of the warm-up phase was not included in the subsequent key generation process.
[0118] A discrete time step of 0.1 ms was used, with a total simulation duration of 500 ms, resulting in 5000 time steps. At each time step t, the state of all 256 neurons was updated. For neuron i, the update of its membrane potential V[i] considered three parts: natural membrane potential decay (typically 15 ms), external input current, and synaptic input. External input current I... exti Set to a constant background current of 5 picoamperes plus ±2 picoamperes of Gaussian noise. Synaptic input I syni Calculate the effect of all presynaptic neurons j: Check whether neuron j fires at time tD[i][j]. If so, add S[i][j] × synaptic current amplitude (10 picoamperes) to I. syni V was calculated.i After that, the threshold V th [i] is compared with the threshold V i [i] at time t. If V th [i], the neuron i fires at time t, records the firing event, and resets V i [i] to -75 mV; if V th [i], V i is kept unchanged. For example, at time step t = 125.7 ms, the calculation process of neuron #89 is as follows: the original membrane potential is -62.3 mV, the natural decay is -63.1 mV, the external input contribution is +0.7 mV, the synaptic input contribution is +4.2 mV, the new membrane potential is -58.2 mV, the threshold is -55.0 mV, and the judgment is not to fire. After the calculation of all 5000 time steps is completed, a 256x5000 pulse matrix is obtained, in which the element value 1 represents firing and 0 represents not firing.
[0119] For each neuron i, the firing time sequence [t1, t2...t n ] in the 500 ms simulation is extracted. The time intervals △t = [t2-t1, t3-t2...t n -t (n-1) ] between adjacent firings are calculated. The time intervals △t are mapped to integer values in the range of 0-255: mapped value = (△t%25.6)x10, where "%" represents the modulo operation. For example, the firing times of neuron #125 are [78.2, 103.5, 142.8, 168.3] ms, the corresponding time intervals are [25.3, 39.3, 25.5] ms, and the mapped values are [253, 137, 255]. Repeat this process for all 256 neurons to obtain a chaotic sequence with a length of about 8192. The chaotic sequence has high unpredictability, and the statistical characteristic analysis results are as follows: entropy value 7.997 / 8, autocorrelation coefficient r <0.01 (delay >2), approximate entropy value 0.9992, linear complexity >4000, binomial distribution test p value 0.478, and run test p value 0.523.
[0120] The environmental noise sources consist of two parts: a CMOS thermal noise sensor and a quantum random number generator. The CMOS sensor, operating at a constant temperature of 25±0.1℃, collects noise generated by the thermal motion of electrons at a sampling rate of 50kHz and a resolution of 16 bits. The quantum random number generator, based on a single-photon detector, captures quantum fluctuation signals and generates a true random bit stream at a rate of 1Mbps. The raw data for the two noise sources are [-0.0023, 0.0017, -0.0009...] and [1, 0, 1, 1, 0, 0, 1...], respectively. The CMOS noise data is preprocessed: high-pass filtering (cutoff frequency 1kHz) removes low-frequency drift, and quantization is performed to floating-point numbers in the range [-1.0, 1.0]. Von Neumann extraction is performed on the quantum bit stream to eliminate potential biases. The processed noise data are then fused in a 1:1 ratio to generate a comprehensive environmental noise sequence. This noise sequence is injected into the neuronal membrane potential through a gain factor of 0.5 mV: V i += 0.5 × noise value. After noise injection, the neural network exhibits richer dynamic behaviors, such as bifurcation, intermittent chaos, and synchronization-desynchronization transitions. The processed noise information is fused with the aforementioned chaotic sequence through a bit-level XOR operation: new sequence[i] = chaotic sequence[i] ⊕ (noise sequence[i] × 255), where “⊕” represents bit-by-bit XOR. The fused sequence passed the NISTSP800-22 randomness test, passing all 15 tests, and the p-values were evenly distributed.
[0121] The system is equipped with a high-precision temperature sensor (accuracy ±0.01℃), a barometric pressure sensor (accuracy ±0.01kPa), and a humidity sensor (accuracy ±0.1%), with a sampling frequency of 10Hz. Sensor data is acquired via a 16-bit ADC to record real-time changes in microscopic environmental parameters. The phase transition detection algorithm employs multi-scale trend analysis to identify parameter abrupt changes. When a temperature change rate exceeds 0.1℃ / second, a pressure change rate exceeds 0.1kPa / second, or a humidity change rate exceeds 1% / second, it is determined to be a microscopic phase transition event. For example, when the temperature rapidly rises from 25.00℃ to 25.15℃, the system captures a phase transition event. Based on the phase transition information, chaotic sequence parameter adjustments are triggered. The adjustments include: overall neuron threshold shift ±(0.1×phase transition amplitude) millivolts; synaptic weight matrix rotation angle = (phase transition time × 360°)%90°; and neuron dynamics parameter change amplitude = ±(phase transition amplitude × 0.05). For example, when a temperature jump of 0.15°C is detected, the neuron threshold shifts by -0.015 mV, the synaptic weight matrix rotates by 42°, and the neuron dynamics parameter changes by -0.0075. This parameter adjustment causes the chaotic system to enter a new dynamic trajectory, resulting in no statistical correlation between consecutively generated key sequences. During a 500-millisecond simulation, an average of 3-5 microscopic phase transition events were captured, each triggering a parameter adjustment.
[0122] According to the adjusted parameters, the neuromorphic chaotic computing unit 500 is re-run for 500 milliseconds, and the neuron firing pattern is recorded. A sliding window method is used to extract a new pulse sequence every 100 milliseconds, seamlessly connecting to the previous sequence. The boundary region (50 milliseconds) of the new and old sequences is smoothed to ensure that the sequence transition does not produce statistical anomalies. The time interval features in the new sequence are extracted and mapped to integer values in the range of 0-255. The newly generated value sequence is bit-level fused with the environmental noise sequence to obtain an updated random chaos sequence. The sequence updating process is implemented in a pipeline on the hardware level, with the new sequence generation and the old sequence usage executed in parallel to ensure continuous availability of key material. The update frequency is 10 times per second, and each update generates about 1024 bits of new key material. The final generated initial key sequence has a length of 8192 bits, which can be regarded as 32 256-bit sub-key blocks, meeting the key requirements of the AES-256 algorithm.
[0123] The generated initial key sequence has undergone strict cryptographic evaluation, with an entropy detection result of 7.999 / 8, close to an ideal random source; all NIST SP800-22 tests are passed with a minimum p-value of 0.237; the linear complexity is 4183, much higher than half the sequence length; no repeating patterns are detected in the periodicity analysis; the differential analysis test shows that a 1-bit change in input results in a 50.02% change in output; the time sequence analysis test shows that even if the first 4096 bits of the key are known, the success rate of predicting the next 4096 bits is only 50.1%, almost equivalent to random guessing; the key recovery attack complexity analysis shows that using the most advanced quantum computing technology, the cracking complexity still reaches 2 128 , far beyond the practical attack capability. When implemented on an Inteli7-9700K processor, the 8192-bit key generation takes about 85 milliseconds, meeting the real-time encryption system requirements.
[0124] In an optional implementation, identity authentication and permission verification are performed based on the dynamic trust score model and the secure communication channel, and when a suspicious operation is detected, an immune response mechanism is triggered to isolate the suspicious node, including:
[0125] The basic communication channel information and identity authentication information are input into a life encoder, which converts the basic communication channel information into entangled state information; the entangled state information and trust score information are input into a gene entanglement unit based on the secure communication channel, and the entangled state information is biologically encoded based on the gene entanglement unit to generate a gene authentication result;
[0126] The genetic authentication result and the secure communication channel are input into a life key defense unit, and the life key defense unit generates defense channel information based on a biological key distribution mechanism, and the defense channel information inherits the biological features of the genetic authentication result;
[0127] The trust score information, the genetic authentication result, and the defense channel information are input into a permission evaluation unit, and the permission evaluation unit calculates a comprehensive permission score according to a preset weight coefficient, compares the comprehensive permission score with a dynamic threshold, and generates a permission verification result;
[0128] A deviation value between a current operation and a normal operation baseline is detected, and when the deviation value exceeds a preset immune threshold, an immune response mechanism is triggered, and the immune response mechanism performs biological isolation on a suspicious node based on the permission verification result, and the biological isolation realizes immune shielding of the suspicious node through the defense channel information.
[0129] When a user requests to access a Flink job file, an identity authentication and permission verification process starts from a life encoder. The life encoder receives two types of input information: basic communication channel information and identity authentication information. The basic communication channel information includes a source IP address "192.168.1.100", a MAC address "00:1A:2B:3C:4D:5E", a communication protocol version "IPv6", a port number "8080", and a timestamp "1718784645". The identity authentication information includes a user ID "user_12345", an encrypted password hash value "$2a$10$d8zk2OG5YVJ...", a security token "eyJhbGciOiJIUzI1NiIs...", and a fingerprint feature parameter "FP:A72B83C5". The life encoder internally implements a DNA encoding conversion module that converts digital information into ATCG base sequences using a quaternary encoding mapping table. For example, the IP address "192.168.1.100" is first converted to binary "11000000.10101000.00000001.01100100", and then each two bits are mapped to a base (00--A, 01--T, 10--C, 11--G) to obtain "GCTAGAGAAAGTAT". Similarly, other information is also converted into corresponding base sequences. The life encoder also includes an RNA transcription unit that splices and transcribes base sequences from different information sources according to priorities to generate an initial RNA sequence "GCTAGAGAAAGTATCGATCG...". Finally, through the base complementary pairing rule (A-T, C-G) and RNA splicing technology, 128-bit entangled state information "0xA7B2C3D4E5F60718293A4B5C6D7E8F90" is generated.
[0130] After the entangled state information is generated, the system transmits it to the genetic entanglement unit through a secure communication channel along with the trust score information. The trust score information is generated by a score generation module that analyzes three dimensions: user historical behavior score, current session security level, and operation type risk coefficient. The historical behavior score is based on the user's operation records in the past 30 days, taking into account operation frequency, resource consumption, and the number of abnormalities, and is scored as 85 points. The current session security level considers connection duration, network stability, and environmental factors, and is scored as 90 points. The operation type risk coefficient is classified according to the degree of influence of the operation on the system, with query operations being low-risk and the coefficient being 2. The combined trust score information is represented as "85_90_2". The genetic entanglement unit first performs sequence segmentation processing, dividing the entangled state information "0xA7B2C3D4E5F60718293A4B5C6D7E8F90" into 8 segments of 16 bits each. Then it performs cross-recombination, converting the trust score information "85_90_2" into the base sequence "ATCGCTAT" and inserting it between the segments. Next, it performs mutation processing, calculating the mutation probability "0.023" based on the current system load "45%" and network delay "12ms", and randomly mutating approximately 2.3% of the bits in the sequence. Finally, through gene expression regulation, the processed sequence is amplified to 512 bits to generate the genetic authentication result "0xB8C9DAEBFC0D1E2F3A4B5C6D7E8F9019283A4B5C6D7E8F9...". This result contains three layers of information: information content, structural features, and variation patterns.
[0131] After the genetic authentication result is generated, the system inputs it into the life key defense unit along with secure communication channel information. The secure communication channel information includes the encryption algorithm type "AES-256", the key length "256 bits", the communication protocol "TLSv1.3", the session identifier "SID:F7A3E9D1", and the channel establishment time "2025-06-19T10:30:45Z". The life key defense unit implements a biological key distribution mechanism that simulates the signal transmission process between biological cells. First, the high-entropy region "0xB8C9DAEB" is extracted from the genetic authentication result as a seed. Then, combined with the current timestamp "1718784645" and the random entropy source "0xD7E8F901", a temporary key pair is generated through an elliptic curve cryptography algorithm. The temporary public key "0xA4B5C6D7..." is broadcast to authorized nodes through a secure channel, and the temporary private key is kept locally. Next, using the principle of quantum key distribution, a unique identifier and session key are assigned to each communication node, forming a key distribution tree. The system fuses the feature sequence of the genetic authentication result with the node identifier to generate a biological key "0xF1E2D3C4B5A6978..." that is adapted to a specific node. This key has a time limit, with a default validity period of 30 minutes, and the system automatically rotates the key material every 5 minutes to ensure forward security. Finally, the life key defense unit outputs 1024-bit defense channel information, including the biological key, access control policy, and secure communication parameters.
[0132] After the defense channel is established, the system inputs the trust score information "85_90_2", the check value "0xB8C9" of the genetic authentication result, and the integrity marker "0xF1E2" of the defense channel information into the permission evaluation unit. The permission evaluation unit contains a multi-dimensional evaluation model that calculates the permissions of the input data. First, the trust score information is standardized, converting "85_90_2" to a standard score "0.85_0.90_0.20". Then, the validity of the genetic authentication result is verified, and the integrity score "0.92" of the check value "0xB8C9" is calculated. Next, the security level of the defense channel information is checked, and according to the integrity marker "0xF1E2", it is rated as "high security" with a corresponding score of "0.95". The permission evaluation unit assigns weight coefficients of 0.4, 0.35, and 0.25 to the three dimensions respectively, and calculates the comprehensive permission score as (0.85×0.4)+(0.92×0.35)+(0.95×0.25)=0.895, i.e. 89.5 points. The dynamic threshold of the system is dynamically adjusted by the security monitoring center according to the current security situation, and the current value is 75 points. Comparing the comprehensive permission score of 89.5 points with the dynamic threshold of 75 points, the permission verification result of "pass" is generated, and the verification details "AUTH:user_12345:PASS:89.5:75" are recorded.
[0133] The system continuously detects deviations of user operations from the normal operation baseline through the operation monitoring module. The normal operation baseline is generated by the data statistics module, which analyzes the operation records of all users in the past 90 days, including daily operation frequency distribution "5-15 times / day", single operation duration "30-180 seconds", data access volume "0.5-5 MB / time", access pattern "metadata area accounts for 70%, configuration area accounts for 20%, data area accounts for 10%" and other characteristics. The system performs real-time feature extraction on the current user operation to obtain the feature vector "[15, 25, 8, 0.3, 0.6, 0.1]", which represents the operation frequency of 15 times / 2 minutes, the average operation time of 25 seconds, the data access volume of 8 MB, and the access proportion of metadata, configuration and data area. The normal baseline vector is "[5, 60, 3, 0.7, 0.2, 0.1]". The system calculates the weighted Euclidean distance of the two vectors to obtain the deviation value 25.7, which exceeds the preset immune threshold 20.0.
[0134] When the deviation value exceeds the immune threshold, the system triggers the immune response mechanism. This mechanism consists of a three-level defense system. The first level is passive defense, the system records the abnormal event "EVENT:ANOMALY:user_12345:25.7:20.0" and pushes an alarm to the security log center. The second level is active defense, the system executes temporary restriction measures, reduces user operation permissions, reduces trust score from 85 to 60, limits access frequency to a maximum of 3 times every 5 minutes, and prohibits access to the configuration area. The third level is immune isolation, for high-risk operations or failed permission verification, the system performs biological isolation. The biological isolation process first generates the isolation instruction "ISOLATE:192.168.1.107:HIGH:300", indicating that the node with IP "192.168.1.107" is isolated with high priority for 300 seconds. The isolation executor receives the instruction and immediately interrupts all communication sessions with the node, revokes the temporary key in the secure communication channel, and adds the node identifier to the blockchain blacklist. At the same time, the system broadcasts the isolation notification "ALERT:NODE:192.168.1.107:ISOLATED:HIGH:300" to other nodes in the network to prevent the node from accessing system resources through other paths.
[0135] The system creates an isolation record in the isolation state record table, including the node identifier "192.168.1.107", the user ID "user_12345", the isolation reason "abnormal operation frequency & data access limit exceeded", the isolation level "high", the isolation start time "2025-06-19T10:35:12Z", the isolation duration "300 seconds", and the release condition "administrator review or time expiration". During the isolation period, the system performs a node state check every 60 seconds, recording the node's network activity and the number of attempted connections. After the isolation time ends, the system performs a release isolation process, first verifying the node's security, then recalculating the trust score. If the score is higher than the recovery threshold of 50 points, the communication rights are gradually restored; if the score is lower than the threshold, the isolation time is extended to 600 seconds and the isolation level is increased. For nodes that have triggered isolation multiple times, the system implements a progressive punishment mechanism, doubling the isolation time each time, up to a maximum of 24 hours, and permanently blacklisting after three times.
[0136] Through the above detailed implementation, the system realizes a biological heuristic-based identity authentication and permission verification mechanism, which can accurately identify abnormal operations and take corresponding immune response measures, effectively ensuring the security and integrity of Flink job files. This implementation is suitable for distributed computing environments with high security requirements and has strong practical value.
[0137] Figure 5 The bar chart for performance comparison between the biological intelligent immune authentication of the embodiment of the present application and the traditional scheme is as follows:
[0138] The graph shows the comparison data of three different security defense schemes (traditional encryption authentication scheme, blockchain security scheme and biological intelligent immune scheme) in five key performance indicators. In terms of identity authentication success rate, the three schemes achieve 86.5%, 92.4% and 97.8% respectively, and the biological intelligent scheme has obvious advantages; in terms of abnormal operation detection rate, the traditional scheme is 72.3%, the blockchain scheme is improved to 83.7%, and the biological intelligent scheme is as high as 94.5%; in terms of immune isolation response time, the traditional scheme is only 65.8%, the blockchain scheme is improved to 80.2%, and the biological intelligent scheme leads by a large margin to 92.3%; in terms of security protection coverage, the three schemes are 78.2%, 85.9% and 95.1% respectively; in terms of overall system security, the traditional scheme is 75.6%, the blockchain scheme is improved to 84.3%, and the biological intelligent scheme is as high as 96.2%. Overall, the biological intelligent immune scheme is significantly better than the other two schemes in all five key indicators, fully embodying the technical innovation value and application potential of the scheme in the field of security defense.
[0139] In an alternative embodiment, the biological encoding of the entangled state information based on the gene entanglement unit includes:
[0140] The entangled state information contains the association characteristics between the communication nodes. The gene entanglement unit converts the association characteristics into a gene sequence based on a DNA sequence mapping mechanism. The gene sequence contains the biological characteristics of the entangled state information.
[0141] Based on the gene sequence, an authentication template is constructed, and the authentication template is dynamically encoded by a gene recombination mechanism to generate a gene authentication result. The gene authentication result inherits the biological characteristics of the gene sequence.
[0142] The association characteristics between the communication nodes contained in the entangled state information refer to the association state characteristic data formed between the quantum bits in the quantum communication system. These association characteristics can be represented as a set of bit strings, such as "010110", which describes the quantum entanglement relationship between communication node A and communication node B. In a specific example, the association characteristics can be a set of 32-bit binary strings, such as "01011010110101101011010110101101", where each bit represents the measurement result of a specific quantum state.
[0143] When the above entangled state information is input into the gene entanglement unit, the system will start the DNA sequence mapping mechanism. This mechanism uses base correspondence encoding method to convert binary association characteristics into DNA four-base sequence. The specific mapping rule is: binary "00" is mapped to adenine (A), "01" is mapped to cytosine (C), "10" is mapped to guanine (G), and "11" is mapped to thymine (T). Taking the above 32-bit binary string as an example, the converted DNA sequence is "CATGCATGCATGCATG". This DNA sequence contains the biological characteristics of the original entangled state information and has the stability and diversity characteristics of DNA molecules.
[0144] After the conversion is completed, the system constructs an authentication template based on the DNA sequence. The authentication template construction process involves structuring the arrangement of the DNA sequence according to biological rules. For example, the sequence of 16 bases "CATGCATGCATGCATG" is arranged in a 4x4 matrix to form a two-dimensional structure template. This template serves as the basic data structure for gene authentication and contains the spatial distribution characteristics of the original entangled state information.
[0145] Next, the system dynamically encodes the authentication template through a genetic recombination mechanism. The genetic recombination mechanism simulates the biological DNA recombination process, including cutting, exchanging, and connecting operations. In this embodiment, the genetic recombination follows the following steps: segment the DNA sequence of the authentication template into groups of 4 bases each, obtaining sequence fragments "CATG", "CATG", "CATG", "CATG"; apply a preset recombination rule, such as a cyclic shift rule, to each sequence segment, shifting it 1 bit to the right, obtaining "GCAT", "GCAT", "GCAT", "GCAT"; and finally, re-connect the recombined fragments to form a new DNA sequence "GCATGCATGCATGCAT".
[0146] The recombined DNA sequence is further enhanced in security through secondary encoding. This step uses complementary base replacement technology to replace some bases with their complementary bases (A-T complementary, C-G complementary). For example, replace the bases at even positions in the sequence "GCATGCATGCATGCAT" with their complements, obtaining the sequence "GAATGGATTCATTCAT". This sequence is the final genetic authentication result, which inherits the biological characteristics of the original gene sequence while having higher security and uniqueness.
[0147] In actual application, when the identity of a communication node needs to be verified, the receiving party will re-collect the entangled state information and generate a comparison sequence through the same biological encoding process. The validity of the identity is determined by calculating the similarity of the two DNA sequences. The similarity calculation uses the Hamming distance algorithm, which calculates the number of different bases at corresponding positions in the two sequences. For example, if the reference sequence is "GAATGGATTCATTCAT" and the sequence to be verified is "GAATGGATTCATTCAG", they have 1 different base at one position, and the similarity is (16-1) / 16=93.75%. The system has a preset threshold of 90%, so this verification is passed.
[0148] To improve the anti-interference ability of the system, the genetic entanglement unit also introduces an error correction mechanism. This mechanism is based on the Reed-Solomon encoding principle and adds redundant base information to the DNA sequence. For example, 4 check bases are appended to the 16-base sequence to form an extended sequence "GAATGGATTCATTCATACTG". This allows the system to recover the correct authentication result even if some base information is lost or incorrect.
[0149] In addition, to cope with quantum channel noise interference, the system implements a dynamic base substitution strategy. When the noise level exceeds the preset threshold, the system will automatically adjust the DNA sequence mapping rule, such as adjusting the original mapping rule "00--A, 01--C, 10--G, 11--T" to "00--C, 01--G, 10--T, 11--A", thereby improving the robustness of the encoding. In one test, when the channel noise increased from 5% to 15%, the authentication accuracy of the system was improved from 87% to 94% by dynamically adjusting the mapping rule.
[0150] The physical implementation of the genetic entanglement unit uses a dedicated integrated circuit, including a mapping module, a recombination module, and an encoding module. The mapping module handles the conversion of entangled state information to DNA sequences, the recombination module performs genetic recombination operations, and the encoding module generates the final authentication result. Through hardware acceleration, the system can complete the entire process from entangled state information collection to genetic authentication result generation in milliseconds, meeting the real-time communication authentication requirements.
[0151] The method further comprises:
[0152] In terms of Flink configuration file, a parameter web.upload.dir is added, which is used to specify the HDFS storage path of the service end to accept the job file uploaded by the client. All job files uploaded by users through Flink WebUI are uniformly saved to the subdirectory under the HDFS directory specified by the Flink configuration file (i.e. the value of the above parameter), and the subdirectory name is the MD5 value of the job file.
[0153] The web.upload.dir parameter does not directly conflict with the existing Flink configuration, and it is added as a new configuration item to supplement the default job file storage method of Flink (i.e. the local temporary directory of Flink JobManager). In terms of default value, to avoid problems caused by configuration omissions, the default value is set to / user / flink / jobs_uploads (or adjusted according to the actual deployment environment), but it is strongly recommended to explicitly configure this parameter in actual application, and the parameter value should be a valid HDFS path, and the path should already exist or the Flink process user should have the permission to create it. The path should not contain special characters or spaces.
[0154] The loading time of the web.upload.dir parameter, in the Flink JobManager and TaskManager startup process, the Flink configuration file (such as flink-conf.yaml) is loaded, which includes the web.upload.dir parameter. The JobManager reads this parameter when starting the WebUI service to initialize the storage path of the job file, to determine the HDFS path where the job file should be stored. Flink verifies the validity of the web.upload.dir parameter during startup or when receiving job file uploads, including checking whether the path exists, whether it is writable, etc. If the verification fails, an error message is recorded to the process output log, and the corresponding error prompt information is prompted in the job file upload interface.
[0155] In the Flink WebUI client, when the user sends the job file to the Flink JobManager by clicking the newly added Submit New Job To HDFS button, the JobManager first calculates the MD5 value of the entire JAR file, and then calls the HDFS API to determine whether a directory with the MD5 value exists under the above-defined directory on HDFS.
[0156] Among them, when establishing a connection with HDFS, the FileSystem.get(new Configuration()) API is used to obtain the HDFS file system instance. If HDFS is configured with Kerberos authentication, the Kerberos parameter value needs to be set in the Flink configuration file in advance and the corresponding Kerberos ticket is used for identity verification. If HDFS is configured with superuser permissions, ensure that the user running the JobManager has appropriate permissions to access HDFS. Check and create the storage directory: call the HDFS API to determine whether a directory with the MD5 value exists under the specified HDFS directory (specified by the web.upload.dir parameter in the Flink configuration file). If it does not exist, use the mkdirs() method to create the directory.
[0157] If the above directory does not exist, create a directory with the above MD5 value under the directory defined by the Flink specified parameter, then save the job file to this directory and return the MD5 value of the current file to the front end as the unique identifier of the job. Among them, the copyFromLocalFile() API of HDFS is used to copy the local job file to the newly created directory on HDFS.
[0158] If an MD5 collision occurs when uploading a job file, the MD5 value of the entire file + the current timestamp in milliseconds is used as the directory name and job unique identifier.
[0159] After successful saving, the MD5 value of the current file is returned to the front end as the unique identifier of the job.
[0160] When the user clicks the job start button, the Flink JobManager first obtains the job file stored in the above directory on HDFS according to the MD5 value and the newly added save directory parameter value passed by the page through the HDFS API. If the file is mistakenly deleted, it will prompt that the file has been deleted. If it exists, the job file is grabbed to the system temporary directory to facilitate subsequent filling of complete job submission commands.
[0161] Add a download button to the job list. When the user clicks download, the Flink JobManager first obtains the job file stored in the above directory on HDFS according to the MD5 value and the newly added save directory parameter value passed by the page through the HDFS API. If the file is mistakenly deleted, it will prompt that the file has been deleted. If it exists, the job file is transmitted to the user's browser in the form of a file stream.
[0162] For the monitoring page, modify the uploaded job name indicator value to MD5 value + task submission timestamp to facilitate users to distinguish monitoring indicators through the MD5 value and job submission time displayed in the job list.
[0163] Example 1: Single job deployment
[0164] I. Configuration parameter settings
[0165] · Flink configuration file modification: In the Flink configuration file (such as flink-conf.yaml), add the parameter web.upload.dir and set its value to a valid path on HDFS, for example:
[0166] web.upload.dir: hdfs: / / namenode:8020 / user / flink / jobs_uploads.
[0167] · HDFS connection configuration: Ensure that the Flink cluster can connect to HDFS and configure the necessary Kerberos authentication parameters (if HDFS enables Kerberos authentication).
[0168] II. User uploads job files
[0169] User operation: User selects a local job file (e.g. JAR package) and uploads it through the job upload interface of Flink WebUI.
[0170] System processing:
[0171] 1. MD5 calculation: After receiving the upload request, Flink JobManager calculates the MD5 value of the job file.
[0172] 2. Directory check and creation: According to the value of the web.upload.dir parameter, JobManager calls the HDFS API to check whether a directory named after the MD5 value exists on HDFS. If not, create the directory.
[0173] 3. File saving: Save the job file under the newly created MD5 directory, and return the MD5 value to the front end as the unique identifier of the job.
[0174] III. Job file loading and starting
[0175] User operation: User finds the uploaded job in the job list of Flink WebUI and clicks the "Start" button.
[0176] System processing:
[0177] 1. File acquisition: Flink JobManager acquires the job file remotely through the HDFS API according to the MD5 value of the job and the value of the web.upload.dir parameter.
[0178] 2. Temporary storage: Grab the job file to the system temporary directory for subsequent filling of complete job submission commands.
[0179] 3. Job starting: Execute the job submission command to start the Flink job.
[0180] IV. Final display
[0181] Job list display: In the job list of Flink WebUI, display the uploaded job, including job name, MD5 value, submission time, etc.
[0182] Monitoring page display: In the monitoring page, display the running state and indicators of the job, and the job name is displayed as MD5 value + task submission timestamp, which is convenient for users to distinguish.
[0183] Example two: JobManager failover
[0184] I. Fault occurrence
[0185] • Scenario description: The JobManager in the Flink cluster stops running due to downtime, and needs to be failover.
[0186] II. Failover process
[0187] New JobManager startup: During the failover process, a new JobManager instance is started.
[0188] Configuration synchronization: The new JobManager obtains the configuration information of the Flink cluster from the configuration center or shared storage, including the web.upload.dir parameter value.
[0189] Job file recovery:
[0190] 1. Directory traversal: The new JobManager traverses the job file storage directory on HDFS according to the web.upload.dir parameter value.
[0191] 2. Job information loading: Load the information of all job files, including MD5 value, upload time, etc.
[0192] III. Job recovery and user operation: The user logs in to the Flink WebUI and views the list of recoverable jobs.
[0193] System processing:
[0194] 1. Job selection: The user selects the job to be recovered and clicks the "Recover" button.
[0195] 2. File acquisition: The new JobManager acquires the job file remotely through the HDFS API according to the MD5 value of the job.
[0196] 3. Temporary storage and startup: The job file is fetched to the system temporary directory, and the complete job submission command is filled in, and the Flink job is started.
[0197] IV. Final display
[0198] Job list display: In the job list of Flink WebUI, the recovered job is displayed, and the original job name, MD5 value, etc. information is retained.
[0199] • Monitoring page display: In the monitoring page, the running state, indicators, etc. of the recovered job are displayed, and the job name display method is the same as in Example One.
[0200] Example Three: Cluster expansion
[0201] I. Expansion preparation
[0202] • Scenario description: The Flink cluster needs to be scaled out by adding new nodes to increase processing capacity.
[0203] II. Configuration synchronization and node joining
[0204] • Configuration synchronization: Ensure that the newly added node has consistent configuration information with other nodes in the Flink cluster, including the value of the web.upload.dir parameter.
[0205] • Node joining: Add the new node to the Flink cluster and make it part of the cluster.
[0206] III. Job migration and loading
[0207] User operation: Users select to migrate part of the job to the new node according to their needs.
[0208] System processing:
[0209] 1. Job selection: Users select the job to be migrated in the job list of Flink WebUI.
[0210] 2. File acquisition: Flink JobManager acquires job files remotely through HDFS API based on the MD5 value of the job.
[0211] 3. Temporary storage and submission: Grab the job file to the system temporary directory of the new node and submit it to the new node for running.
[0212] IV. Final display
[0213] • Job list display: In the job list of Flink WebUI, display the migrated job and retain the original job name, MD5 value, etc. Users can view and manage these jobs on the new node as needed.
[0214] Monitoring page display: In the monitoring page, display the running status, indicators, etc. of the migrated job on the new node, and the job name display method is the same as in the first embodiment.
[0215] In the second aspect of the embodiment of the application, a persistent Flink job file loading and display system is provided, comprising:
[0216] A first unit is configured to obtain a file identifier, a version marker and metadata information of a Flink job file, perform sharded storage on the Flink job file based on the file identifier, the version marker and the metadata information, and generate a file shard set and file fingerprint information.
[0217] a second unit configured to construct a bio-intelligent inspired blockchain consensus mechanism based on the set of file fragments and the file fingerprint information, the blockchain consensus mechanism identifying abnormal file operations through immune memory cells and establishing a dynamic trust score model of file operation behavior based on an antigen-antibody reaction principle;
[0218] a third unit configured to deploy a blockchain node network and generate a dynamic encryption key based on the abnormal file operation and the dynamic trust score model, the dynamic encryption key being used to establish a secure communication channel in the blockchain node network, the secure communication channel being used to implement encrypted transmission and storage of file fragments;
[0219] a fourth unit configured to perform identity authentication and permission verification based on the dynamic trust score model and the secure communication channel when a file access request is received, and trigger an immune response mechanism to isolate a suspicious node when a suspicious operation is detected;
[0220] a fifth unit configured to perform legality verification on an update operation using the file fingerprint information and the dynamic trust score model when a file version is updated, and store file fragment information of a new version to the blockchain node network through the secure communication channel after the verification is passed.
[0221] In a third aspect of the embodiments of the present application, an electronic device is provided, comprising:
[0222] a processor;
[0223] a memory for storing processor-executable instructions;
[0224] The processor is configured to invoke the instructions stored in the memory to execute the method described above.
[0225] In a fourth aspect of the embodiments of the present application, a computer readable storage medium is provided, which stores computer program instructions, and the computer program instructions are executed by a processor to implement the method described above.
[0226] The present application can be a method, device, system and / or computer program product. The computer program product can include a computer readable storage medium having computer readable program instructions thereon, which are used to implement various aspects of the present application.
[0227] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit the present application; although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions recorded in the above embodiments can be modified, or some or all of the technical features can be replaced by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for loading and displaying persistent Flink job files, characterized in that, include: Obtain the file identifier, version tag, and metadata information of the Flink job file. Based on the file identifier, version tag, and metadata information, perform shard storage on the Flink job file to generate a file shard set and file fingerprint information. Based on the file fragment set and the file fingerprint information, a blockchain consensus mechanism inspired by biological intelligence is constructed. The blockchain consensus mechanism identifies abnormal file operations through immune memory cells and establishes a dynamic trust scoring model for file operation behavior based on the antigen-antibody reaction principle. Based on the aforementioned abnormal file operations and the dynamic trust scoring model, a blockchain node network is deployed and a dynamic encryption key is generated. This dynamic encryption key is used to establish a secure communication channel within the blockchain node network. The secure communication channel is used to implement encrypted transmission and storage of file fragments, including: Obtain the feature matrix of abnormal file operations. The feature matrix includes operation type, abnormal score, timestamp and operation location information. Calculate the dynamic trust score based on the feature matrix and construct the dynamic trust score as a multidimensional feature tensor. Based on the operation position information in the feature matrix and the multidimensional feature tensor, the node deployment density of the operation position is calculated according to the anomaly score and the component value of the multidimensional feature tensor corresponding to each operation position. The node deployment density is obtained by a weighted combination of the anomaly score weight coefficient and the feature tensor weight coefficient. The multidimensional feature tensor is reconstructed to obtain a key seed matrix. The multidimensional feature tensor is then transformed in shape along a preset dimension, which is determined by the node deployment density. The key seed matrix is then input into a chaotic mapping model for nonlinear transformation. The control parameters of the chaotic mapping model are dynamically adjusted according to the dynamic trust score to obtain an initial key sequence. The information entropy value of the initial key sequence is calculated. The information entropy value is obtained by calculating the logarithmic product of the probabilities of each character in the initial key sequence. The iteration number of the chaotic mapping model is adjusted according to the information entropy value to generate a dynamic encryption key that meets a preset entropy threshold. Upon receiving a file access request, identity authentication and permission verification are performed based on the dynamic trust scoring model and the secure communication channel. When a suspicious operation is detected, an immune response mechanism is triggered to isolate the suspicious node. When a file version is updated, the update operation is verified using the file fingerprint information and the dynamic trust scoring model. After verification, the file fragment information of the new version is stored in the blockchain node network through the secure communication channel.
2. The method according to claim 1, characterized in that, Based on the file fragment set and the file fingerprint information, a bio-inspired blockchain consensus mechanism is constructed. This consensus mechanism identifies abnormal file operations through immune memory cells and establishes a dynamic trust scoring model for file operation behavior based on the antigen-antibody reaction principle, including: Based on the file fragment set and the file fingerprint information, the file fragment set is divided into multiple feature subsets. An immune memory cell model is constructed based on the feature hash values and correlations in the feature subsets. The immune memory cell model includes a feature recognition layer, a pattern matching layer, and a response output layer. The feature recognition layer extracts behavioral features of file operations based on the feature hash values. The pattern matching layer establishes a feature similarity calculation matrix based on the correlations. The response output layer generates an immune response result based on the feature similarity calculation matrix. File operation information is fused with the behavioral features to generate an antigen feature vector, and the feature patterns of the pattern matching layer are represented as antibody feature vectors. The binding strength between the antigen feature vector and the antibody feature vector is calculated based on the feature similarity calculation matrix, and the binding strength is obtained by calculating the cosine similarity between the feature vectors. The immune response result and the binding strength are weighted and combined to obtain the immune response coefficient, and the immune response coefficient is mapped to a dynamic trust scoring model.
3. The method according to claim 2, characterized in that, The feature recognition layer extracts behavioral features of file operations based on the feature hash value, the pattern matching layer establishes a feature similarity calculation matrix based on the correlation, and the response output layer generates an immune response result based on the feature similarity calculation matrix, including: In the feature recognition layer, the feature hash value is input to the prefrontal pattern calculation unit, which extracts the behavioral features of file operations based on the spontaneous firing mechanism of neurons and generates an initial feature vector. The initial feature vector is input into a dynamic adaptive network, which enhances the initial feature vector based on the synaptic plasticity mechanism to generate an enhanced feature vector. The enhanced feature vector inherits the temporal and spatial features of the initial feature vector and has dynamic adaptability. A feature similarity calculation matrix is constructed based on the enhanced feature vector. An antigen-antibody dynamic balance mechanism is introduced into the feature similarity calculation matrix. The feature similarity calculation matrix is input into the immune memory unit. The immune memory unit optimizes the feature similarity calculation matrix based on the immune tolerance mechanism to generate an optimized feature pattern. A bee colony collaborative network is constructed based on the optimized feature patterns. The bee colony collaborative network includes multiple collaborative computing nodes, and adjacent collaborative computing nodes exchange feature information through a pheromone transfer mechanism. A group decision result is generated based on the feature information exchange results in the bee colony collaborative network, and a final immune response result is generated based on the comparison results of the group decision result and a preset immune tolerance threshold.
4. The method according to claim 1, characterized in that, The key seed matrix is input into a chaotic mapping model for nonlinear transformation. The control parameters of the chaotic mapping model are dynamically adjusted according to the dynamic trust score, resulting in an initial key sequence including: The key seed matrix is converted into neuron input potential values, each neuron input potential value corresponds to a neuron location, and a neuromorphic chaos computing unit is constructed based on the neuron input potential values. The neuromorphic chaos computing unit includes membrane potential and synaptic weights. The firing threshold and synaptic weights of the neuromorphic chaotic computing unit are adjusted according to the dynamic trust score. A spiking neural network is constructed based on the adjusted firing threshold and synaptic weights. The spiking neural network has spatiotemporal dynamic characteristics. The state information of the spiking neural network is input into the neuromorphic chaos computing unit, the time-series evolution process of the spiking neural network is calculated by the neuromorphic chaos computing unit, and the output signal of the spiking neural network is decoded to obtain a chaotic sequence, the chaotic sequence inheriting the dynamic characteristics of the spiking neural network; The environmental noise information generated by the motion of microscopic particles is acquired, and the environmental noise information is input into the neuromorphic chaotic computing unit for processing. The processed noise information is then fused with the chaotic sequence to obtain a chaotic sequence with enhanced randomness. The phase transition information in the motion of the microscopic particles is detected, and the parameters of the chaotic sequence are adjusted based on the phase transition information. The randomness-enhanced chaotic sequence is updated according to the adjusted parameters to generate an initial key sequence.
5. The method according to claim 1, characterized in that, Identity authentication and permission verification are performed based on the dynamic trust scoring model and the secure communication channel. When a suspicious operation is detected, an immune response mechanism is triggered to isolate the suspicious node, including: Basic communication channel information and identity authentication information are input into the life encoder, which converts the basic communication channel information into entangled state information. Based on the secure communication channel, the entangled state information and trust scoring information are input into the gene entanglement unit, which performs biological encoding on the entangled state information to generate a gene authentication result. The gene authentication result and the secure communication channel are input into the life key defense unit. The life key defense unit generates defense channel information based on the biological key distribution mechanism. The defense channel information inherits the biological characteristics of the gene authentication result. The trust score information, the gene authentication result, and the defense channel information are input into the permission evaluation unit. The permission evaluation unit calculates a comprehensive permission score based on a preset weight coefficient, compares the comprehensive permission score with a dynamic threshold, and generates a permission verification result. The deviation between the current operation and the normal operation baseline is detected. When the deviation exceeds a preset immune threshold, an immune response mechanism is triggered. The immune response mechanism performs biological isolation on suspicious nodes based on the permission verification result. The biological isolation achieves immune shielding of the suspicious nodes through the defense channel information.
6. The method according to claim 5, characterized in that, Based on the gene entanglement unit, the entangled state information is biologically encoded to generate gene authentication results, including: The entangled state information contains the association features between communication nodes. The entangled state information is input into the gene entanglement unit, which converts the association features into gene sequences based on a DNA sequence mapping mechanism. The gene sequences contain the biological features of the entangled state information. An authentication template is constructed based on the gene sequence. The authentication template is dynamically encoded through a gene recombination mechanism to generate a gene authentication result. The gene authentication result inherits the biological characteristics of the gene sequence.
7. A persistent Flink job file loading and display system, used to implement the method of any one of claims 1-6, characterized in that, include: The first unit is used to obtain the file identifier, version tag, and metadata information of the Flink job file, and perform fragmented storage on the Flink job file based on the file identifier, version tag, and metadata information to generate a file fragment set and file fingerprint information; The second unit is used to construct a bio-inspired blockchain consensus mechanism based on the file fragment set and the file fingerprint information. The blockchain consensus mechanism identifies abnormal file operations through immune memory cells and establishes a dynamic trust scoring model for file operation behavior based on the antigen-antibody reaction principle. The third unit is used to deploy a blockchain node network and generate a dynamic encryption key based on the abnormal file operation and the dynamic trust scoring model. The dynamic encryption key is used to establish a secure communication channel in the blockchain node network. The secure communication channel is used to realize the encrypted transmission and storage of file fragments. The fourth unit is used to perform identity authentication and permission verification based on the dynamic trust scoring model and the secure communication channel when a file access request is received. When a suspicious operation is detected, an immune response mechanism is triggered to isolate the suspicious node. The fifth unit is used to verify the legitimacy of the update operation by using the file fingerprint information and the dynamic trust scoring model when the file version is updated. After the verification is successful, the file fragment information of the new version is stored in the blockchain node network through the secure communication channel.
8. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 6.
Citation Information
Patent Citations
An intelligent management method and system for bidding information based on artificial intelligence
CN119743333A
Computer network security monitoring system and method
CN120090801A