A multi-modal AI fusion network information security real-time defense system and method

The real-time network information security defense system, which integrates multimodal AI, constructs an anti-self-oscillation verification closed loop by using disturbance commands and suppression signals. This solves the problems of verification feedback failure and system self-oscillation caused by disturbance signals mixing into background traffic, and achieves precise targeting and defense against network threats.

CN122226311APending Publication Date: 2026-06-16JIANGSU SANDBOX TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JIANGSU SANDBOX TECH CO LTD
Filing Date
2026-01-27
Publication Date
2026-06-16

Smart Images

  • Figure CN122226311A_ABST
    Figure CN122226311A_ABST
Patent Text Reader

Abstract

The application discloses a network information security real-time defense system and method based on multi-modal AI fusion, and relates to the technical field of network information security.The system generates a locking window and generates an inhibition signal while generating a disturbance instruction through a multi-modal cooperative controller, and constructs a controlled verification channel; the alignment unit accurately screens out target data matching time and protocol type from real-time multi-modal data during the validity period of the inhibition signal according to reference information; and the quantitative unit calculates the consistency and completeness characteristic values of the target data, and solves the confidence index.The application realizes causal locking and quantitative evaluation of cross-protocol layer feedback signals while preventing verification actions from triggering system false alarms, and achieves deterministic defense against fuzzy state threats.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of network information security technology, and in particular to a multimodal AI-integrated real-time network information security defense system and method. Background Technology

[0002] In the process of proactively verifying threats in ambiguous network states, defense systems need to inject perturbation signals into network connections to induce feedback. However, this mechanism faces a technical dilemma: the injected perturbation signals and the induced cross-layer feedback are mixed in with regular business traffic, making it impossible for the defense system's perception logic to distinguish between native anomalies and verification feedback. This signal confusion not only makes verification evidence difficult to extract due to spatiotemporal misalignment, but also causes the defense system to erroneously identify its own proactive verification behavior as an attack, thereby triggering logical self-excitation and failure of system functions. Summary of the Invention

[0003] This invention provides a multimodal AI-integrated real-time network information security defense system and method to solve the technical problems of verification feedback extraction failure and system logic self-oscillation caused by the mixing of background traffic with disturbance signals when the multimodal defense system performs active verification.

[0004] In view of the above problems, the present invention provides a multimodal AI fusion network information security real-time defense system, the system comprising: A multimodal cooperative controller is configured to: generate a disturbance command and reference information when it is determined that the network connection meets a preset triggering condition; and calculate the lock window duration and generate a suppression signal while generating the disturbance command; A disturbance injection unit, connected to the multimodal cooperative controller, is configured to send disturbance data packets to the network connection in response to the disturbance command; An alignment unit, connected to the multimodal cooperative controller, is configured to receive the reference information and, during the period of receiving the suppression signal, filter out data whose time difference and protocol type match the reference information from real-time collected network traffic data and application log data as target data and output it to the quantization unit; A quantization unit, connected to the alignment unit, is configured to receive the target data, calculate a consistency feature value characterizing the degree of consistency between the target data and the expected result of the disturbance command, and a completeness feature value characterizing the completeness of the target data, and calculate a confidence index based on the consistency feature value and the completeness feature value. The multimodal collaborative controller is also configured to receive the confidence index and generate a defense execution command when the index meets preset defense conditions.

[0005] This invention also provides a real-time network information security defense method based on multimodal AI fusion, comprising the following steps: When the network connection is determined to meet the preset triggering conditions, a disturbance command and reference information are generated, the lock window duration is calculated, and a suppression signal is generated. In response to the disturbance command, a disturbance data packet is sent to the network connection; During the period when the suppression signal is in effect, based on the reference information, data with matching time difference and protocol type are selected from real-time collected network traffic data and application log data as target data; Calculate a consistency feature value that characterizes the degree of consistency between the target data and the expected results, and a completeness feature value that characterizes the degree of completeness of the target data; A confidence index is calculated based on the consistency feature value and the integrity feature value, and a defense operation is performed when the index meets the preset defense conditions.

[0006] The technical solution provided in this application has at least the following technical effects: By generating a suppression signal while generating a disturbance command, and using reference information to perform causal alignment and quantization calculations on real-time acquired multimodal data, this invention constructs an anti-self-oscillation verification closed loop independent of conventional monitoring logic. This solution physically isolates secondary anomalies caused by verification traffic, eliminates the risk of false alarms, and ensures that in complex network timing environments, the system can accurately locate and parse cross-protocol layer verification feedback signals, achieving objective confirmation of network threats. Attached Figure Description

[0007] Figure 1 A schematic diagram of the structure of a multimodal AI fusion network information security real-time defense system provided in an embodiment of the present invention; Figure 2 This is a flowchart illustrating a real-time network information security defense method based on multimodal AI fusion, as provided in an embodiment of the present invention. Detailed Implementation

[0008] The above technical solutions will now be described in detail with reference to the accompanying drawings and specific embodiments to provide a better understanding of them. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. It should be understood that the present invention is not limited to the exemplary embodiments used only to explain the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention. Furthermore, it should be noted that, for ease of description, only the parts related to the present invention are shown in the drawings, not all of them.

[0009] Example 1 This embodiment provides a real-time network information security defense system that integrates multimodal AI. For example... Figure 1 As shown, the system mainly consists of a multimodal cooperative controller, a disturbance injection unit, an alignment unit, and a quantization unit. Real-time defense against network threats is achieved by constructing a verification closed loop to prevent self-oscillation.

[0010] Specifically, the multimodal collaborative controller, as the core hub of the system, is configured to execute the following logic: when the network connection meets preset triggering conditions, it generates a disturbance command and reference information; and while generating the disturbance command, it calculates the lock window duration and generates a suppression signal. A disturbance injection unit, connected to the multimodal collaborative controller, responds to the disturbance command and sends disturbance data packets to the network connection. An alignment unit receives the reference information and, while receiving the suppression signal, marks data whose time difference and protocol type match the reference information as target data and outputs it for real-time collected network traffic data and application log data, while discarding or temporarily storing unmatched data. A quantization unit, connected to the alignment unit, receives the target data, calculates a consistency feature value characterizing the degree of consistency between the target data and the expected result of the disturbance command, and a completeness feature value characterizing the completeness of the target data, and calculates a confidence index based on these two feature values. Finally, the multimodal collaborative controller receives the confidence index and generates a defense execution command when the index meets preset defense conditions.

[0011] Based on the above system architecture, this embodiment also provides a real-time network information security defense method using multimodal AI fusion. For example... Figure 2 As shown, the method includes the following steps: During the operation of the multimodal AI-integrated real-time network information security defense system, the multimodal collaborative controller continuously monitors the security status data of network connections. When the malicious confidence score of a network connection falls into a preset intermediate value range, the multimodal collaborative controller determines that the network connection is in an ambiguous state. For network connections in an ambiguous state, the multimodal collaborative controller initiates active verification logic, first generating a perturbation command. The perturbation command contains specific operation identifiers for the target Internet Protocol address and port number, used to instruct the perturbation injection unit to execute specific network protocol stack operations.

[0012] Along with the generation of the disturbance command, the multimodal cooperative controller synchronously generates fingerprint information for subsequent feedback matching. The fingerprint information consists of three core data fields: disturbance transmission timestamp, estimated round-trip time, and expected feedback type. The disturbance transmission timestamp records the system moment the disturbance command was issued; the estimated round-trip time is estimated based on historical transmission control protocol handshake records of the network connection; and the expected feedback type defines a specific protocol behavior pattern that has a logical causal relationship with the disturbance command.

[0013] To ensure the verification process does not trigger false alarms in the defense system, the multimodal cooperative controller calculates the lock window duration while generating disturbance commands. The lock window duration calculation relies on two dynamic parameters: the smoothed round-trip time (RTT) and the change in RTT. The RTT is iteratively updated using an exponentially weighted moving average algorithm; that is, the new RTT equals the old RTT multiplied by a preset retention coefficient plus the current measurement multiplied by a preset update coefficient. The change in RTT reflects the degree of fluctuation in network latency.

[0014] After obtaining the latest smoothed round-trip time and its change, the multimodal cooperative controller calculates the lock window duration using a weighted summation method. Specifically, the lock window duration equals the smoothed round-trip time plus four times the change in smoothed round-trip time. Setting the weight of the change in smoothed round-trip time to four times aims to cover the maximum random jitter that may occur during network transmission, ensuring that the verification feedback data falls entirely within the time range defined by the lock window duration.

[0015] After calculating the lock window duration, the multimodal collaborative controller generates a corresponding suppression signal and sends it to the alignment unit. The suppression signal triggers the alignment unit to enter the verification feedback data marking mode. In the verification feedback data marking mode, the alignment unit suspends the transmission of data to the regular threat detection module and instead performs feature scanning on the real-time collected network traffic data and application log data based on the received fingerprint information, thereby constructing an anti-self-oscillation verification channel independent of the regular monitoring process.

[0016] After receiving a disturbance command from the multimodal cooperative controller, the disturbance injection unit parses the operation identifier contained in the disturbance command. Depending on the operation identifier, the disturbance injection unit selects to perform different types of cross-protocol layer disturbance operations, such as network layer truncation or application layer forgery.

[0017] Implementation of network layer truncation perturbation strategy: When a perturbation command instructs the execution of the operation to discard the last Transmission Control Protocol (TCP) fragment in the Hypertext Transfer Protocol (HTTP) request message sequence, the perturbation injection unit first performs deep packet inspection on the current session data stream of the target network connection. By parsing the content length field in the HTTP header, the perturbation injection unit calculates the total data size of the current request message. Subsequently, based on the maximum segment length of the current network link, the perturbation injection unit calculates the total number of TCP fragments into which the request message is divided and the sequence number range of the last fragment.

[0018] During data transmission, the perturbation injection unit forwards all preceding data fragments normally until it detects a data packet with a sequence number matching the last fragment. At this point, the perturbation injection unit performs a drop action, preventing the last fragment from reaching the target server. Since the target server's application layer program relies on the content length field to determine whether the request has been fully received, the absence of the last fragment will cause the server's receive buffer to enter a waiting state. After the waiting time exceeds the server's preset read timeout threshold, the server's application layer program will throw a read timeout exception or log an error indicating an incomplete request body. This application layer timeout phenomenon, directly induced by network layer packet loss, constitutes a cross-layer causal feature for subsequent verification.

[0019] Implementation of application-layer forgery perturbation strategies: When a perturbation command instructs the execution of a Hypertext Transfer Protocol (HTTP) request packet containing a keep-alive header field and a reset flag, the perturbation injection unit constructs a special application-layer request message. This request message explicitly declares the connection attribute as keep-alive in its header field, conveying to the target server the intention to maintain the Transmission Control Protocol (TCP) connection state after the response is completed.

[0020] After the request message body is completely sent to the target server, the perturbation injection unit, without waiting for a server response, immediately constructs and sends a Transmission Control Protocol (TCP) data packet with a reset flag. The sending of the reset flag forcibly interrupts the current TCP connection. For servers implementing a standard protocol stack, when the application layer attempts to reuse the connection for a response write-back based on a keep-alive commitment, an input / output anomaly is triggered because the underlying connection has been broken. This leads the server-side operating system kernel to send a reset message or multiple retransmissions of the termination message. This protocol state conflict, created by the application layer's false commitment and the network layer's forced interruption, forces the target server to expose the true feedback of its protocol stack processing logic, providing deterministic observational evidence for subsequent intent quantification.

[0021] After receiving the suppression signal from the multimodal cooperative controller, the alignment unit enters the verification feedback data filtering state. In this state, the alignment unit performs a strict spatiotemporal causal matching operation on the real-time collected network traffic data and application log data based on the received reference information.

[0022] Anchoring of the disturbance transmission timestamp: To ensure the baseline accuracy of subsequent spatiotemporal distance calculations, the multimodal cooperative controller records a high-precision disturbance transmission timestamp at the physical moment the disturbance command is actually sent to the network interface card. The disturbance transmission timestamp is sampled using a system clock with microsecond-level precision and synchronized with the timestamp system of the network traffic acquisition probe. This disturbance transmission timestamp, as one of the core fields of the reference information, is transmitted to the alignment unit for storage, establishing the time origin of the causal verification process. Simultaneously, this reference information is uniquely bound to the five-tuple information of the target network connection. The five-tuple information includes the source Internet Protocol address, destination Internet Protocol address, source port, destination port, and protocol type, thereby ensuring the independence of the verification channel in a high-concurrency network environment.

[0023] Calculation and threshold determination of spatiotemporal distance: For each data record collected under the verification feedback data filtering state, the alignment unit first extracts the receiving timestamp of the data record. Then, the alignment unit performs a spatiotemporal distance calculation. The calculation logic is as follows: subtract the sum of the disturbance transmission timestamp and the estimated round-trip time from the receiving timestamp, and take the absolute value of this difference as the time difference.

[0024] After calculating the time difference, the alignment unit compares it with a preset time window threshold. The time window threshold is a tolerance range set based on the jitter statistics of the network environment, for example, set to twenty milliseconds. If the calculated time difference is less than the time window threshold, it indicates that the arrival time of the data record falls within the theoretical arrival time range of the feedback signal triggered by the disturbance command, conforming to causal timing logic. The time window threshold is introduced to cover random queuing delays and processing delay fluctuations on the network transmission path, preventing the missed detection of valid feedback signals due to network jitter.

[0025] Protocol type matching and target data tagging: Based on the time difference condition being met, the alignment unit further performs a protocol type matching check. The alignment unit parses the actual protocol type of the data record, for example, extracting Hypertext Transfer Protocol status codes or error codes from application logs, or Transmission Control Protocol flags from network traffic. The alignment unit then compares the parsed actual protocol type with the expected feedback type defined in the reference information.

[0026] The alignment unit determines a data record match is successful only when the time difference is less than the time window threshold and the actual protocol type is exactly the same as the expected feedback type. Successfully matched data records are marked as target data by the alignment unit and output to the quantization unit for further processing. For data records that fail to meet both the time difference and protocol type matching conditions simultaneously, the alignment unit discards or temporarily stores them during the suppression signal's active period, preventing them from being sent to the regular threat detection process, thus achieving strict data isolation between the verification and monitoring channels.

[0027] Vectorized computation of consistent eigenvalues: The quantization unit first maps the target data into a first vector. Specifically, the quantization unit uses a pre-trained intent feature extraction model to convert key fields in the target data, such as Hypertext Transfer Protocol status codes, error description text, or combinations of network layer flags, into numerical vectors in a high-dimensional space.

[0028] In a preferred embodiment, the intent feature extraction model can encode the log text using a pre-trained Word2Vec model or a BERT model. In another embodiment that does not rely on complex neural networks, the quantization unit uses one-hot encoding or a pre-defined hash map to directly map specific HTTP status codes (such as 404, 500) or TCP flags (such as RST, FIN) to preset standard basis vectors.

[0029] In a simplified implementation, the intent feature extraction model is configured as a pre-defined lookup table to directly map common Hypertext Transfer Protocol status codes or Transmission Control Protocol flags to predefined fixed vectors. Simultaneously, the quantization unit acquires a pre-defined second vector corresponding to the perturbation command. This second vector is a standard feature vector pre-encoded and stored during the system initialization phase for the expected feedback results of each perturbation strategy.

[0030] After acquiring the first and second vectors, the quantization unit calculates the distance or similarity between the two vectors. One specific implementation is to calculate the cosine distance, i.e., subtracting the cosine similarity of the two vectors from a given value. To map the calculation result to a standard interval, the quantization unit normalizes the distance value to obtain a consistency feature value. The magnitude of the consistency feature value reflects the degree of semantic consistency between the actual captured target data and the expected result. When the consistency feature value approaches zero, it indicates that the semantic features of the target data highly match the expected result, meaning the feedback signal conforms to the expected perturbation command.

[0031] Statistical calculation of completeness feature values: Parallel to vectorized computation, the quantization unit performs statistical calculations of the integrity feature values. The quantization unit counts the actual number of data packets or log entries contained in the target data. Simultaneously, the quantization unit queries the expected number of data packets defined in the reference information. The expected number of data packets refers to the number of standard responses that a single disturbance action should induce in an ideal network environment.

[0032] The quantization unit calculates the ratio of the actual number of data packets to the expected number of data packets, defining this ratio as the completeness feature value. The completeness feature value measures the completeness of the verification feedback evidence. When the completeness feature value is close to one, it indicates that the system has captured a complete verification feedback chain, and the evidence is highly credible; when the completeness feature value is much smaller than one, it indicates that the feedback signal is missing, which may reduce the reliability of subsequent decisions.

[0033] Nonlinear fusion calculation of confidence index: Based on the calculated consistency and completeness feature values, the quantization unit uses a non-linear weighted formula to calculate the confidence index. The formula for calculating the confidence index is: In this formula, Represents a confidence level indicator. Represents a consistent feature value. This represents the completeness feature value. The first term in the formula... The square of the consistency eigenvalue is introduced. This square operation nonlinearly amplifies the consistency deviation; that is, when the feedback signal deviates slightly from the expectation, this term increases rapidly, thus reflecting sensitivity to abnormal responses. The second term in the formula... A normalized exponential function based on completeness eigenvalues ​​is introduced. This function implements a non-linear penalty for missing evidence; that is, when the completeness eigenvalue deviates from the value of one, the value of this term increases exponentially. Parameters , and These are preset positive real coefficients used to adjust the weighting of different features on the final result. (Function) The hyperbolic tangent function is used to map the weighted summation result within the parentheses to the interval between zero and one. The lower the calculated confidence index value, the more consistent the verification feedback is with the expected result and the more complete the evidence, indicating a higher degree of confirmation of the attack behavior by the system. The coefficient... The sensitivity of the confidence index to data integrity was adjusted; a larger [indicator / measure] was used. The value will cause in When slightly missing The value rises rapidly, thereby achieving strict constraints on the integrity of verification evidence and ensuring that defense is only carried out against high-confidence threats in complex network environments, reflecting the system's extremely low false alarm rate.

[0034] After receiving the confidence index output by the quantization unit, the multimodal collaborative controller executes the defense decision logic. The multimodal collaborative controller reads the system's preset defense conditions, such as setting a confidence threshold. When the received confidence index is less than the confidence threshold (0.2), the multimodal collaborative controller determines that the target object of the current network connection exhibits abnormal behavior characteristics that conform to the expected disturbance command, and the verification evidence is complete and reliable, thereby confirming that the network connection poses a substantial security threat.

[0035] Based on the above judgment results, the multimodal collaborative controller generates defense execution instructions. These instructions contain operation codes such as blocking, resetting, or adding to a blacklist of the target network connection. These instructions are sent to the network firewall or traffic scrubbing device, immediately interrupting the communication of the malicious connection, thereby achieving real-time blocking of network threats.

[0036] In parallel with the defense execution, the multimodal collaborative controller monitors the timing status of the lock window duration. When the timing from the start of the self-generated suppression signal reaches the lock window duration, the multimodal collaborative controller automatically cancels the suppression signal. The cancellation of the suppression signal triggers the alignment unit to exit the verification feedback data marking mode and resumes normal forwarding and processing of real-time collected network traffic data and application log data. At this point, the defense system concludes the active verification process for this network connection, seamlessly switching back to the normal passive monitoring mode, completing a full self-oscillation prevention closed-loop verification operation.

[0037] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A multimodal AI-integrated real-time network information security defense system, characterized in that, The system includes: A multimodal cooperative controller is configured to: generate a disturbance command and reference information when it is determined that the network connection meets a preset triggering condition; and calculate the lock window duration and generate a suppression signal while generating the disturbance command; A disturbance injection unit, connected to the multimodal cooperative controller, is configured to send disturbance data packets to the network connection in response to the disturbance command; An alignment unit, connected to the multimodal cooperative controller, is configured to receive the reference information and, during the period of receiving the suppression signal, filter out data whose time difference and protocol type match the reference information from real-time collected network traffic data and application log data as target data and output it to the quantization unit; A quantization unit, connected to the alignment unit, is configured to receive the target data, calculate a consistency feature value characterizing the degree of consistency between the target data and the expected result of the disturbance command, and a completeness feature value characterizing the completeness of the target data, and calculate a confidence index based on the consistency feature value and the completeness feature value. The multimodal collaborative controller is also configured to receive the confidence index and generate a defense execution command when the index meets preset defense conditions.

2. The system according to claim 1, characterized in that, The alignment unit is configured as follows: The reference information is stored, including the disturbance transmission timestamp, estimated round-trip time, and expected feedback type. Obtain the received timestamp and actual protocol type of the collected data; Calculate the absolute difference between the received timestamp and the sum of the disturbance transmission timestamp and the estimated round-trip time; When the absolute difference is less than a preset time window threshold and the actual protocol type is consistent with the expected feedback type, the operation of filtering out matching data as target data and outputting it is performed.

3. The system according to claim 1, characterized in that, The quantization unit is configured as follows: Map the target data to a first vector, and obtain a preset second vector; Calculate the distance between the first vector and the second vector, or calculate the similarity between the first vector and the second vector, invert or complement the similarity, and then normalize it to obtain the consistency feature value. Count the number of data packets contained in the target data to obtain the preset expected number of data packets; The ratio of the number of data packets to the expected number of data packets is calculated as the integrity feature value.

4. The system according to claim 3, characterized in that, The quantification unit calculates the confidence index using the following formula. : in, The consistency feature value, The completeness feature value is, and The value range is [0,1], and the value 1 represents that the data packet is complete; , and The coefficients are preset positive real numbers; It is the hyperbolic tangent function.

5. The system according to claim 1, characterized in that, The multimodal cooperative controller is configured as follows: Obtain the smoothed round-trip time of the network connection and its variation. The sum of the smooth round-trip time and four times the amount of change is calculated as the duration of the locking window; Within the lock window duration after the suppression signal is generated, the suppression signal is sent to the alignment unit.

6. The system according to claim 1, characterized in that, The disturbance command instructs the disturbance injection unit to perform one of the following operations: The operation flag for discarding the last TCP fragment packet in the HTTP request message sequence; An operation identifier for sending an HTTP request packet that includes the Keep-Alive header field and the RST flag.

7. A real-time network information security defense method based on multimodal AI fusion, characterized in that, Includes the following steps: When the network connection is determined to meet the preset triggering conditions, a disturbance command and reference information are generated, the lock window duration is calculated, and a suppression signal is generated. In response to the disturbance command, a disturbance data packet is sent to the network connection; During the period when the suppression signal is in effect, based on the reference information, data with matching time difference and protocol type are selected from real-time collected network traffic data and application log data as target data; Calculate a consistency feature value that characterizes the degree of consistency between the target data and the expected results, and a completeness feature value that characterizes the degree of completeness of the target data; A confidence index is calculated based on the consistency feature value and the integrity feature value, and a defense operation is performed when the index meets the preset defense conditions.