Network security active dialing model construction method and system and computer device

By combining natural language understanding and deep quantization with pointer head scanning and multi-factor decoding quantization, a probe model based on trust-importance is constructed, which solves the problems of insufficient deep semantic parsing and obscure results in existing technologies, and realizes efficient and interpretable network threat identification and response.

CN121486098BActive Publication Date: 2026-04-24宁波市互联网信息办公室 +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
宁波市互联网信息办公室
Filing Date
2026-01-08
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing network security probing technologies lack deep semantic analysis capabilities at the data perception and understanding level. Their decision-making logic is simplistic and relies on fixed thresholds, resulting in obscure and difficult-to-understand outputs. This leads to poor ability to identify new types of attacks, low efficiency, and a high false alarm rate.

Method used

By employing natural language understanding and classification, deep quantization, four-level approximation interval coding, pointer head scanning, and multi-factor decoding quantization, a probe model based on trust-importance is constructed to generate intuitive natural language output.

Benefits of technology

It enables accurate perception and representation of network data, improves the accuracy and speed of dialing tests, reduces the false alarm rate, and enhances the robustness of the system and the interpretability of the results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121486098B_ABST
    Figure CN121486098B_ABST
Patent Text Reader

Abstract

The application provides a network security active dialing test large model construction method and system and computer equipment. The method comprises the following steps: collecting network data, performing natural language understanding and classification, converting unstructured network data into structured data with security labels, and performing primary quantization; encoding and deep quantization are performed on the network data after primary quantization, the data boundary symbol is generated, and multiple trust degrees are fused to calculate the semantic correlation trust degree; dialing is performed based on the multiple trust degrees and importance degrees obtained by scanning to obtain data dialing test results and dialing test markers; single-factor inverse quantization is performed on the boundary symbol and the code, and multi-factor comprehensive inverse coding quantization is performed, the inverse coding quantization mapping relationship is given, and the comprehensive inverse coding quantization value is calculated; the natural language sequence pair database is obtained as a standard expert database by training the weighted vector of the comprehensive inverse coding quantization value. The application improves the accuracy of dialing test, reduces the false positive rate and the false negative rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of network security technology, and in particular to a method, system and computer equipment for constructing a large-scale proactive network security probe model. Background Technology

[0002] With the rapid development of internet technology, cyberspace faces increasingly complex and ever-changing security threats, such as malicious attacks, data breaches, and abnormal traffic. To ensure the stable operation of network systems, proactive detection technology, as a proactive security monitoring method, is widely used in network performance monitoring and security situation awareness. This technology proactively probes and analyzes network services, ports, and traffic by simulating user behavior or sending probe packets, aiming to discover potential risks before security threats cause actual damage.

[0003] However, existing probing solutions combining intelligent technologies such as natural language processing still face fundamental challenges. First, at the data perception and understanding level, existing methods mostly employ shallow word frequency statistics or static template matching, failing to achieve deep semantic analysis and quantification of network data (such as logs, comment text, and protocol content), resulting in poor identification of new attacks and disguised threats. Second, at the core decision-making logic level, the probing engine's intelligence is insufficient; its judgments often rely on manually preset single thresholds, failing to build a flexible decision-making model that dynamically integrates multi-dimensional factors such as semantic trust and the importance of inter-target relationships, leading to a blind and inefficient probing process. Finally, at the result generation and interaction level, system outputs are mostly obscure encoded or binary alarm signals, creating a significant gap with human experts' natural language cognition, resulting in poor interpretability and severely hindering security personnel's rapid understanding and response to threats. Summary of the Invention

[0004] Therefore, it is necessary to provide a method, system, and computer equipment for constructing a large-scale proactive network security probe model to address the aforementioned technical issues.

[0005] In a first aspect, embodiments of the present invention propose a method for constructing a large-scale proactive network security probe model, the method comprising:

[0006] Collect network data from different networks, perform natural language understanding and classification, transform unstructured network data into structured data with security labels, perform preliminary quantization, and store it in a classification database;

[0007] The network data after initial quantization is encoded and deep quantized. Data delimiters are generated by a four-level approximation interval coding method. Multiple trust levels are fused to calculate the semantic association trust level, which is then stored in a designated trust level library.

[0008] The pointer read head scans the data along a prescribed scanning route and with prescribed delimiters. Based on the multiple trust levels and importance obtained from the scan, the data is tested, and the test results are obtained and test markers are added.

[0009] Determine the delimiter and code of each data point under the test state, perform single-factor inverse quantization on the delimiter and code, then perform multi-factor comprehensive inverse quantization, give the inverse quantization mapping relationship, and calculate the comprehensive inverse quantization value;

[0010] The weighted vector of the comprehensive decoded quantization value is trained to obtain a database of decoded quantization and corresponding natural language order pairs. The natural language order pair database serves as a standard expert database and is used to test the natural language output of the decoded quantization.

[0011] In some embodiments, the encoding and deep quantization of the initially quantized network data, and the generation of data delimiters using a four-level approximation interval coding method, include:

[0012] Calculate the probability of each symbol appearing in the data after primary quantization;

[0013] Different flags are set according to the magnitude of the probability;

[0014] The probability of a certain interval is approximated by a fourth-order calculation to obtain the boundary marker of the data.

[0015] In some embodiments, the pointer read head scans the data using the trust level, performs a dial test based on the semantic association trust level of the data according to the current state of the data being read, obtains the data dial test result, and adds a dial test flag.

[0016] The state transition method for pointer reading is defined as follows: ,in, For fuzzy transfer functions, yes The extension satisfies and , This indicates the current state of the data being read. , , It's data. and It is the level of trust in the state transition.

[0017] In some embodiments, if the data being read is located at a specific delimiter, then the current data is read. The method is as follows:

[0018] ;

[0019] If the state being read is at the start or end point, the method to read the current data (ω, μ) is as follows:

[0020] ;

[0021] in, The initial and intermediate states to be read. It is the start marker. It is a terminator. Given a finite-length string composed of input symbols, A string of stack symbols, Let F be a finite set of states, and let F be the set of trigger states for the test result. and It is the membership degree of the transfer. For the membership degree set, For all possible stack symbol strings corresponding The maximum membership degree of the transition.

[0022] In some embodiments, the step of conducting a probe based on the semantic association trust level of the data, obtaining the probe result, and adding a probe identifier includes:

[0023] The semantic association trust score is compared with the established classification database to obtain the importance between the two.

[0024] The test is conducted based on the stated importance.

[0025] In some embodiments, the marking method of the dial indicator includes:

[0026] set up ,satisfy:

[0027]

[0028] and

[0029] If:

[0030]

[0031] but That is, the judgment result, marked by a delimiter, in which , For a pre-set threshold, For the set of all possible natural language security categories, for A subset of.

[0032] In some embodiments, the step of performing single-factor inverse quantization on the delimiter and the code, followed by multi-factor comprehensive inverse quantization, providing an inverse quantization mapping relationship, and calculating the comprehensive inverse quantization value includes:

[0033] Establish factor set With decoded set ;

[0034] Use weights To describe the weight of each factor, for each factor A separate reverse encoding As a set of factors to decoded set A mapping , by mapping Inducible factor set to decoded set A relationship Then by relationship Inducible factor set to decoded set A linear transformation:

[0035]

[0036] in, B represents the inverse encoding matrix, and B represents the comprehensive inverse quantization value.

[0037] In some embodiments, the inverse quantization value is integrated. The inverse encoding matrix is , Weight set of each factor for:

[0038] , ;

[0039] in, This is the weight estimate for the k-th factor.

[0040] Secondly, embodiments of the present invention propose a network security proactive probe large-scale model construction system, the system comprising:

[0041] The classification and primary quantization module is used to collect network data from different networks, perform natural language understanding and classification, transform unstructured network data into structured data with security labels, perform primary quantization, and store it in the classification database.

[0042] The trust quantization module is used to encode and deeply quantize the network data after the initial quantization. It generates data delimiters through a four-level approximation interval coding method, integrates multiple trust degrees, calculates the semantic association trust degree, and stores it in a specified trust degree library.

[0043] The pointer read head state transition module is used to scan the data by the pointer read head according to the specified scanning route and the specified boundary markers, and to perform testing based on the multiple trust levels and importance obtained from the scanning, so as to obtain the data testing results and add testing markers.

[0044] The network dialing test module is used to determine the delimiter and code of each data in the dialing test state, perform single-factor inverse quantization on the delimiter and code, then perform multi-factor comprehensive inverse quantization, give the inverse quantization mapping relationship, and calculate the comprehensive inverse quantization value;

[0045] The dequantization module is used to train the weighted vector of the comprehensive decoded quantization value to obtain the decoded quantization and the corresponding natural language sequence pair database. The natural language sequence pair database serves as a standard expert database and is used to test the natural language output of the decoded quantization.

[0046] Thirdly, embodiments of the present invention provide a computer device including a memory and a processor, wherein the memory stores a computer program and the processor executes the steps described in the first aspect.

[0047] The beneficial effects of the above methods, systems, and computer equipment are as follows:

[0048] By combining a two-level quantization mechanism of "natural language understanding classification and primary quantization" with "deep quantization and semantic association trust calculation", the problem of insufficient deep semantic understanding of network data and coarse quantization in existing technologies is solved, and the accurate perception and characterization of ever-changing threats is realized, thereby improving the accuracy of probe detection from the source.

[0049] By introducing a two-factor joint detection decision model based on "trust level-importance", the problem of existing technologies relying on fixed thresholds and having a single and rigid decision-making mechanism is solved, realizing intelligent and focused proactive detection, and significantly reducing false alarm rate and false negative rate.

[0050] By employing a pointer read head state transition mechanism to perform flexible scanning of the encoded signal, the problem of poor adaptability of traditional rigid rule engines in complex network environments is solved, and the robustness of the system in the face of noisy data and unknown threats is enhanced.

[0051] By constructing a "standard sequence pair expert database" using "multi-factor comprehensive reverse coding quantization", the problem of obscure and difficult-to-understand output results and poor interpretability of existing dialing systems is solved. It realizes the reverse mapping of machine coding into intuitive natural language, providing a clear and operable basis for security decision-making.

[0052] By employing specific methods such as four-level approximation interval coding and pointer-directed scanning, the problems of bloated traditional system architecture and low efficiency when processing massive amounts of data are solved, achieving a significant improvement in testing speed while ensuring accuracy. Attached Figure Description

[0053] Figure 1 This is a flowchart illustrating a method for constructing a large-scale network security proactive probe model in one embodiment;

[0054] Figure 2 This is a flowchart illustrating a method for generating data delimiters in one embodiment;

[0055] Figure 3 This is a schematic diagram illustrating the reading process of the pointer read head in one embodiment;

[0056] Figure 4 This is a flowchart illustrating a dialing test method in one embodiment;

[0057] Figure 5 This is a schematic diagram of the structure of a network security proactive probe large model construction system in one embodiment. Detailed Implementation

[0058] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are merely some examples or embodiments of the present invention. For those skilled in the art, the present invention can be applied to other similar scenarios based on these drawings without creative effort. Unless obvious from the context or otherwise specified, the same reference numerals in the drawings represent the same structures or operations.

[0059] As indicated in this invention and the claims, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" do not specifically refer to the singular and may also include the plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of explicitly identified steps and elements, which do not constitute an exclusive list, and the method or apparatus may also include other steps or elements.

[0060] While this invention makes various references to certain modules in an apparatus according to embodiments of the invention, any number of different modules can be used and run on a computing device and / or processor. Modules are merely illustrative, and different aspects of the apparatus and methods may use different modules.

[0061] It should be understood that when a unit or module is described as "connected" or "coupled" to other units, modules, or blocks, it may refer to a direct connection or coupling, or communication with other units, modules, or blocks, or the presence of intermediate units, modules, or blocks, unless the context explicitly indicates otherwise. The term "and / or" as used herein may include any and all combinations of one or more of the related listed items.

[0062] like Figure 1 As shown, this embodiment of the invention provides a method for constructing a large-scale proactive network security probing model, including the following steps:

[0063] S102: Collect network data from different networks, perform natural language understanding and classification, transform unstructured network data into structured data with security labels, perform preliminary quantization, and store it in the classification database;

[0064] S104, the network data after primary quantization is encoded and deep quantized. Data delimiters are generated by the four-level approximation interval coding method, and multiple trust degrees are fused to calculate the semantic association trust degree, which is then stored in the specified trust degree library.

[0065] S106, The pointer read head scans the data according to the prescribed scanning route and the prescribed delimiters, and performs probing based on the multiple semantic association trust and importance obtained from the scan, and obtains the data probing results and adds probing markers;

[0066] S108, determine the delimiter and code of each data in the dialing test state, perform single-factor inverse quantization on the delimiter and the code, then perform multi-factor comprehensive inverse coding quantization, give the inverse coding quantization mapping relationship, and calculate the comprehensive inverse coding quantization value;

[0067] S110, The weighted vector of the comprehensive decoded quantization value is trained to obtain the decoded quantization and the corresponding natural language sequence pair database. The natural language sequence pair database serves as a standard expert database and is used to test the natural language output of the decoded quantization.

[0068] This embodiment solves the problems of insufficient deep semantic understanding and coarse quantization of network data in existing technologies by combining "natural language understanding classification and primary quantization" with "deep quantization and semantic association trust calculation". It achieves accurate perception and characterization of ever-changing threats and improves the accuracy of probe detection from the source.

[0069] This embodiment solves the problems of existing technologies relying on fixed thresholds and having a single, rigid decision-making mechanism by introducing a two-factor joint detection decision model based on "trust level-importance". It achieves intelligent and focused proactive detection, significantly reducing the false alarm rate and false negative rate.

[0070] This embodiment uses a pointer read head state transition mechanism to perform flexible scanning of the encoded signal, which solves the problem of poor adaptability of traditional rigid rule engines in complex network environments and enhances the robustness of the system when facing noisy data and unknown threats.

[0071] This embodiment utilizes "multi-factor comprehensive reverse coding quantization" to construct a "standard sequence pair expert database," which solves the problem of obscure and difficult-to-understand output results and poor interpretability of existing testing systems. It realizes the reverse mapping of machine coding into intuitive natural language, providing a clear and operable basis for security decision-making.

[0072] This embodiment utilizes specific methods such as four-level approximation interval coding and pointer-directed scanning to solve the problems of bloated traditional system architecture and low efficiency when processing massive amounts of data. While ensuring accuracy, it achieves a significant improvement in dialing speed.

[0073] In step S102, based on website or network data analysis, the data on the webpage is transformed into structured data, and these structured data are labeled as secure, insecure, and with different security levels. Multiple keyword databases with different characteristics are established, such as usernames, passwords, ports, module entry protocols, module exit protocols, URLs, and various sensitive words, to obtain different security detection categories and create a classification database.

[0074] Structured data from the previous layer cannot be directly input into the next layer; it requires quantification. Data quantification transforms raw network data into semantic association trust calculations. Specific operations include: ① segmenting comment data into words, dividing each Chinese character in a user comment into an independent word; ② removing punctuation, emoticons, special characters, and stop words from sentences; ③ restoring different word forms of the same phrase to their original form, avoiding excessive consideration of word form influence; ④ querying and removing a large amount of similar and duplicate content; and ⑤ specifying the required window length to ensure that each comment has the same sentence size.

[0075] In step S104, as Figure 2 As shown, the encoding and deep quantization of the initially quantized network data, and the generation of data delimiters using a four-level approximation interval coding method, include:

[0076] S202, calculate the probability of each symbol information appearing in the data after primary quantization;

[0077] S204, set different flags according to the probability;

[0078] S206, perform a four-level probability approximation calculation on a certain interval to obtain the data boundary marker.

[0079] For example, if the probability of an account being insecure is 0.2, its flag is set to... Then its probability p( =0.2. Then, through a four-level approximation operation, i.e. , ;further, , ;then, , ;at last, , Any real number within the last pair of intervals [0.06752, 0.0688] can represent the corresponding identifier. The encoding is performed to make the code length as short as possible, taking three decimal places, i.e., using 0.068 as the identifier. Decoding is simply the inverse operation of the above encoding.

[0080] set up For the same type of natural language dataset, for any natural language item in the set Different experts can provide insights into this natural language processing project. Understanding lends a certain degree of trust .set up and yes Two independent trust levels on a subset.

[0081] right , and ,calculate:

[0082] or , (1)

[0083] in, and yes A subset of , representing the degree of trust and The set of propositions supported. The conflict coefficient represents... and The degree of conflict between them. This is a normalization factor used to normalize the synthesized trust level to... Interval.

[0084] Then semantic association trust degree for:

[0085] (2)

[0086] yes A non-empty subset of represents the propositions supported by the combined trust level.

[0087] In equation (2), if ,but Determine a semantic association trust level; if Then it is believed , A contradiction occurs when two semantic terms cannot be related, meaning they are unrelated. In this case, these two semantic terms should be assigned to different databases. For combinations of multiple trust levels, the trust levels can be paired, and the two semantic terms can be assigned to the same database.

[0088] In step S106, the pointer read head uses the semantic association trust level output from the previous layer to scan different encoded signals. This embodiment shows the pointer read head's scanning state transitions... For fuzzy transition functions, secure probes are implemented. The pointer read state transition method is defined as follows: ,here, yes The extension satisfies and , This indicates the current state of the data being read. , , It's data. and It is the level of trust in the state transition.

[0089] The process of the pointer read head reading signal is as follows: Figure 3 As shown, Figure 3 medium state This is the starting state of the process of the pointer reading head reading data. , It reads the intermediate state; double circle This indicates the status of the read result. Only the path leading to the final read result status is shown here. (Input data...) or With a certain degree of membership From state arrive The change is represented as from arrive The mark , or A directed arc. The pointer (denoted as FPDA) can ultimately read certain data in the following two ways:

[0090] (I) If the read state is at a specific delimiter, then the current data is read. The method is as follows:

[0091] (3)

[0092] (II) If the state being read is at the start or end point, i.e., an empty string, then the method to read the current data (ω, μ) is as follows:

[0093] (4)

[0094] in, The initial and intermediate states to be read. It is the start marker. It is a terminator. Given a finite-length string composed of input symbols, A string of stack symbols, Let F be a finite set of states, and let F be the set of trigger states for the test result. and It is the membership degree of the transfer. For the membership degree set, For all possible stack symbol strings corresponding The maximum membership degree of the transition.

[0095] In step S106, as Figure 4 As shown, the probe test based on the semantic association trust level of the data, obtaining the data probe test results and adding probe test markers includes:

[0096] S402, compare the semantic association trust with the established classification database to obtain the importance between the two;

[0097] S404, Perform a test based on the stated importance.

[0098] The state transition scanning method in the previous layer reads the quantified semantic association trust level and compares it with the classification database established in the first layer. This comparison determines which semantic importance interval the currently read trust level belongs to, thus enabling probing of specific data. Therefore, it is necessary to first define the importance between the trust level of the currently read data and the quantified category library.

[0099] The current signal confidence level and the quantized category library are taken as two objectives. and Define the target set Importance is ,here The number.

[0100] if ,but and It is an independent item, indicating and Their occurrence does not affect each other. If ,but and It is negatively correlated, which means that if or Presentation, then the target or It's unlikely to happen. If... ,but and It is positively correlated, which means that if in the read region... Presentation, then the target It may also occur.

[0101] according to and Different delimiters are assigned based on the relevance. If and Related, assigning the same encoding but different header characters. If and Unrelated targets are assigned different codes and different header characters. The pointer read head's judgment on the probing of two adjacent targets is determined by the semantic association trust level and the importance between the two targets. If Then dial the test And add a dial indicator, for example But if and Then dial the test And add a dial indicator, for example Simultaneous testing of multiple targets is determined by multiple levels of trust and importance. The targets are paired up according to the method of assigning delimiters to the two types of targets, and then tested sequentially.

[0102] The delimiter is marked by the following method: ,satisfy:

[0103] (5)

[0104] and (6)

[0105] If:

[0106] (7)

[0107] but This is the judgment result, marked by a delimiter. Among them... , The threshold is set in advance. It is the set of all possible natural language security categories. for subsets, for example ={malice} ={Suspicious}.

[0108] In step S108, at the dequantization layer, the target of the measurement is finally output based on the delimiter, such as the flow rate level. However, the output of the previous layer is an encoded number, so it is necessary to dequantize the number back into natural language. Dequantizing an encoded number into natural language often involves multiple factors or indicators. In this case, it is required to make a comprehensive quantization of the target based on these multiple factors, rather than quantifying the target based on only one factor.

[0109] The comprehensive quantification method is a highly effective multi-factor inverse quantification method that comprehensively de-encodes the coded numbers resulting from the influence of multiple factors. The specific method is as follows:

[0110] set up for Factors (or indicators), for A type of reverse encoding.

[0111] Because various factors occupy different positions and have different effects, weights can be used. To describe importance, it is a set of factors. A subset of. For each factor A separate reverse encoding It can be seen as arrive A mapping ,Depend on Can induce arrive A decoded quantization mapping relationship Then by Can induce arrive A linear transformation:

[0112] (8)

[0113] It is a decoded set A subset of this is the synthetic decoding. Here, the synthetic decoding quantization value is... , , It is a composition operation, for example Calculation.

[0114] Constructing a comprehensive reverse coding decision model, , , These are the three elements of this model.

[0115] The method and steps of comprehensive reverse coding decision-making are:

[0116] Establish factor set With decoded set ;

[0117] Establish a comprehensive inverse encoding matrix;

[0118] For each factor First, establish a single-factor inverse coding: ;

[0119] Right now ( )express Factors The inverse encoding is performed, thus obtaining the single-factor inverse encoding matrix. .

[0120] Based on the weight of each factor Comprehensive reverse encoding: yes A subset of . Operation It is defined based on the requirements of the actual problem. Based on the calculation... Different definitions of can yield different results. Here, we define . The operation is The corresponding dequantization model is as follows:

[0121] , (9)

[0122] Due to the result of comprehensive reverse encoding The value is determined only by and ( This method, which involves a specific operation within a given context, focuses on the primary factors, while other factors have little impact on the result. Sometimes, the dequantization result can be difficult to discern. In such cases, the following dequantization model is used:

[0123] (10)

[0124] In practical applications, if the main factor plays a dominant role in the comprehensive reverse coding, it is recommended to adopt (9) or (10). When model (9) fails, (10) can be adopted.

[0125] In comprehensive decoding and dequantization, weights are crucial, reflecting the status or role of each factor in the comprehensive dequantization process, and directly affecting the result of comprehensive dequantization.

[0126] Weights assigned based on expert experience can reflect the actual situation to a certain extent, and the results of reverse encoding are relatively consistent with reality. However, they are often subjective and cannot objectively reflect the actual situation; the reverse encoding results may be "distorted." Here, a method for determining weights is presented as follows:

[0127] In the problem of comprehensive decoding and dequantization, if the comprehensive dequantization value is known... The inverse encoding matrix is So, what is the weight distribution of each factor? yes:

[0128] , (11)

[0129] in, This is the weight estimate for the k-th factor.

[0130] This section presents an approximate method for weight processing, with a set of selectable weight allocation schemes. We started from... Choose the optimal weight allocation. This makes it possible for the The determined comprehensive decoding and dequantization and The closest match is obtained, ultimately yielding the corresponding natural language.

[0131] In step S110, all data scanned by the read head is decoded and its corresponding natural language is stored as index pairs <α, β> in a designated database as a standard expert library. Here, α is the delimiter and its decoded value, and β is the dequantized natural language. Based on this standard expert library, the output of the model is the dequantized corresponding delimiter and its encoded natural language.

[0132] Here is a comprehensive example of inverse quantification, which uses four types of data corresponding to four indicators as an example for the inverse quantification of the target. , denoted as The corresponding factor set Inverse quantization set .

[0133] Single factor for category 1 Implementing single-factor reverse quantification means that if 20% of the reverse quantification is positive, 50% is neutral, 20% is negative, and 10% is adverse, then the conclusion is: positive. .

[0134] Similarly, let's assume: neutral ;negative ;bad

[0135] Therefore, there is an inverse quantization matrix. .

[0136] Since the four indicator factors corresponding to different data types are different, their weights are also different. Let the weight given to type 1 by the inverse quantification method be... Therefore, based on the inverse quantization of the four types of data for each category, the comprehensive inverse quantization of category 1 can be obtained as follows: The composition operation here Perform the calculation according to formula (8). The inverse quantification is represented as follows: "positive" at 20%; "neutral" at 30%; "negative" at 40%; and "bad" at 10%. Based on the principle of maximum membership, the conclusion is that the output inverse quantification result is "negative".

[0137] It should be understood that although the steps in the flowchart above are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowchart above may include multiple steps or stages, which are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps.

[0138] In one embodiment, such as Figure 5 As shown in the figure, this invention provides a network security proactive probe testing large model construction system, the system comprising:

[0139] The classification and primary quantization module 502 is used to collect network data from different networks, perform natural language understanding and classification, transform unstructured network data into structured data with security labels, perform primary quantization, and store it in the classification database.

[0140] The trust quantization module 504 is used to encode and deeply quantize the network data after the initial quantization, generate data delimiters through a four-level approximation interval coding method, fuse multiple trust degrees, calculate the semantic association trust degree, and store it in a specified trust degree library.

[0141] The pointer read head state transition module 506 is used to scan the data by the pointer read head according to the prescribed scanning route and the prescribed boundary markers, perform dialing tests based on the multiple trust levels and importance obtained from the scan, obtain the data dialing test results and add dialing test markers;

[0142] The network dialing test module 508 is used to determine the delimiter and code of each data in the dialing test state, perform single-factor inverse quantization on the delimiter and the code, then perform multi-factor comprehensive inverse coding quantization, give the inverse coding quantization mapping relationship, and calculate the comprehensive inverse coding quantization value;

[0143] The dequantization module 510 is used to train the weighted vector of the comprehensive decoded quantization value to obtain the decoded quantization and the corresponding natural language sequence pair database. The natural language sequence pair database serves as a standard expert database and is used to test the natural language output of the decoded quantization.

[0144] In one embodiment, the present invention provides a computer device including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps in any of the above embodiments of the network security proactive probe large model construction method.

[0145] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical storage, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0146] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0147] The above embodiments merely illustrate several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A method for constructing a large-scale proactive network security probing model, characterized in that, The method includes: Collect network data from different networks, perform natural language understanding and classification, transform unstructured network data into structured data with security labels, perform preliminary quantization, and store it in a classification database; The network data after initial quantization is encoded and deep quantized. Data delimiters are generated by a four-level approximation interval coding method. Multiple trust levels are fused to calculate the semantic association trust level, which is then stored in a designated trust level library. The pointer read head scans the data according to the specified trust level and the specified delimiter. Based on the current state of the data being read, the data is tested based on the multiple trust levels and importance obtained from the scan, and the test results are obtained and a test marker is added. Determine the delimiter and code of each data point under the test state, perform single-factor inverse quantization on the delimiter and code, then perform multi-factor comprehensive inverse quantization, give the inverse quantization mapping relationship, and calculate the comprehensive inverse quantization value; The weighted vector of the comprehensive decoded quantization value is trained to obtain a database of decoded quantization and corresponding natural language order pairs. The natural language order pair database serves as a standard expert database for probing the natural language output of decoded quantization. The process of encoding and deep quantizing the initially quantized network data, and generating data delimiters using a four-level approximation interval coding method, includes: Calculate the probability of each symbol appearing in the data after primary quantization; Different flags are set according to the magnitude of the probability; The probability is approximated to a certain interval using a four-level approximation method to obtain the data's boundary marker.

2. The method according to claim 1, characterized in that, The state transition method for reading by the pointer read head is defined as follows: ,in, For fuzzy transfer functions, yes The extension satisfies and , This indicates the current state of the data being read. , , It's data. and It is the level of trust in the state transition.

3. The method according to claim 2, characterized in that, If the data being read is located at a specific delimiter, then the current data is read. The method is as follows: ; If the state being read is at the starting or ending point, the method for reading the current signal (ω, μ) is as follows: ; in, The initial and intermediate states to be read. It is the start marker. It is a terminator. Given a finite-length string composed of input symbols, A string of stack symbols, Let F be a finite set of states, and let F be the set of trigger states for the test result. and It is the membership degree of the transfer. For the membership degree set, For all possible stack symbol strings corresponding The maximum membership degree of the transition.

4. The method according to claim 2, characterized in that, The process of conducting a probe based on the semantic association trust level of this data, obtaining the probe results, and adding probe markers includes: The semantic association trust score is compared with the established classification database to obtain the importance between the two. The test is conducted based on the stated importance.

5. The method according to claim 4, characterized in that, The marking methods for dial indicator symbols include: set up ,satisfy: and If: but That is, the judgment result, marked by a delimiter, in which , For a pre-set threshold, For the set of all possible natural language security categories, for A subset of.

6. The method according to claim 1, characterized in that, The process of performing single-factor inverse quantization on the delimiter and the code, followed by multi-factor comprehensive inverse quantization, providing the inverse quantization mapping relationship, and calculating the comprehensive inverse quantization value includes: Establish factor set With decode set ; Use weights To describe the weight of each factor, for each factor A separate reverse encoding As a set of factors to decoded set A mapping , by mapping Inducible factor set to decoded set A relationship Then by relationship Inducible factor set to decoded set A linear transformation: in, B represents the inverse encoding matrix, and B represents the comprehensive inverse quantization value.

7. The method according to claim 6, characterized in that, Comprehensive inverse quantization value The inverse encoding matrix is , Weight set of each factor for: , ; in, This is the weight estimate for the k-th factor. Indicates the j-th inverse code Factors The reverse encoding performed, This represents the j-th comprehensive inverse quantization value.

8. A network security proactive probe testing large-scale model construction system, characterized in that, The system includes: The classification and primary quantization module is used to collect network data from different networks, perform natural language understanding and classification, transform unstructured network data into structured data with security labels, perform primary quantization, and store it in the classification database. The trust quantization module is used to encode and deeply quantize the network data after initial quantization, generate data delimiters through a four-level approximation interval coding method, and fuse multiple trust levels to calculate the semantic association trust level, which is then stored in a designated trust level library. The process of encoding and deeply quantizing the network data after initial quantization and generating data delimiters through a four-level approximation interval coding method includes: calculating the probability of each symbol information appearing in the initially quantized data; setting different delimiters based on the probability magnitude; and performing a four-level approximation interval calculation on the probability to obtain the data delimiters. The pointer read head state transition module is used to scan the data by the pointer read head according to the specified scan route and specified delimiters using the trust level. Based on the current state of the read data, it performs a test based on the multiple trust levels and importance obtained from the scan, obtains the data test result, and adds a test flag. The network dialing test module is used to determine the delimiter and code of each data in the dialing test state, perform single-factor inverse quantization on the delimiter and code, then perform multi-factor comprehensive inverse quantization, give the inverse quantization mapping relationship, and calculate the comprehensive inverse quantization value; The dequantization module is used to train the weighted vector of the comprehensive decoded quantization value to obtain the decoded quantization and the corresponding natural language sequence pair database. The natural language sequence pair database serves as a standard expert database and is used to test the natural language output of the decoded quantization.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Federal learning communication method, system and device

    CN119544695A

  • Cloud data anomaly detection and safety response system based on artificial intelligence

    CN120614208A