Commercial password detection data identification and correlation analysis method based on large model

By building a lightweight memory storage module and a dynamic memory recycling mechanism, combined with the BERT architecture fine-tuning model and redundant data temporary storage module, the problem of inaccurate redundant data processing in commercial password detection is solved, and efficient and accurate data identification and association analysis are achieved.

CN120705519AActive Publication Date: 2025-09-26ZHIXUN CIPHER (SHANGHAI) TESTING TECH CO LTD +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511203279.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-27
Publication Date
2025-09-26
Estimated Expiration
2045-08-27

AI Technical Summary

Technical Problem

Existing commercial password detection methods based on large models lack effective mechanisms for judging and processing redundant data, resulting in low detection accuracy and high false alarm rate, making it difficult to meet high-precision requirements.

Method used

Build a lightweight memory storage module, set up a dynamic memory recovery mechanism and redundant data confirmation cycle, use a dedicated model fine-tuned based on the BERT architecture to extract key entities and contextual logical relationships, dynamically generate memory indexes, and optimize the retrieval process through a timeout mechanism and redundant data temporary storage module.

Benefits of technology

Significantly reduce memory usage, improve the efficiency and accuracy of redundant data detection, reduce the probability of false positives, optimize the retrieval process, and improve the accuracy and precision of commercial password detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120705519A_ABST
    Figure CN120705519A_ABST
Patent Text Reader

Abstract

The invention discloses a commercial password detection data identification and association analysis method based on a large model, and relates to the technical field of commercial password detection, and the method comprises the steps: constructing a lightweight memory storage module, and setting a dynamic memory recovery mechanism; performing real-time processing on the commercial password detection data by using the large model; automatically retrieving the associated memory through a memory index generated by a large model; comparing an analysis result output by the large model with a memory unit in a lightweight memory storage module; the types of misinformation cases are collected and analyzed, and the memory index is optimized according to the analysis result; according to the method, the redundant data detection duration and the redundant data determination period which are dynamically changed are introduced, a redundant data temporary storage module is arranged, and the detection recovery duration of the redundant data determination period is reserved, so that the misjudgment probability of the redundant data is reduced, and the detection efficiency is improved. And the probability of false alarm of the commercial password detection result caused by misjudgment of redundant data is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of commercial password detection, and in particular to a commercial password detection data identification and association analysis method based on a large model. Background Art

[0002] In the work of commercial password detection, it is usually necessary to identify, analyze and verify a large amount of data related to commercial passwords to ensure that the use of commercial passwords complies with relevant standards and specifications; traditional commercial password detection data processing methods usually rely on fixed rules and templates, and have poor adaptability to complex text data and dynamically changing scenarios; at the same time, in the process of processing password detection data, it is easy to have problems such as fictitious error information, large memory usage and inaccurate grasp of the logical relationship of data context. These problems will lead to low accuracy of detection results, resulting in a high false alarm rate, which cannot meet the high precision requirements of commercial password detection.

[0003] With the development of big model technology, big models have demonstrated powerful processing capabilities in natural language processing and data recognition. However, in the field of commercial password detection, existing big model-based processing methods do not have further redundant data confirmation mechanisms and honor data detection recovery mechanisms when designing lightweight storage modules for the judgment and processing of redundant data. In addition, the fixed detection time for redundant data may lead to inaccurate data judgment results, which further leads to low accuracy and high false alarm rate of commercial password detection results. In addition, there is a lack of special design for the characteristics of password detection data, especially in terms of memory storage, association retrieval and error interception. It is difficult to efficiently and accurately complete the identification and association analysis of commercial password detection data. Summary of the Invention

[0004] The purpose of the present invention is to provide a commercial password detection data identification and association analysis method based on a large model to solve the problems raised in the prior art.

[0005] To achieve the above objectives, the present invention provides the following technical solution: a method for identifying and analyzing commercial password detection data based on a large model, the method comprising the following steps: S1: Build a lightweight memory storage module and set up a dynamic memory recycling mechanism for the memory storage module. When designing the dynamic memory recycling mechanism, set the redundant data confirmation cycle and dynamic detection duration adjustment algorithm, and set up a redundant data temporary storage module; S2: Use a large model to process commercial password detection data in real time, extract key entities and contextual logical relationships in the data, and dynamically generate memory indexes based on the extracted logical relationships; S3. When analyzing a new commercial password detection text, automatically retrieve the associated memory related to the current commercial password detection text in the lightweight memory module through the memory index generated by the large model; S4: Compare the analysis results output by the large model with the associated memory in the lightweight memory storage module. If the results are inconsistent with the information in the memory unit, intercept the results and mark the inconsistent points. S5: Collect false alarm cases, analyze the types and causes of false alarm cases, and optimize the memory index based on the analysis results.

[0006] Furthermore, in step S1: the constructed lightweight memory storage module includes an entity storage unit, a logical relationship storage unit and an index storage unit; the entity storage unit stores key entities extracted from commercial cipher texts through a hash table structure; the logical relationship storage unit stores contextual logical relationships between key entities in the form of triples; the index storage unit stores information in the memory index table through a B+ tree structure; the module adopts an efficient storage structure, which can significantly reduce memory usage while ensuring data storage integrity; at the same time, a dynamic memory recovery mechanism is set to regularly detect the access frequency of stored data and clean up redundant data that has not been accessed for a long time, thereby freeing up memory space.

[0007] Furthermore, in step S1: the dynamic memory recycling mechanism performs access detection on the lightweight memory storage module at regular intervals, and sets a redundant data confirmation cycle. In one redundant data confirmation cycle, a total of m redundant data detections are performed, and the number of unaccessed data detected each time is recorded. {n1, n2, ... n m}, record the data with 0 access times detected in m redundant data detections as a set {A1, A2, ..., A m}, after completing m redundant data detections, find the determined redundant data set A, , where i = 1, 2, ... m, represents the number of m redundant data detections; A represents the set of data determined to be redundant within a redundant data determination cycle; The duration between each redundant data detection within a redundant data determination cycle It is a dynamically changing value. The data access duration selected by the first and second redundant data detection is set to Δt0. Starting from the third redundant data detection, the data access duration selected by each redundant data detection is The change is based on the change in the number of data with 0 access times detected in the previous monitoring, and is calculated according to the following formula: ; Wherein, j=3, 4, ..., m, represents the number of each redundant data detection starting from the third redundant data detection; Indicates the data access duration selected by each redundant data detection starting from the third redundant data detection. Based on the calculated dynamic change value of the data access duration selected by each redundant data detection, the cycle length of the entire redundant data determination cycle can be calculated: ; Among them, T represents the cycle length of a redundant data determination cycle; a dynamic memory recovery mechanism is set up to delete redundant data and release memory space. When designing the dynamic memory recovery mechanism, a dynamically changing redundant data detection time and redundant data determination cycle are introduced, wherein the redundant data detection time is adjusted according to the changing trend of the number of redundant data detected in the previous two redundant data detection processes. The redundant data detection time can be automatically and dynamically adjusted according to the change in the number of detected data, so that the number in the redundant data set detected each time is relatively balanced, so that the result of the redundant data obtained by the intersection at the end of the final redundant data determination cycle is more accurate; and the monitoring time is adjusted in real time according to the number of redundant data monitored. When the number of data detected decreases, the detection time is automatically increased, and the number of redundant data detections in the case of low redundant data volume is reduced, thereby improving the redundant data detection efficiency and the accuracy of the detection results.

[0008] Furthermore, a redundant data temporary storage module is provided. After a redundant data determination cycle ends, the redundant data in the determined redundant data set A is temporarily stored in the redundant data temporary storage module. In the next redundant data determination cycle, the redundant data stored in the redundant data temporary storage module can also be used for comparison with the new commercial cipher text data. The data stored in the redundant data temporary storage module is processed according to the comparison result. The comparison result is analyzed as follows: If the comparison result of the redundant data stored in the redundant data temporary storage module is a successful comparison, that is, if data having the same logical relationship between the key entity and the entity as the new commercial cipher text data is retrieved from the redundant data, that is, if the associated memory is retrieved, then the comparison result is a successful comparison, and the successfully matched redundant data is extracted from the redundant data temporary storage module and re-stored in the lightweight memory storage module, and is no longer treated as redundant data, thereby completing the detection and recovery of redundant data; If the comparison result of the redundant data stored in the redundant data temporary storage module is a comparison failure, that is, if no data with the same logical relationship between the key entities and entities as the new commercial password text data is retrieved in the redundant data, that is, no associated memory is retrieved, then the comparison result is a comparison failure, and after the next redundant data determination cycle ends, the redundant data that has not been detected and recovered will be deleted, thereby freeing up memory space; a redundant data temporary storage unit is set up to give data confirmed as redundant data a detection and recovery opportunity, and a detection and recovery time of a redundant data determination cycle is reserved. After the large model retrieval triggers the timeout mechanism for the retrieval of the associated memory, the retrieval of the redundant data stored in the data temporary storage unit can be triggered. This design provides an opportunity for detection and recovery for redundant data, reduces the probability of misjudgment of redundant data, and reduces the probability of false positives in commercial password detection results due to misjudgment of redundant data.

[0009] Furthermore, in step S2: a dedicated model fine-tuned based on the BERT architecture is used to process commercial cryptography detection data in real time to extract key entities and contextual logical relationships in the data. The model incorporates a large amount of commercial cryptography detection-related text data during training, which has stronger domain adaptability. The key entities include the name of the commercial cryptography algorithm, the model of the cryptographic device, the key length, the encryption method, the cryptographic application scenario, and the compliance certification status. The contextual logical relationships include the association relationship, temporal relationship, causal relationship, and constraint relationship between entities. In the process of extracting key entities and contextual logical relationships from the data, named entity recognition technology and relationship extraction technology are used, combined with professional dictionaries in the field of commercial cryptography, including commonly used algorithms, equipment, terminology, etc. in this field. Model parameters are optimized through multiple rounds of iterative training to improve the accuracy of key entity and contextual logical relationship extraction. Fuzzy and ambiguous information is marked and further confirmed in combination with subsequent association analysis. Based on the extracted key entities and contextual logical relationships, a memory index table is dynamically generated. The memory index table is centered on the key entities and records the logical relationship information corresponding to the entities and their location information in the text for subsequent rapid retrieval.

[0010] Furthermore, in step S3, the process of detecting new commercial cipher text data includes the following steps: S3-1. After pre-processing the new commercial cipher text data, use the same large model and technology as the key information extraction step to extract key entities in the data. The extracted key entities include entity name, type, and occurrence location; S3-2. Based on the extracted key entities, the B+ tree structure of the memory index table is traversed, and associated memories related to the new commercial password text data in the memory index table are retrieved through information comparison, and the analysis and comparison results between the new commercial password data and the retrieved associations and memories are output; S3-3. Set a timeout mechanism for the retrieval process. If the retrieval time exceeds the set threshold, the current retrieval is terminated and the retrieval results are analyzed and processed to optimize the retrieval process.

[0011] Furthermore, in step S3-3: a search timeout mechanism is set for the search process, and a search time threshold Δt is set for the search process. max , compare and analyze the real-time retrieval time Δt′ with the set threshold, and determine whether to trigger the timeout mechanism based on the retrieval results and the real-time retrieval time; If the real-time retrieval time Δt′<Δt max , if it is determined that the associated memory related to the new commercial password text data has been retrieved from the memory index table, the timeout mechanism will not be triggered and the analysis of the retrieval results will be directly entered; If the real-time retrieval time Δt′≥Δt max , if it is determined that no associated memory related to the new commercial password text data is retrieved from the memory index table, or no associated memory that completely matches the new commercial password is retrieved, a timeout mechanism is triggered; and after the timeout mechanism is triggered, a redundant data temporary storage module retrieval is started for the current commercial password text data, and a timeout mechanism is also set in the process of retrieving the redundant data stored in the redundant data temporary storage module; if an associated memory is retrieved in the redundant data temporary storage module, the commercial password text data corresponding to the retrieved associated memory is extracted from the redundant data temporary storage module and re-stored in the lightweight memory storage module.

[0012] Furthermore, in step S4: the analysis and comparison results between the new commercial cryptographic data output by the large model and the retrieved associations and memories include the differences in the core attributes of the key entities, that is, the deviation lengths between the key lengths corresponding to the algorithms, and the logical relationships between the entities; the analysis results output by the large model are compared with the specific contents of the association memories in the lightweight memory storage module, including the deviation lengths between the key lengths corresponding to the algorithms and the logical relationships between the entities; the comparison method adopts a hierarchical comparison strategy, first comparing the core attributes of the key entities, that is, the key lengths corresponding to the algorithms, and then comparing the key logical relationships between the entities; For the detection of core attributes of key entities, a contradiction judgment threshold is set. For key length data, the key length contradiction threshold corresponding to the algorithm is set to L0. The deviation length between the key length in the new commercial cipher text extracted in real time by the large model and the key length of the retrieved associated memory is recorded as L. By analyzing and comparing the relationship between L and L0, it is determined whether there is a contradiction. For logical relationship analysis, if there is a direct conflict in the logical relationship between entities, it is directly determined that there is a contradiction. For example, the logical relationship between the AES algorithm and the SM4 algorithm stored in the memory unit is: first use the AES algorithm and then upgrade to the SM4 algorithm. However, the logical relationship between the two entities in the detection result of the large model is: first use the SM4 algorithm and then upgrade to the AES algorithm. In this case, it is determined that there is a direct conflict in the logical relationship and it is determined that there is a contradiction. The analysis results are as follows: If L≤L0 and there is no direct conflict in the logical relationship, it is determined that the real-time commercial password does not contradict the retrieved associated memory, and the output continues to the subsequent process; If L>L0 or there is a direct conflict in the logical relationship, it is determined that there is a contradiction between the real-time commercial password and the retrieved associated memory; For analysis results that are judged to be contradictory, the analysis results will be intercepted, the output to the subsequent process will be suspended, the contradiction points will be marked through a structured form, and the contradiction information will be temporarily stored in the pending queue, waiting for review by the management personnel; when the management personnel reviews, they will check all the information in the structured form to determine whether the intercepted result is a false alarm; after review, if it is determined to be a false alarm, the system will automatically include the case in the training set of the memory index optimization step, and release the interception, allowing the result to enter the subsequent process; if it is determined to be a valid contradiction, it will be processed according to the instructions of the management personnel, such as correcting the output results of the large model and supplementing the memory unit information; the corrected large model output results need to be compared with the memory unit again, and only after confirming that there is no contradiction can the subsequent process be entered; the supplemented memory unit information will be updated to the lightweight memory storage module for subsequent comparison and analysis.

[0013] Furthermore, in step S5: false positive cases are collected. The types of false positive cases include correct results that are mistakenly intercepted, wrong results that are not intercepted, and cases of associated memory retrieval errors. The collected false positive cases are used to automatically adjust the weight of the memory index. The weight calculation is performed using a dynamic adjustment formula: ; in, represents the new memory index weight after adjustment; q represents the memory index weight before adjustment; s represents the set adjustment coefficient, which is related to the degree of influence of false alarms on the detection results; Indicates the false alarm frequency of each memory index entry; Indicates the total usage frequency of each memory index entry; The adjustment coefficient s is the system default setting. The degree of influence of false alarms on the detection results varies according to the type of false alarm case. The two types of false alarm cases, namely, the correct results that were mistakenly intercepted and the wrong results that were not intercepted, have a high degree of influence on the detection results. Therefore, when adjusting the memory index weights for these two types of false alarm cases, a high adjustment coefficient is set, such as s=0.3; the impact of the associated memory retrieval error on the monitoring results is low, so a low adjustment coefficient is set for the false alarm case of the associated memory retrieval error, such as s=0.1; the weight adjustment process is executed regularly. During the adjustment process, the false alarm frequency and total usage frequency of each memory index entry are first counted, and then the new memory index weight is calculated according to the dynamic adjustment formula. Finally, the memory index weight in the memory index table is updated, and the updated memory index weight is fed back to the large model.

[0014] Compared with the prior art, the present invention has the following beneficial effects: The present application constructs a lightweight memory storage module, which adopts an efficient storage structure and can significantly reduce memory usage while ensuring data storage integrity; and the present application also sets up a dynamic memory recovery mechanism for deleting redundant data to release memory space. When designing the dynamic memory recovery mechanism, a dynamically changing redundant data detection time and redundant data determination cycle are introduced, wherein the redundant data detection time is adjusted according to the changing trend of the number of redundant data detected in the previous two redundant data detection processes, and the redundant data detection time can be automatically dynamically adjusted according to the change in the number of detected data, so that the number in the redundant data set detected each time is relatively balanced, so that the result of the redundant data obtained by the intersection at the end of the final redundant data determination cycle is more accurate; and the monitoring time is adjusted in real time according to the number of redundant data monitored, and when the number of data is detected to drop, the detection time is automatically increased, and the number of redundant data detections in the case of low redundant data volume is reduced, thereby improving the detection efficiency of redundant data and the accuracy of detection results. Accuracy; and the present application sets up a redundant data temporary storage unit to give data that is confirmed as redundant data an opportunity for detection and recovery, and reserves a detection and recovery time of a redundant data determination cycle. After the large model retrieval triggers the timeout mechanism for the retrieval of the associated memory, it can trigger the retrieval of the redundant data stored in the data temporary storage unit. This design provides an opportunity for detection and recovery for redundant data, reduces the probability of misjudgment of redundant data, and reduces the probability of false positives in commercial password detection results due to misjudgment of redundant data; the present application sets a timeout mechanism for the retrieval process and optimizes the retrieval process; the present application uses a dedicated model based on BERT architecture fine-tuning to process commercial password detection data in real time, extract key entities and contextual logical relationships in the data, and the model incorporates a large amount of commercial password detection-related text data during the training process, and has stronger domain adaptability; the present application calculates new weights based on the dynamic adjustment formula, and finally updates the weight values ​​in the memory index table; in this way, the triggering accuracy of the memory index is optimized and the fictitious error rate is reduced. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 This is a flow chart of a method for commercial password detection data identification and association analysis based on a large model of the present invention; Figure 2 This is a schematic diagram of the structure of a lightweight memory storage module for a large-model-based commercial password detection data identification and association analysis method of the present invention; Figure 3 The present invention is a flow chart of a dynamic memory recovery mechanism for a large-model-based commercial password detection data identification and association analysis method. DETAILED DESCRIPTION

[0016] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0017] like Figure 1-Figure 3 As shown, the present invention provides a technical solution, a commercial password detection data identification and association analysis method based on a large model, such as Figure 1 As shown, the method includes the following steps: S1: Build a lightweight memory storage module and set up a dynamic memory recycling mechanism for the memory storage module. When designing the dynamic memory recycling mechanism, set the redundant data confirmation cycle and dynamic detection duration adjustment algorithm, and set up a redundant data temporary storage module; S2: Use a large model to process commercial password detection data in real time, extract key entities and contextual logical relationships in the data, and dynamically generate memory indexes based on the extracted logical relationships; S3. When analyzing a new commercial password detection text, automatically retrieve the associated memory related to the current commercial password detection text in the lightweight memory module through the memory index generated by the large model; S4: Compare the analysis results output by the large model with the associated memory in the lightweight memory storage module. If the results are inconsistent with the information in the memory unit, intercept the results and mark the inconsistent points. S5: Collect false alarm cases, analyze the types and causes of false alarm cases, and optimize the memory index based on the analysis results.

[0018] In step S1: the constructed lightweight memory storage module includes an entity storage unit, a logical relationship storage unit and an index storage unit; the entity storage unit stores key entities extracted from commercial cipher texts through a hash table structure; the logical relationship storage unit stores contextual logical relationships between key entities in the form of triples; the index storage unit stores information in the memory index table through a B+ tree structure; at the same time, a dynamic memory recycling mechanism is set to regularly detect the access frequency of the stored data and clean up redundant data that has not been accessed for a long time, thereby freeing up memory space.

[0019] In step S1: the dynamic memory recycling mechanism performs access detection on the lightweight memory storage module regularly and sets a redundant data confirmation cycle. In one redundant data confirmation cycle, a total of m redundant data detections are performed, and the number of unaccessed data detected each time is recorded. m}, record the data with 0 access times detected in m redundant data detections as a set {A1, A2, ..., A m}, after completing m redundant data detections, find the determined redundant data set A, , where i = 1, 2, ... m, represents the number of m redundant data detections; A represents the set of data determined to be redundant within a redundant data determination cycle; The duration between each redundant data detection within a redundant data determination cycle It is a dynamically changing value. The data access duration selected by the first and second redundant data detection is set to Δt0. Starting from the third redundant data detection, the data access duration selected by each redundant data detection is The change is based on the change in the number of data with 0 access times detected in the previous monitoring, and is calculated according to the following formula: ; Wherein, j=3, 4, ..., m, represents the number of each redundant data detection starting from the third redundant data detection; Indicates the data access duration selected by each redundant data detection starting from the third redundant data detection. Based on the calculated dynamic change value of the data access duration selected by each redundant data detection, the cycle length of the entire redundant data determination cycle can be calculated: ; Wherein, T represents the period length of a redundant data determination period.

[0020] A redundant data temporary storage module is provided. After a redundant data determination cycle ends, the redundant data in the determined redundant data set A is temporarily stored in the redundant data temporary storage module. In the next redundant data determination cycle, the redundant data stored in the redundant data temporary storage module can also be used for comparison with the new commercial cipher text data. The data stored in the redundant data temporary storage module is processed according to the comparison result. The comparison result is analyzed as follows: If the comparison result of the redundant data stored in the redundant data temporary storage module is successful, the successfully matched redundant data is extracted from the redundant data temporary storage module and re-stored into the lightweight memory storage module, and is no longer used as redundant data, thereby completing the detection and recovery of redundant data; If the comparison result of the redundant data stored in the redundant data temporary storage module is a comparison failure, and after the next redundant data determination cycle ends, the redundant data that has not been detected and recovered will be deleted, thereby releasing memory space.

[0021] In step S2: a dedicated model fine-tuned based on the BERT architecture is used to process commercial cryptographic detection data in real time to extract key entities and contextual logical relationships in the data; key entities include the name of the commercial cryptographic algorithm, cryptographic device model, key length, encryption method, cryptographic application scenario and compliance certification status; contextual logical relationships include association relationships, temporal relationships, causal relationships and constraint relationships between entities; a memory index table is dynamically generated based on the extracted key entities and contextual logical relationships; the memory index table takes the key entities as the core, records the logical relationship information corresponding to the entities and the position information in the text, for subsequent rapid retrieval.

[0022] In step S3: the process of detecting new commercial cipher text data includes the following steps: S3-1. After pre-processing the new commercial cipher text data, use the same large model and technology as the key information extraction step to extract key entities in the data. The extracted key entities include entity name, type, and occurrence location; S3-2. Based on the extracted key entities, the B+ tree structure of the memory index table is traversed, and associated memories related to the new commercial password text data in the memory index table are retrieved through information comparison, and the analysis and comparison results between the new commercial password data and the retrieved associations and memories are output; S3-3. Set a timeout mechanism for the retrieval process. If the retrieval time exceeds the set threshold, the current retrieval is terminated and the retrieval results are analyzed and processed to optimize the retrieval process.

[0023] In step S3-3: set a retrieval timeout mechanism for the retrieval process and set a retrieval time threshold Δt for the retrieval process max , compare and analyze the real-time retrieval time Δt′ with the set threshold, and determine whether to trigger the timeout mechanism based on the retrieval results and the real-time retrieval time; If the real-time retrieval time Δt′<Δt max , if it is determined that the associated memory related to the new commercial password text data has been retrieved from the memory index table, the timeout mechanism will not be triggered and the analysis of the retrieval results will be directly entered; If the real-time retrieval time Δt′≥Δt max, if it is determined that no associated memory related to the new commercial password text data is retrieved from the memory index table, or no associated memory that completely matches the new commercial password is retrieved, a timeout mechanism is triggered; and after the timeout mechanism is triggered, a redundant data temporary storage module retrieval is started for the current commercial password text data, and a timeout mechanism is also set in the process of retrieving the redundant data stored in the redundant data temporary storage module; if an associated memory is retrieved in the redundant data temporary storage module, the commercial password text data corresponding to the retrieved associated memory is extracted from the redundant data temporary storage module and re-stored in the lightweight memory storage module.

[0024] In step S4: the analysis and comparison results between the new commercial cryptographic data output by the large model and the retrieved associations and memories include the differences in the core attributes of the key entities, that is, the deviation lengths between the key lengths corresponding to the algorithms, and the logical relationships between the entities; the analysis results output by the large model are compared with the specific contents of the association memories in the lightweight memory storage module, including the deviation lengths between the key lengths corresponding to the algorithms and the logical relationships between the entities; the comparison method adopts a hierarchical comparison strategy, first comparing the core attributes of the key entities, that is, the key lengths corresponding to the algorithms, and then comparing the key logical relationships between the entities; For the detection of core attributes of key entities, a contradiction determination threshold is set. For key length data, the key length contradiction threshold corresponding to the algorithm is set to L0. The deviation length between the key length in the new commercial cipher text extracted in real time by the large model and the key length in the retrieved associated memory is recorded as L. By analyzing and comparing the relationship between L and L0, it is determined whether there is a contradiction. For logical relationship analysis, if there is a direct conflict in the logical relationship between entities, it is directly determined to be a contradiction. The analysis results are as follows: If L≤L0 and there is no direct conflict in the logical relationship, it is determined that the real-time commercial password does not contradict the retrieved associated memory, and the output continues to the subsequent process; If L>L0 or there is a direct conflict in the logical relationship, it is determined that there is a contradiction between the real-time commercial password and the retrieved associated memory; For analysis results that are judged to be contradictory, the analysis results will be intercepted, the output to the subsequent process will be suspended, the contradiction points will be marked through a structured form, and the contradiction information will be temporarily stored in the pending queue, waiting for review by the management personnel; when the management personnel review, they will check all the information in the structured form and judge whether the intercepted result is a false alarm based on their own professional knowledge and relevant materials; after review, if it is judged to be a false alarm, the system will automatically include the case in the training set of the memory index optimization step, and release the interception, allowing the result to enter the subsequent process; if it is judged to be a valid contradiction, it will be processed according to the instructions of the management personnel; the output results of the corrected large model need to be compared with the memory unit again, and only after confirming that there is no contradiction can the subsequent process be entered; the supplemented memory unit information will be updated to the lightweight memory storage module for subsequent comparison and analysis.

[0025] In step S5: false positive cases are collected. The types of false positive cases include correct results that are mistakenly intercepted, incorrect results that are not intercepted, and cases of associated memory retrieval errors. The collected false positive cases are used to automatically adjust the weight of the memory index. The weight calculation adopts the dynamic adjustment formula for calculation: ; in, represents the new memory index weight after adjustment; q represents the memory index weight before adjustment; s represents the set adjustment coefficient, which is related to the degree of influence of false alarms on the detection results; Indicates the false alarm frequency of each memory index entry; Indicates the total usage frequency of each memory index entry; The degree of influence of false alarms on detection results varies according to the type of false alarm cases. The two types of false alarm cases, namely, correct results that were mistakenly intercepted and erroneous results that were not intercepted, have a high degree of influence on detection results. Therefore, a high adjustment coefficient is set when adjusting the memory index weight for these two types of false alarm cases. The degree of influence of associated memory retrieval errors on monitoring results is low. Therefore, a low adjustment coefficient is set for false alarm cases such as associated memory retrieval errors. The weight adjustment process is executed regularly. During the adjustment process, the false alarm frequency and total usage frequency of each memory index entry are first counted, and then the new memory index weight is calculated according to the dynamic adjustment formula. Finally, the memory index weight in the memory index table is updated, and the updated memory index weight is fed back to the large model.

[0026] Example 1: In step S1: the dynamic memory recycling mechanism regularly performs access detection on the lightweight memory storage module and sets a redundant data confirmation cycle. In one redundant data confirmation cycle, a total of m=3 redundant data detections are performed, and the number of unaccessed data detected in each detection is recorded as a set {n1=12, n2=10, n3=8}. The data detected as having 0 access times in the m=3 redundant data detections is recorded as a set {A1, A2, A3}. After the three redundant data detections are completed, the determined redundant data set A is obtained. ; A represents a set of data that is determined to be redundant data within a redundant data determination cycle; The duration between each redundant data detection within a redundant data determination cycle It is a dynamically changing value. The data access duration selected by the first and second redundant data detection is set to Δt0. Starting from the third redundant data detection, the data access duration selected by each redundant data detection is The change is based on the change in the number of data with 0 access times detected in the previous monitoring, and is calculated according to the following formula: ; Wherein, j represents the number of each redundant data detection starting from the third redundant data detection; Indicates the data access duration selected by each redundant data detection starting from the third redundant data detection. Based on the calculated dynamic change value of the data access duration selected by each redundant data detection, the cycle length of the entire redundant data determination cycle can be calculated: =1.07h; ; Wherein, T represents the period length of a redundant data determination period; T=3.07h.

[0027] A redundant data temporary storage module is provided. After a redundant data determination cycle ends, the redundant data in the determined redundant data set A is temporarily stored in the redundant data temporary storage module. In the next redundant data determination cycle, the redundant data stored in the redundant data temporary storage module can also be used for comparison with the new commercial cipher text data. The data stored in the redundant data temporary storage module is processed according to the comparison result. The comparison result is analyzed as follows: If the comparison result of the redundant data stored in the redundant data temporary storage module is successful, the successfully matched redundant data is extracted from the redundant data temporary storage module and re-stored into the lightweight memory storage module, and is no longer used as redundant data, thereby completing the detection and recovery of redundant data; If the comparison result of the redundant data stored in the redundant data temporary storage module is a comparison failure, and after the next redundant data determination cycle ends, the redundant data that has not been detected and recovered will be deleted, thereby releasing memory space.

[0028] In step S3-3: set a retrieval timeout mechanism for the retrieval process and set a retrieval time threshold Δt for the retrieval process max = 60s, compare and analyze the real-time search duration Δt′ with the set threshold, and determine whether to trigger the timeout mechanism based on the search results and the real-time search duration; If the real-time retrieval time Δt′<Δt max , if it is determined that the associated memory related to the new commercial password text data has been retrieved from the memory index table, the timeout mechanism will not be triggered and the analysis of the retrieval results will be directly entered; If the real-time retrieval time Δt′≥Δt max , if it is determined that no associated memory related to the new commercial password text data is retrieved from the memory index table, or no associated memory that completely matches the new commercial password is retrieved, a timeout mechanism is triggered; and after the timeout mechanism is triggered, a redundant data temporary storage module retrieval is started for the current commercial password text data, and a timeout mechanism is also set in the process of retrieving the redundant data stored in the redundant data temporary storage module; if an associated memory is retrieved in the redundant data temporary storage module, the commercial password text data corresponding to the retrieved associated memory is extracted from the redundant data temporary storage module and re-stored in the lightweight memory storage module.

[0029] In step S4: the analysis and comparison results between the new commercial cryptographic data output by the large model and the retrieved associations and memories include the differences in the core attributes of the key entities, that is, the deviation lengths between the key lengths corresponding to the algorithms, and the logical relationships between the entities; the analysis results output by the large model are compared with the specific contents of the association memories in the lightweight memory storage module, including the deviation lengths between the key lengths corresponding to the algorithms and the logical relationships between the entities; the comparison method adopts a hierarchical comparison strategy, first comparing the core attributes of the key entities, that is, the key lengths corresponding to the algorithms, and then comparing the key logical relationships between the entities; For the detection of core attributes of key entities, a contradiction determination threshold is set. For key length data, the key length contradiction threshold corresponding to the algorithm is set to L0 = 32 bytes. The deviation length between the key length in the new commercial cipher text extracted in real time by the large model and the key length of the retrieved associated memory is recorded as L. By analyzing and comparing the relationship between L and L0, it is determined whether there is a contradiction. For logical relationship analysis, if there is a direct conflict in the logical relationship between entities, it is directly judged as a contradiction. The analysis results are as follows: If L≤L0 and there is no direct conflict in the logical relationship, it is determined that the real-time commercial password does not contradict the retrieved associated memory, and the output continues to the subsequent process; If L>L0 or there is a direct conflict in the logical relationship, it is determined that there is a contradiction between the real-time commercial password and the retrieved associated memory.

[0030] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above and that the invention can be embodied in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the invention is defined by the appended claims, not the foregoing description, and all variations within the meaning and range of equivalents of the claims are intended to be included therein. Any reference sign in a claim should not be construed as limiting the claim to which it relates.

Claims

1. A method for identifying and analyzing commercial password detection data based on a large model, characterized by: The method comprises the following steps: S1: Build a lightweight memory storage module and set up a dynamic memory recycling mechanism for the memory storage module. When designing the dynamic memory recycling mechanism, set the redundant data confirmation cycle and dynamic detection duration adjustment algorithm, and set up a redundant data temporary storage module; S2: Use a large model to process commercial password detection data in real time, extract key entities and contextual logical relationships in the data, and dynamically generate memory indexes based on the extracted logical relationships; S3. When analyzing a new commercial password detection text, automatically retrieve the associated memory related to the current commercial password detection text in the lightweight memory module through the memory index generated by the large model; S4: Compare the analysis results output by the large model with the associated memory in the lightweight memory storage module. If the results are inconsistent with the information in the memory unit, intercept the results and mark the inconsistent points. S5: Collect false alarm cases, analyze the types and causes of false alarm cases, and optimize the memory index based on the analysis results.

2. The method for identifying and analyzing commercial password detection data based on a large model according to claim 1, characterized in that: In step S1: the constructed lightweight memory storage module includes an entity storage unit, a logical relationship storage unit and an index storage unit; the entity storage unit stores key entities extracted from commercial cipher texts in a hash table structure; the logical relationship storage unit stores contextual logical relationships between key entities in the form of triples; The index storage unit stores information in the memory index table through a B+ tree structure; at the same time, a dynamic memory recycling mechanism is set up to regularly detect the access frequency of the stored data and clean up redundant data that has not been accessed for a long time, thereby freeing up memory space.

3. The method for identifying and analyzing commercial password detection data based on a large model according to claim 2, characterized in that: In step S1: the dynamic memory recycling mechanism regularly performs access detection on the lightweight memory storage module and sets a redundant data confirmation cycle. In one redundant data confirmation cycle, a total of m redundant data detections are performed, and the number of unaccessed data detected each time is recorded. m }, record the data with 0 access times detected in m redundant data detections as a set {A1, A2, ..., A m }, after completing m redundant data detections, find the determined redundant data set A, , where i = 1, 2, ... m, represents the number of m redundant data detections; A represents the set of data determined to be redundant within a redundant data determination cycle; In a redundant data determination cycle, the data access duration selected by the first and second redundant data detections is set to Δt0; starting from the third redundant data detection, the data access duration selected by each redundant data detection is set to , calculated according to the following formula: ; Wherein, j=3, 4, ..., m, represents the number of each redundant data detection starting from the third redundant data detection; Indicates the data access duration selected by each redundant data detection starting from the third redundant data detection. Based on the calculated dynamic change value of the data access duration selected by each redundant data detection, the cycle length of the entire redundant data determination cycle can be calculated: ; Wherein, T represents the period length of a redundant data determination period.

4. The method for identifying and analyzing commercial password detection data based on a large model according to claim 3, characterized in that: A redundant data temporary storage module is provided. After a redundant data determination cycle ends, the redundant data in the determined redundant data set A is temporarily stored in the redundant data temporary storage module. In the next redundant data determination cycle, the redundant data stored in the redundant data temporary storage module can also be used for comparison with the new commercial cipher text data. The data stored in the redundant data temporary storage module is processed according to the comparison result. The comparison result is analyzed as follows: If the comparison result of the redundant data stored in the redundant data temporary storage module is successful, the successfully matched redundant data is extracted from the redundant data temporary storage module and re-stored into the lightweight memory storage module, and is no longer used as redundant data, thereby completing the detection and recovery of redundant data; If the comparison result of the redundant data stored in the redundant data temporary storage module is a comparison failure, and after the next redundant data determination cycle ends, the redundant data that has not been detected and recovered will be deleted, thereby releasing memory space.

5. The method for identifying and analyzing commercial password detection data based on a large model according to claim 1, characterized in that: In step S2: a dedicated model fine-tuned based on the BERT architecture is used to process commercial cryptographic detection data in real time to extract key entities and contextual logical relationships in the data; the key entities include the name of the commercial cryptographic algorithm, the model of the cryptographic device, the key length, the encryption method, the cryptographic application scenario, and the compliance certification status; the contextual logical relationships include the association relationship, temporal relationship, causal relationship, and constraint relationship between entities; based on the extracted key entities and contextual logical relationships, a memory index table is dynamically generated; The memory index table takes key entities as the core, records the logical relationship information corresponding to the entities and the location information in the text for subsequent rapid retrieval.

6. The method for identifying and analyzing commercial password detection data based on a large model according to claim 1, characterized in that: In step S3: the process of detecting new commercial cipher text data includes the following steps: S3-1. After pre-processing the new commercial cipher text data, use the same large model and technology as the key information extraction step to extract key entities in the data. The extracted key entities include entity name, type, and occurrence location; S3-2. Based on the extracted key entities, the B+ tree structure of the memory index table is traversed, and associated memories related to the new commercial password text data in the memory index table are retrieved through information comparison, and the analysis and comparison results between the new commercial password data and the retrieved associations and memories are output; S3-3. Set a timeout mechanism for the retrieval process. If the retrieval time exceeds the set threshold, the current retrieval is terminated and the retrieval results are analyzed and processed to optimize the retrieval process.

7. The method for identifying and analyzing commercial password detection data based on a large model according to claim 6, characterized in that: In step S3-3: set a retrieval timeout mechanism for the retrieval process and set a retrieval time threshold Δt for the retrieval process max , compare and analyze the real-time retrieval time Δt′ with the set threshold, and determine whether to trigger the timeout mechanism based on the retrieval results and the real-time retrieval time; If the real-time retrieval time Δt′<Δt max , if it is determined that the associated memory related to the new commercial password text data has been retrieved from the memory index table, the timeout mechanism will not be triggered and the analysis of the retrieval results will be directly entered; If the real-time retrieval time Δt′≥Δt max , if it is determined that no associated memory related to the new commercial password text data is retrieved from the memory index table, or no associated memory that completely matches the new commercial password is retrieved, a timeout mechanism is triggered; and after the timeout mechanism is triggered, a redundant data temporary storage module is started to search for the current commercial password text data, and a timeout mechanism is also set during the process of retrieving the redundant data stored in the redundant data temporary storage module; If the associated memory is retrieved in the redundant data temporary storage module, the commercial password text data corresponding to the retrieved associated memory is extracted from the redundant data temporary storage module and re-stored in the lightweight memory storage module.

8. The method for identifying and analyzing commercial password detection data based on a large model according to claim 1, characterized in that: In step S4: the analysis and comparison results between the new commercial cryptographic data output by the large model and the retrieved associations and memories include the differences in the core attributes of the key entities, that is, the deviation lengths between the key lengths corresponding to the algorithms, and the logical relationships between the entities; the analysis results output by the large model are compared with the specific contents of the association memories in the lightweight memory storage module, including the deviation lengths between the key lengths corresponding to the algorithms and the logical relationships between the entities; the comparison method adopts a hierarchical comparison strategy, first comparing the core attributes of the key entities, that is, the key lengths corresponding to the algorithms, and then comparing the key logical relationships between the entities; For the detection of core attributes of key entities, a contradiction determination threshold is set. For key length data, the key length contradiction threshold corresponding to the algorithm is set to L0. The deviation length between the key length in the new commercial cipher text extracted in real time by the large model and the key length in the retrieved associated memory is recorded as L. By analyzing and comparing the relationship between L and L0, it is determined whether there is a contradiction. For logical relationship analysis, if there is a direct conflict in the logical relationship between entities, it is directly determined to be a contradiction. The analysis results are as follows: If L≤L0 and there is no direct conflict in the logical relationship, it is determined that the real-time commercial password does not contradict the retrieved associated memory, and the output continues to the subsequent process; If L>L0 or there is a direct conflict in the logical relationship, it is determined that there is a contradiction between the real-time commercial password and the retrieved associated memory; For analysis results that are judged to be contradictory, the analysis results will be intercepted, the output to the subsequent process will be suspended, the contradiction points will be marked through a structured form, and the contradiction information will be temporarily stored in the pending queue, waiting for review by the management personnel; when the management personnel review, they will check all the information in the structured form to determine whether the intercepted result is a false alarm; after review, if it is judged to be a false alarm, the system will automatically include the case in the training set of the memory index optimization step, and release the interception, allowing the result to enter the subsequent process; if it is judged to be a valid contradiction, it will be processed according to the instructions of the management personnel; the output results of the corrected large model need to be compared with the memory unit again, and only after confirming that there is no contradiction can the subsequent process be entered; the supplemented memory unit information will be updated to the lightweight memory storage module for subsequent comparison and analysis.

9. The method for identifying and analyzing commercial password detection data based on a large model according to claim 1, characterized in that: In step S5: false positive cases are collected. The types of false positive cases include correct results that are mistakenly intercepted, incorrect results that are not intercepted, and cases of associated memory retrieval errors. The collected false positive cases are used to automatically adjust the memory index weight. The weight calculation adopts the dynamic adjustment formula for calculation: ; in, represents the new memory index weight after adjustment; q represents the memory index weight before adjustment; s represents the set adjustment coefficient, which is related to the degree of influence of false alarms on the detection results; Indicates the false alarm frequency of each memory index entry; Indicates the total usage frequency of each memory index entry; The degree of influence of false alarms on detection results varies according to the type of false alarm cases. The two types of false alarm cases, namely, correct results that were mistakenly intercepted and erroneous results that were not intercepted, have a high degree of influence on detection results. Therefore, a high adjustment coefficient is set when adjusting the memory index weight for these two types of false alarm cases. The degree of influence of associated memory retrieval errors on monitoring results is low. Therefore, a low adjustment coefficient is set for false alarm cases such as associated memory retrieval errors. The weight adjustment process is executed regularly. During the adjustment process, the false alarm frequency and total usage frequency of each memory index entry are first counted, and then the new memory index weight is calculated according to the dynamic adjustment formula. Finally, the memory index weight in the memory index table is updated, and the updated memory index weight is fed back to the large model.

Citation Information

Patent Citations

  • Commercial password application compliance evaluation method and device

    CN117714120A

  • Commercial password transformation method and system for power monitoring

    CN119544195A

  • Automatic tax declaration and recheck system

    CN120471722A

  • Automatic programming method based on human-computer interaction

    WO2024021312A1