Network connection management system and method based on Internet of Things equipment
By building a three-level database architecture and a multi-level composite retrieval mechanism, the accuracy and efficiency issues of the IoT device file rejection system are solved, and efficient file classification management and threat identification are achieved.
Patent Information
- Application Number
- CN202510678742.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-05-26
AI Technical Summary
The existing file rejection system of IoT devices has problems such as low rejection judgment accuracy, ineffective reuse of historical data, and lack of classification logic in file storage, which leads to misjudgment, missed judgment and low processing efficiency.
A three-level database architecture is constructed, and a multi-level composite retrieval mechanism is adopted. Through a multi-field joint indexing strategy and a dual verification mechanism of file content relevance/similarity, combined with a dynamic benchmark value judgment algorithm, refined classification management of files is achieved.
It significantly improves file retrieval efficiency, reduces false positives and missed positives, optimizes storage efficiency and system response speed, and enhances adaptability to new threats.
Smart Images

Figure CN120675742A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of Internet of Things security technology, and specifically discloses a network connection management system and method based on Internet of Things devices. Background Art
[0002] Existing document rejection systems typically use a single matching rule, resulting in a simple system architecture and low deployment and maintenance costs. This matching mechanism, based on a static rule base, requires no complex calculations and can complete basic feature comparisons in a short period of time. It doesn't require a multi-level database or historical data analysis framework, and has low requirements for storage and computing resources, making it suitable for lightweight applications. However, this system has the following issues:
[0003] (1) The accuracy of rejection judgment is low, and it is easy to cause misjudgment or omission due to file name disguise or content modification:
[0004] The current system's identification of files mainly relies on basic feature matching, such as fixed suffix names and simple content keywords. It is difficult to deal with file name obfuscation methods, such as the use of deceptive naming such as ".txt.exe" double suffix, Unicode special character disguise, space padding, and content-level evasion techniques, such as nested encrypted compressed packages, text encoding conversion, and fragmentation to hide malicious code; the detection mechanism lacks deep analysis capabilities. For example, insufficient binary feature verification of the true type of the file causes some malicious files to bypass detection through format disguise; at the same time, the judgment logic based on the static rule library cannot effectively identify variant content that has undergone local code obfuscation or semantic replacement, resulting in normal files being intercepted due to accidental rule triggering, or high-risk files being missed due to delayed feature library updates.
[0005] (2) Historical data is not effectively reused, and matching rules cannot be dynamically optimized through similarity analysis:
[0006] The system has not established a framework for quantitative storage and similarity calculation of file features, resulting in massive historical samples being archived in isolated form, making it impossible to extract common features through cluster analysis, such as similar malicious code fragments and high-frequency risk structure patterns; rule iteration relies on manual experience summary and lacks an automated association learning mechanism, making it difficult to identify implicit associations between new attack methods and historical data, such as code reuse features of similar payloads in different attack events; in addition, a time-series-based threat evolution model has not been constructed, making it impossible to predict and intercept in advance the variant files generated by attackers through small-scale incremental modifications, resulting in defense rules lagging behind the evolution of attack technology for a long time.
[0007] (3) File storage lacks classification logic, resulting in low efficiency in subsequent processing:
[0008] After files are stored in the database, a single storage strategy is adopted, such as storing them only by timestamp or upload order. A multi-dimensional classification index is not established based on file attributes, content characteristics, or processing status. This results in subsequent operations, such as batch retrieval of high-risk documents and correlation analysis of homologous attack files, requiring traversal of the entire data set, and response speed decreases sharply as data volume increases. Furthermore, there is no collaborative mechanism with business processes, such as automatically routing suspected sensitive files to dedicated sandboxes for testing, which increases the burden of manual review and makes it difficult to support classification-based priority scheduling, such as prioritizing high-risk files.
[0009] Therefore, it is necessary to invent a network connection management system and method based on Internet of Things devices to solve the above problems. Summary of the Invention
[0010] In order to overcome the above-mentioned defects of the prior art, the present invention provides a network connection management system and method based on Internet of Things devices, which constructs a three-level database architecture; a rejected file library, a threat file library, and a pending review library, and adopts a multi-level composite retrieval mechanism to realize intelligent file classification management; the core process includes: rejected file retrieval based on multi-field joint index, historical matching verification of sender information, and a double verification mechanism of file content relevance / similarity, and innovatively configures a 72-hour automatic cleanup strategy for the threat file database; technical features cover multi-dimensional retrieval fields such as file attributes, terminal identification, IP address, etc., and adopts a notification mechanism that combines data fingerprint feature code with verification QR code, and cooperates with a dynamic benchmark value judgment algorithm to realize risk classification and disposal; effectively solves the problems mentioned in the background technology.
[0011] To achieve the above objectives, the present invention provides the following technical solution: a network connection management system and method based on IoT devices, specifically comprising the following steps:
[0012] S1. The receiving end receives a data text set receiving request and sends the data text set to the cloud management platform; the cloud management platform searches the data text set and performs a composite condition search on the data text set based on a multi-field joint index strategy in the rejected file database;
[0013] S2. If historical data with similar data features is found, a rejection instruction is executed and the data text set is stored in the first database; if historical data with similar data features is not found, a match is performed to match the file name with the sender information in the third database;
[0014] S3. Perform a secondary match based on the primary match result; if the primary match contains historical data with similar data features, perform name-content relevance matching, scan the data text set content, perform relevance matching with the historical rejected file content, and calculate the matching value; if the primary match does not contain similar data features, perform similarity matching, scan the file content, perform content similarity matching on the file content in the first database, and calculate the similarity value;
[0015] S4. If the relevance reaches the benchmark value, the data text set is stored in the third database. If the relevance does not reach the benchmark value, the data text set is further matched for similarity; the data text set is classified according to the size of the similarity value and stored in the database;
[0016] S5, classifying the data text sets for similarity matching according to the size of the similarity values, and storing the data text sets in a database;
[0017] S6. The cloud management platform sends notification information to the receiving end according to the storage location of the data text set.
[0018] Preferably, the multiple fields include file attributes, sending terminal identification, IP address, and digital certificate information; the file attributes include file name, file size, and data text set keywords extracted after scanning the data text set content.
[0019] Preferably, the method of classification according to the size of similarity is to store the data text set in the third database when the similarity value S∈[0,0.3), store the data text set in the second database when S∈[0.3,0.7), and store the data text set in the first database when S∈[0.7,1].
[0020] Preferably, the similarity calculation formula is: In the formula, R represents the similarity value, w represents the keyword, F represents the file name keyword set, C represents the file content word set, TF-IDF(w,T) represents the TF-IDF weight of word w in text T, TF-IDF(w,T) = TF(w,C) × IDF(w,D), D in the formula w represents the number of documents in the database containing keyword w, and N represents the total number of documents in the database; In the formula, C w represents the number of times keyword w appears in content C, and Ctotal represents the total number of words in content C.
[0021] Preferably, the content relevance is calculated as follows: D irepresents the i-th historical file in the database, DB1 represents the rejected file database, V represents the entire vocabulary, and C represents the content of the current file to be analyzed.
[0022] Preferably, the first database is a rejected file database; the second database is a threat file database, which is configured with a data isolation storage area and an automatic cleanup strategy, and the storage period does not exceed 72 hours; the third database is a pending file database.
[0023] Preferably, the threshold setting formula is R 基准 =R 初始 +α·ΔR 误判 +β·ΔR 漏检 +γ·F 人工 , where αβγ are weight coefficients, which are determined by Bayesian optimization. Typical values are α=0.3, β=0.5, γ=0.2, ΔR 误判 Represents the misjudgment rate adjustment item, and the calculation formula is In the formula, A represents the misjudgment rate, ε represents the smoothing factor, such as 0.001, to avoid zero division errors, and ΔR 漏检 Represents the misjudgment rate adjustment item, and the calculation formula is In the formula, B represents the missed detection rate, F 人工 Represents the manual feedback correction term, and the calculation formula is In the formula, Q represents the number of files identified as malicious in the manual review results, P represents the number of files mislabeled as normal, M represents the total number of reviews, and R represents the number of files identified as malicious in the manual review results. 初始 The calculation formula is obtained through statistics from the historical malicious file sample library: R 初始 =μ 恶意 -k·σ 恶意 , in the formula μ 恶意 represents the mean value of the name-content correlation of known malicious files, σ 恶意 represents the standard deviation; k represents the confidence factor, which is determined based on expert advice and historical experience.
[0024] Preferably, the similar data features are a combined feature group of file attribute hash value, sending terminal device fingerprint, and network behavior feature vector.
[0025] Preferably, the notification information includes a data fingerprint feature code, a disposal suggestion, a storage partition identifier and a verification QR code.
[0026] Preferably, the network connection management system based on Internet of Things devices includes a cloud management platform, a receiving end, and a database.
[0027] Technical effects and advantages of the present invention:
[0028] 1. By building a three-level database architecture and a multi-level composite retrieval mechanism, refined classification management of files is achieved. Based on a multi-field joint indexing strategy, retrieval efficiency is significantly improved, high-risk files are quickly identified and isolated, and the cost of manual intervention is reduced.
[0029] 2. A dual verification mechanism combining file content relevance matching and similarity matching, combined with a dynamic threshold judgment algorithm, effectively distinguishes malicious files from misidentified files, reducing the probability of missed detections and misidentifications. This dual verification covers historical data features and real-time content analysis, enhancing adaptability to new threats.
[0030] 3. The threat file database is configured with a 72-hour automatic cleanup policy, combined with data isolation storage areas to prevent redundant data from occupying resources for a long time, while reducing the impact of old threats on the system and significantly optimizing storage efficiency and system response speed. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] The present invention is further described with reference to the accompanying drawings. However, the embodiments in the accompanying drawings do not constitute any limitation to the present invention. A person skilled in the art can obtain other drawings based on the following drawings without creative effort.
[0032] Figure 1 It is a step diagram of the present invention.
[0033] Figure 2 It is the overall flow chart of the present invention. DETAILED DESCRIPTION
[0034] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0035] The present invention provides a network connection management system and method based on Internet of Things devices, which is characterized by specifically comprising: a cloud management platform, a receiving end, and a database;
[0036] In a more specific application of the present invention, the cloud management platform is used to perform multi-dimensional analysis on the received data text set, specifically including: file attributes, sender information, content characteristics, and realize automated risk determination through composite condition retrieval and dynamic matching algorithms; based on the matching results, such as similarity and relevance, a hierarchical storage strategy is triggered to distinguish malicious files, potential threats, and files awaiting review;
[0037] The database adopts a three-level database architecture. The first database is used to permanently store malicious data text sets. The second database is a data isolation storage area for storing data text sets with unknown sources or posing threats. The third database is used to store data text sets that need to be manually reviewed.
[0038] like Figure 1 As shown, the specific implementation of the present invention includes the following steps:
[0039] S1. The receiving end receives a data text set receiving request and sends the data text set to the cloud management platform; the cloud management platform searches the data text set and performs a composite condition search on the data text set based on a multi-field joint index strategy in the rejected file database;
[0040] In the above step S1, the multiple fields include file attributes, sending terminal identification, IP address, and digital certificate information; the file attributes include file name, file size, and data text set keywords extracted after scanning the data text set content;
[0041] After receiving a data text set transmission request initiated by an external IoT device, the receiving end immediately uploads the complete data package, including metadata and content, to the cloud management platform. The transmission uses a TLS encrypted channel to ensure data integrity verification and source authentication. The cloud management platform extracts multi-dimensional features of the data text set, including: file name, file size, MIME type, keyword set, sender IP address, device fingerprint, and digital certificate serial number. Based on the extracted multi-dimensional features, an index is constructed and composite conditional search is performed.
[0042] It should be further explained that the specific implementation of the compound condition search is as follows:
[0043] In the first database, combined Boolean queries of the Elasticsearch cluster are used to perform joint queries on the extracted multi-dimensional feature fields, including file name fuzzy matching, Levenshtein distance ≤ 2, keyword set Jaccard similarity ≥ 0.6, IP address segment attribution matching, CIDR block comparison, device fingerprint exact matching, and real-time certificate blacklist query;
[0044] S2. If historical data with similar data features is found, a rejection instruction is executed and the data text set is stored in the first database; if historical data with similar data features is not found, a match is performed to match the file name with the sender information in the third database;
[0045] It should be further explained that if historical data with similar features is retrieved, the multi-dimensional features of the data text set are extracted and the data text set is stored in the corresponding database field of the database;
[0046] S3. Perform a secondary match based on the primary match result; if the primary match contains historical data with similar data features, perform name-content relevance matching, scan the data text set content, perform relevance matching with the historical rejected file content, and calculate the matching value; if the primary match does not contain similar data features, perform similarity matching, scan the file content, perform content similarity matching on the file content in the first database, and calculate the similarity value;
[0047] In the above step S3, the correlation matching value calculation formula is: In the formula, R represents the similarity value, w represents the keyword, F represents the file name keyword set, C represents the file content word set, TF-IDF(w,T) represents the TF-IDF weight of word w in text T, TF-IDF(w,T) = TF(w,C) × IDF(w,D), D in the formula w represents the number of documents in the database containing keyword w, and N represents the total number of documents in the database; In the formula, C w represents the number of times keyword w appears in content C, and Ctotal represents the total number of words in content C.
[0048] The similarity value calculation formula is: D i represents the i-th historical file in the database, DB1 represents the rejected file database, V represents the entire vocabulary, and C represents the content of the current file to be analyzed.
[0049] Example of calculating correlation value: Scenario:
[0050] File name: Equipment Failure Report;
[0051] Extract keyword F = {equipment, fault, report},
[0052] Content C: total word count 200, equipment appears 2 times, fault appears 5 times, and report appears 0 times;
[0053] Database D: Total number of documents N = 1000, of which: equipment appears in 300 documents, fault appears in 50 documents, and report appears in 200 documents;
[0054] Calculation process:
[0055] TF(equipment, C)=2 / 200=0.01,
[0056] TF(fault, C)=5 / 200=0.025,
[0057] TF(report, C)=0 / 200=0,
[0058] IDF(device,D)=log(1000 / (300+1))≈log(3.32)≈1.20,
[0059] IDF(fault,D)=log(1000 / (50+1))≈log(19.61)≈2.97,
[0060] IDF(report,D)=log(1000 / (200+1))≈log(4.98)≈1.61,
[0061] TF-IDF(device,C)=0.01×1.20=0.012,
[0062] TF-IDF(fault,C)=0.025×2.97=0.074,
[0063] TF-IDF(report,C)=0×1.61=0,
[0064] R=0.012+0.074=0.086,
[0065] Conclusion: The file name and file content are not highly correlated, and similarity matching is required for this data text set;
[0066] Similarity value calculation example:
[0067] Scenario:
[0068] The current file C has the keyword and TF-IDF weight {motor: 0.8, overheating: 1.2, shutdown: 0.9}. Database DB1 contains two historical files: D1: {motor: 1.0, fault: 0.7, fire: 1.4}, D2: {overheating: 0.9, shutdown: 1.1, maintenance: 0.5};
[0069] Calculation process:
[0070] Calculate the similarity between the current file and D1
[0071] Calculate the similarity between the current file and D2
[0072] S=max(0.25,0.78)=0.78,
[0073] Conclusion: The similarity between the current file and the file in the first database is 0.78;
[0074] S4. If the relevance reaches the benchmark value, the data text set is stored in the third database. If the relevance does not reach the benchmark value, the data text set is further matched for similarity; the data text set is classified according to the size of the similarity value and stored in the database;
[0075] It should be further explained that the reference value setting formula is R 基准 =R 初始 +α·ΔR 误判 +β·ΔR 漏检 +γ·F 人工 , where αβγ are weight coefficients, which are determined by Bayesian optimization. Typical values are α=0.3, β=0.5, γ=0.2, ΔR 误判 Represents the misjudgment rate adjustment item, and the calculation formula is In the formula, A represents the misjudgment rate, ε represents the smoothing factor, such as 0.001, to avoid zero division errors, and ΔR 漏检 Represents the misjudgment rate adjustment item, and the calculation formula is In the formula, B represents the missed detection rate, F 人工 Represents the manual feedback correction term, and the calculation formula is In the formula, Q represents the number of files identified as malicious in the manual review results, P represents the number of files mislabeled as normal, M represents the total number of reviews, and R represents the number of files identified as malicious in the manual review results. 初始 It represents the initial baseline value, which is obtained through statistics of historical malicious file sample libraries. The calculation formula is: R 初始 =μ 恶意 -k·σ 恶意 , in the formula μ 恶意 represents the mean value of the name-content correlation of known malicious files, σ 恶意 represents the standard deviation; k represents the confidence factor, which is determined based on expert advice and historical experience.
[0076] Example:
[0077] Example of initial baseline calculation:
[0078] If μ 恶意 =0.8,σ 恶意 =0.1, take k=1,
[0079] Then R 初始 =0.8-1×0.1=0.7;
[0080] Example of calculating manual feedback correction items:
[0081] If 100 files are reviewed, 80 of them are confirmed to be malicious and 10 are mistakenly marked as normal,
[0082] Then Fmanual = (80-10) / 100 = 0.7;
[0083] Benchmark value calculation example:
[0084] R initial = 0.7; recent false positive rate = 1%, missed detection rate = 1%, F manual = 0.7;
[0085] α=0.4,β=0.4,γ=0.2;
[0086] R 基准 =0.7+0.4·ln[1 / (0.01+0.001)]+0.4·-ln[1 / (0.01+0.001)]+0.2·
[0087] 0.7=0.84;
[0088] Conclusion: The adjusted baseline value is 0.84;
[0089] S5, classifying the data text sets for similarity matching according to the size of the similarity values, and storing the data text sets in a database;
[0090] It should be further explained that the classification method according to the size of the similarity is that when the similarity value S∈[0,0.3), the data text set is stored in the third database; when S∈[0.3,0.7), the data text set is stored in the second database; when S∈[0.7,1], the data text set is stored in the first database;
[0091] S6. The cloud management platform sends notification information to the receiving end according to the storage location of the data text set.
[0092] The cloud management platform dynamically generates a notification containing the following key information based on the final storage location of the data text set:
[0093] Data fingerprint feature code: A unique identifier generated based on the hash value or feature vector of the file content, used for subsequent tracking and verification.
[0094] Disposal suggestions: If stored in the first database: recommend that the receiving end immediately terminate the transmission and mark the sender as a high-risk terminal; if stored in the second database: prompt the receiving end to temporarily isolate the file and conduct dynamic monitoring before automatically cleaning it within 72 hours; if stored in the third database: notify the receiving end that manual intervention and review are required;
[0095] Storage partition identification: clearly mark the database to which the file belongs;
[0096] Verification QR code: Contains an encrypted metadata link. After scanning, the recipient can obtain file details, processing logs, or submit a review request.
[0097] The overall flow chart of the invention is as follows: Figure 2 As shown;
[0098] Finally, it should be noted that the above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or make equivalent substitutions for some of the technical features therein. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A network connection management method based on Internet of Things devices, characterized in that: The specific steps include: S1. The receiving end receives a data text set receiving request and sends the data text set to the cloud management platform; the cloud management platform searches the data text set and performs a composite condition search on the data text set based on a multi-field joint index strategy in the rejected file database; S2. If historical data with similar data features is found, a rejection instruction is executed and the data text set is stored in the first database; if historical data with similar data features is not found, a match is performed to match the file name with the sender information in the third database; S3. Perform a secondary match based on the primary match result; if the primary match contains historical data with similar data features, perform name-content relevance matching, scan the data text set content, perform relevance matching with the historical rejected file content, and calculate the matching value; if the primary match does not contain similar data features, perform similarity matching, scan the file content, perform content similarity matching on the file content in the first database, and calculate the similarity value; S4. If the relevance reaches the benchmark value, the data text set is stored in the third database. If the relevance does not reach the benchmark value, the data text set is further matched for similarity; the data text set is classified according to the size of the similarity value and stored in the database; S5, classifying the data text sets for similarity matching according to the size of the similarity values, and storing the data text sets in a database; S6. The cloud management platform sends notification information to the receiving end according to the storage location of the data text set.
2. A network connection management method based on an Internet of Things device according to claim 1, characterized in that: The multiple fields include file attributes, sending terminal identification, IP address, and digital certificate information; the file attributes include file name, file size, and data text set keywords extracted after scanning the data text set content.
3. A network connection management method based on an Internet of Things device according to claim 2, characterized in that: The classification method based on the size of the similarity is that when the similarity value is S∈[0,0.3), the data text set is stored in the third database; when S∈[0.3,0.7), the data text set is stored in the second database; when S∈[0.7,1], the data text set is stored in the first database.
4. A network connection management method based on an Internet of Things device according to claim 1, characterized in that: The similarity calculation formula is: In the formula, R represents the similarity value, w represents the keyword, F represents the file name keyword set, C represents the file content word set, TF-IDF(w,T) represents the TF-IDF weight of word w in text T, TF-IDF(w,T) = TF(w,C) × IDF(w,D), D in the formula w represents the number of documents in the database containing keyword w, and N represents the total number of documents in the database; In the formula, C w represents the number of times keyword w appears in content C, and Ctotal represents the total number of words in content C.
5. The method for managing network connections based on an Internet of Things device according to claim 1, wherein: The content relevance is calculated as follows: In the formula, S represents D i represents the i-th historical file in the database, DB1 represents the rejected file database, V represents the entire vocabulary, and C represents the content of the current file to be analyzed.
6. The method for managing network connections based on an Internet of Things device according to claim 1, wherein: The first database is a rejected file database; the second database is a threat file database, which is configured with a data isolation storage area and an automatic cleanup strategy; and the third database is a pending file database.
7. The method for managing network connections based on an Internet of Things device according to claim 1, wherein: The reference value setting formula is R 基准 =R 初始 +α·ΔR 误判 +β·ΔR 漏检 +γ·F 人工 , where αβγ are weight coefficients, which are determined by Bayesian optimization. Typical values are α=0.3, β=0.5, γ=0.2, ΔR 误判 Represents the misjudgment rate adjustment item, and the calculation formula is In the formula, A represents the misjudgment rate, ε represents the smoothing factor, In the formula, B represents the missed detection rate, F 人工 Represents the manual feedback correction term, and the calculation formula is In the formula, Q represents the number of files identified as malicious in the manual review results, P represents the number of files mislabeled as normal, M represents the total number of reviews, and R represents the number of files identified as malicious in the manual review results. 初始 The calculation formula is obtained through statistics from the historical malicious file sample library: R 初始 =μ 恶意 -k·σ 恶意 , in the formula μ 恶意 represents the mean value of the name-content correlation of known malicious files, σ 恶意 represents the standard deviation; k represents the confidence factor, which is determined based on expert advice and historical experience.
8. The method for managing network connections based on IoT devices according to claim 1, wherein: The similar data features are a combined feature group of file attribute hash value, sending terminal device fingerprint, and network behavior feature vector.
9. The method for managing network connections based on IoT devices according to claim 1, wherein: The notification information includes a data fingerprint feature code, disposal suggestions, storage partition identifier and verification QR code.
10. The network connection management system based on Internet of Things devices according to claim 1, characterized in that: Including cloud management platform, receiving end, and database.
Citation Information
Patent Citations
Data set retrieval method and system
CN111026710A
Multi-thread data retrieval method based on AI technology and access method of retrieved data
CN113742292A
Database joint index establishment method and device, equipment and storage medium
CN114817243A
Similar file detection method and system, electronic equipment and storage medium
CN115145872A
File retrieval method and system based on OCR (Optical Character Recognition) technology
CN117390214A