Address standardization and dynamic management method
By establishing a hierarchical entry framework and natural language processing technology, combining geographic information database and dynamic encryption strategy, the problem of inaccurate standardization and insufficient security protection of address information is solved, and the accurate standardization of address information and safe and reliable transmission management is achieved.
Patent Information
- Application Number
- CN202510865284.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-06-26
AI Technical Summary
There are problems of inaccurate standardization and insufficient security protection in the processing of existing address information, especially in the chaos in data structure caused by the diversity and irregularity of address information, and the security and integrity in the transmission process are difficult to guarantee.
By establishing a hierarchical entry framework, using natural language processing technology to perform word segmentation and matching, constructing authentic vectors and authentic attribute vectors, combining with geographical information databases for semantic association, identifying unreliable elements and correcting them, and dynamically adjusting encryption strategies based on historical security data and real-time risk assessment.
It realizes accurate and standardized processing of address information and safe and reliable transmission management, effectively solves the problems of fuzzy addresses and incorrect expressions, and ensures the security and integrity of the data.
Smart Images

Figure CN120371939A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of information processing, and particularly to an address standardization and dynamic management method. Background Art
[0002] In the digital age, address information, as key data, is widely used in various business scenarios. However, due to the diversity and non-standardization of address expressions, as well as possible errors in the data collection process, the original address information often has many problems, such as unclear hierarchical division, chaotic data structure, incorrect expressions, etc., seriously affecting the effective utilization and processing efficiency of address information.
[0003] Most of the existing address standardization methods adopt simple rule matching or templatized processing, which are difficult to cope with complex and changeable address expression forms, unable to accurately identify and process fuzzy and non-standard address information, resulting in inaccurate standardization results. At the same time, in the process of address information transmission and management, traditional technologies lack effective security protection mechanisms and cannot dynamically adjust protection strategies according to security risks at different time periods, making it difficult to ensure the security and integrity of address information during transmission and storage.
[0004] Therefore, the present invention proposes an address standardization and dynamic management method. Summary of the Invention
[0005] The present invention provides an address standardization and dynamic management method to solve the problems of inaccurate standardization and insufficient security protection in existing address information processing, achieve precise standardization processing of address information, and ensure the secure and reliable transmission and management of address information through a dynamic security protection mechanism.
[0006] An address standardization and dynamic management method of the present invention includes: Step 1: Input the collected address information into a hierarchical dictionary structure to perform hierarchical division on the address information, temporarily store the un-successfully divided part in a first storage block, and temporarily store the successfully divided part in a second storage block; Step 2: Convert the unstructured data in the second storage block into initial structured data according to the conversion standard of each division level; Step 3: Split the initial structured data according to the expression combination of the corresponding division level, and respectively judge the authenticity of each split structured data with the corresponding expression standard under the expression combination to construct a authenticity vector and a authenticity attribute vector, where the elements of the authenticity vector and the authenticity attribute vector correspond one by one; Step 4: Analyze the existence reliability of the same corresponding elements in the authenticity vector and the authenticity attribute vector, and screen out unreliable elements depending on the existence reliability; Step 5: Establish an association relationship between the unreliable element and the unpartitioned part of the first storage block, and retrieve the historical standard level description depending on the level to which the unreliable element belongs, and correct the expression corresponding to the unreliable element to obtain a standardized address, where the standardized address is obtained by sequentially structuring the description from a high level to a low level, and the corresponding last level is the lowest level in the collected address information; Step 6: Determine the historical security information of each transmission interface for transmitting the standardized address and the receiving interface of the big data platform during a specified time period, and establish a security protection mechanism and dynamically manage it.
[0007] Preferably, input the collected address information into a hierarchical entry framework to perform hierarchical division on the address information, including: Establish a hierarchical entry framework, where each level contains corresponding standard entries and a thesaurus, and the hierarchical entry framework includes at least country, province, city, district, county, street, community / village, building, unit, floor, and house number; Use natural language processing technology to perform word segmentation on the collected address information to obtain a number of independent words; Match each independent word with the standard entries and thesaurus in the hierarchical entry framework, and perform the matching sequentially from a high level to a low level to obtain an independent word group at the level where the matching is successful; Determine the remaining words that have not been successfully matched among the number of independent words, and respectively semantically associate each remaining word with the independent word group at the level where the matching is successful, and combine the relevant data in the geographic information database to determine the possible levels to which they belong; Temporarily store the independent word group at the level where the matching is successful in the second storage block; Temporarily store the independent words at the possible levels to which they belong in the first storage block.
[0008] Preferably, construct a true / false vector and a true / false attribute vector, including: Determine the first quantity of the split structure data and the combined quantity of the expression combinations at the corresponding division level to obtain a first matching coefficient; According to the second matching coefficient between each split structure data in the corresponding division level and each expression standard in the expression combination, and determine whether the largest second matching coefficient is 1; If it is 1, determine that the corresponding split structure data is true; Otherwise, determine that the corresponding split structure data is false, and analyze all the second matching coefficients involved, and combine the historical matching failure event set based on each expression standard at the corresponding division level to determine the true / false attribute of the corresponding split structure data, where the true / false attribute is related to the possible reasons for the corresponding split structure data being false; Construct a true / false vector based on the largest second matching coefficient corresponding to each split structure data. Meanwhile, construct a true / false attribute vector based on the true / false attribute corresponding to each split structure data.
[0009] Preferably, determining the true / false attribute corresponding to the split structure data includes: Input each historical matching failure event in the historical matching failure event set into an event analysis model to obtain a first failure factor. Meanwhile, retrieve the matching process log of each historical matching failure event from the historical diary record database and input it into a log analysis model to obtain a second failure factor, and establish a failure factor table corresponding to the historical matching failure event; Conduct an aggregation analysis on all failure factor tables to establish a complete description framework; Perform a unique matching judgment on all the second matching coefficients corresponding to the split structure data that is false, that is, judge whether the second quantity of the second matching coefficient with a value of 0 is N - 1. If so, regard the second matching coefficient that is not 0 as the third matching coefficient, and combine the expression standards of the first matching coefficient and the third matching coefficient to construct a search index, where N represents the number of all second matching coefficients under the corresponding split structure data; If not, determine the coefficient ratio of the second matching coefficient that is not 0, and combine the expression standard under the second matching coefficient that is not 0 and the first matching coefficient to construct a search index; Based on the search index, search and locate the possible causes in the complete description framework to obtain the true / false attribute corresponding to the split structure data.
[0010] Preferably, analyze the existence reliability of the same corresponding elements in the true / false vector and the true / false attribute vector, and filter out unreliable elements depending on the existence reliability, including: Determine the error factor of each true / false attribute;
[0011] Among them, is the error factor of the w-th true / false attribute; n represents the number of categories of possible causes corresponding to the w-th true / false attribute; represents the total number of categories of all possible causes; represents the category weight of the i-th category of reasons under the w-th true / false attribute; represents the error probability that the i-th category of reasons under the w-th true / false attribute causes non-structured data to be converted into structured data; Determine the existence reliability of the true / false value and the error factor of the same corresponding element;
[0012] Among them, is the existence reliability of the element corresponding to the w-th authenticity attribute; represents the authenticity value of the element corresponding to the w-th authenticity attribute; If the existence reliability is less than the preset reliability, the corresponding split structure data is regarded as an unreliable element.
[0013] Preferably, the description corresponding to the unreliable element is corrected, including: Extract independent words consistent with the corresponding level from the unpartitioned part of the first storage block according to the level to which the unreliable element belongs, and regard them as the first words; Obtain the association relationship according to the semantic relevance, grammatical structure of the description of the unreliable element and each first word, and the position and association of the unreliable element and the first word in the address logic framework judged according to the overall information logic relationship set by the corresponding partition level; Based on the historical standard level description, determine the historical relevant correction rules, and combine the association relationship to correct the description corresponding to the unreliable element.
[0014] Preferably, a security protection mechanism is established, including: Analyze the historical security information of each transmission interface and receiving interface in a specified time period, determine the set of danger types and the danger probability distribution of each danger type in the corresponding time period, and obtain the information danger combination at each time point; Extract the input behavior sequence of the login end of the corresponding transmission interface at each time point of the receiving interface and each transmission interface in a specified time period, and obtain the behavior danger combination at each time point; Generate the first security evaluation value, the second security evaluation value and the comprehensive security evaluation value according to the information danger combination and behavior danger combination at the same time point, and screen and retain the maximum security evaluation value; Draw a curve for the retained maximum security evaluation value in the order of time points and divide the value level according to the security protection standard to obtain several curve segments; Generate a security encryption factor according to the current security configuration information of the login end of the transmission interface involved in each curve segment and in combination with the value level, and set the encryption method for the corresponding continuous sub-time period according to the security encryption factor, where the encryption method is to perform information coding scrambling on the input address information based on the chaos level of the security encryption factor; Sort all encryption methods in the order of continuous sub-time periods to obtain the initial protection mechanism.
[0015] Preferably, after obtaining the initial protection mechanism, it further includes: Determine the segment length of each continuous sub-time period and the resource consumption for changing the encryption method in adjacent continuous sub-time periods; If each resource loss is less than the preset loss, consider the initial protection mechanism as a secure protection mechanism; Otherwise, lock the first sub - time period where the resource loss is greater than or equal to the preset loss. If the security encryption factor of the first sub - time period is greater than that of the previous sub - time period, continue to keep the corresponding encryption method unchanged; Otherwise, change the encryption method of the first sub - time period to that of the previous sub - time period to obtain a secure protection mechanism.
[0016] Compared with the prior art, the beneficial effects of the present application are as follows: Through multi - level matching, authenticity judgment, and intelligent correction, problems such as fuzzy addresses and incorrect expressions can be effectively solved. Based on historical security data and real - time risk assessment, dynamic adjustment of the encryption strategy is realized to ensure data security.
[0017] Other features and advantages of the present invention will be described in the subsequent specification, and, in part, will become apparent from the specification or will be understood by implementing the present invention. The objectives and other advantages of the present invention can be achieved and obtained by the structures specifically pointed out in the written specification and the drawings.
[0018] The technical solution of the present invention will be further described in detail below through the drawings and embodiments. Description of the Drawings
[0019] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention and do not constitute a limitation to the present invention. In the drawings: Figure 1 It is a flowchart of a method for address standardization and dynamic management in an embodiment of the present invention. Detailed Embodiments
[0021] The following describes the preferred embodiments of the present invention with reference to the drawings. It should be understood that the preferred embodiments described herein are only for the purpose of explaining and illustrating the present invention and are not used to limit the present invention.
[0022] A method for address standardization and dynamic management of the present invention, as Figure 1 shown, includes: Step 1: Input the collected address information into the hierarchical entry framework to hierarchically divide the address information. Temporarily store the parts that are not successfully divided in the first storage block and the parts that are successfully divided in the second storage block; Step 2: Convert the non - structured data in the second storage block into initial structured data according to the conversion standard of each division level; Step 3: Split the initial structure data according to the corresponding hierarchical expressions, and judge the authenticity of each split structure data against the corresponding expression standards under the expression combination, and construct a authenticity vector and a authenticity attribute vector, where the elements of the authenticity vector and the authenticity attribute vector correspond one by one; Step 4: Analyze the existence reliability of the same corresponding elements in the authenticity vector and the authenticity attribute vector, and filter out unreliable elements depending on the existence reliability; Step 5: Establish an association relationship between the unreliable elements and the unsuccessfully partitioned part in the first storage block, and retrieve the historical standard hierarchical description depending on the level to which the unreliable elements belong, and correct the expressions corresponding to the unreliable elements to obtain a standardized address, where the standardized address is obtained by sequentially structuring the descriptions from the high level to the low level, and the corresponding last level is the lowest level in the collected address information; Step 6: Determine the historical security information of each transmission interface for transmitting the standardized address and the receiving interface of the big data platform during a specified time period, and establish a security protection mechanism and dynamically manage it.
[0023] In this embodiment, the hierarchical entry framework is a multi-level entry system constructed according to the address administrative level or spatial logic.
[0024] In this embodiment, the hierarchical division is a process of mapping address information to the corresponding administrative or spatial levels according to the hierarchical entry framework. For example, the address No. 88, Jianguo Road, Chaoyang District, Beijing is divided into: country (China), city (Beijing), district (Chaoyang), street (Jianguo Road), house number (No. 88). The bidirectional longest matching algorithm is combined with a rule engine to first match the high level (such as country / province), and then match layer by layer downward, and regular expressions are used to identify digital features such as house numbers and building numbers.
[0025] In this embodiment, the first storage block stores the address segments that are not successfully matched in the hierarchical division (such as fuzzy place names, misspelled parts). For example, the address "Shanghai Puming District Zhong Puming District" is an unsuccessfully divided part and is stored in the first storage block (actually it should be Pudong New Area). The original undivided text is stored using a distributed file system (such as HDFS), and fast retrieval is achieved by combining an indexing technology (such as Elasticsearch).
[0026] In this embodiment, the second storage block stores the structured address data that is successfully matched in the hierarchical division. For example, after the address "Nanshan District, Science and Technology Park South Area, Shenzhen" is successfully divided, "Shenzhen - Nanshan - Science and Technology Park South Area" is stored in the second storage block. A relational database (such as PostgreSQL) is used to store it structurally according to hierarchical fields (country, province, city, etc.), supporting SQL queries and transaction management.
[0027] The conversion standard is the mapping rule for address data at each level from unstructured text to structured fields. For example, the conversion standard at the street level stipulates that XX Road, No. XX needs to be split into the street name and house number fields; the conversion standard at the community level stipulates that XX Garden, Building XX needs to be split into the community name and building number. Using a regular expression template library (such as the Python re module) combined with a rule engine, preset standard splitting patterns for different levels are supported, and custom template extension is also supported.
[0028] Unstructured data is address text that is not organized in a fixed format (such as a freely entered address description). For example, when a user enters "No. 3, Building 3, Zhongguancun South Street, Haidian District, Beijing", it is unstructured data.
[0029] Use a text parser (such as ANTLR) to identify address components in natural language, and combine part-of-speech tagging (such as the NLTK library) to distinguish elements such as place names and numbers.
[0030] In this embodiment, the initial structured data is the fielded address data formed after splitting according to the conversion standard (the authenticity has not been verified yet). For example, the unstructured data "No. 1, Huaxia Road, Zhujiang New Town, Tianhe District, Guangzhou" is converted into the initial structured data: {City: Guangzhou, District: Tianhe, Street: Huaxia Road, House Number: 1}. The structured fields are stored in JSON / XML format, and automatic conversion from text to object is achieved through a data mapping tool (such as MapStruct).
[0031] The expression combination is the standard expression combination at the same address level (for example, the street level may include expression methods such as road, street, avenue, etc.). For example, the expression combination at the street level is {XX Road, XX Street, XX Avenue, XX Lane}. Statistically analyze the high-frequency expression patterns from the historical address database, generate a combination set through a clustering algorithm (such as K-means), and update it regularly according to new data.
[0032] In this embodiment, the expression standard is each legal address expression paradigm in the expression combination.
[0033] The authenticity judgment is the process of judging whether the split structured data conforms to the corresponding expression standard. For example, when comparing the split data "Puming District" with the district expression standard [correct administrative division name], since Puming District does not exist, it is judged as false; "Pudong New Area" conforms to the standard and is judged as true. Use a rule engine (such as Drools) to perform standard matching, and combine a semantic similarity algorithm (such as cosine similarity) to handle fuzzy matching scenarios.
[0034] The true / false vector is a vector composed of the maximum matching coefficients of each split structure data (element value ∈ [0, 1]). For example, after splitting the address, 3 structure data are obtained, and the matching coefficients are 0.95, 0.1, and 1 respectively. Then the true / false vector is [0.95, 0.1, 1]. Use a numpy array to store the vector, which supports vectorized calculations (such as matrix multiplication) to improve calculation efficiency.
[0035] In this embodiment, the true / false attribute vector is a vector that records the specific error reasons for each piece of false data (the elements are string descriptions). Use a List collection to store the string attributes, which is aligned with the index of the true / false vector and supports fast attribute queries.
[0036] In this embodiment, the existence reliability is a comprehensive reliability index calculated by combining the true / false value and the error factor.
[0037] In this embodiment, the historical standard level description is a database that records the historical correct address expressions and correction rules for each level. For example, the historical standard description at the street level includes rules such as Jianguo Road → correct name, wrong road name → commonly miswritten as XX Road, etc. Use a graph database (such as Neo4j) to store the association relationship between the standard expression and the error pattern, which supports semantic similarity queries.
[0038] In this embodiment, the association relationship is the connection between unreliable elements and the unpartitioned part in terms of semantics, syntax, and logic. For example, the unreliable element "wrong name of the community" has a semantic association with the unpartitioned part "XX Garden" (referring to the same community), forms a modification relationship in syntax, and belongs to the same level logically.
[0039] Implementation means: Use dependency syntax analysis (such as StanfordCoreNLP) to parse the syntactic relationship, calculate the semantic association degree through word vector similarity (such as Sentence-BERT), and determine the logical relationship by combining the address level logic rules.
[0040] In this embodiment, the standardized address is a standardized address expression arranged in the order from high level to low level (the lowest level is the lowest level in the original address). For example, the lowest level of the original address is the building number, and the standardized address format is country - province - city - district - street - community - building number. Use a template engine (such as Velocity) to splice the standardized fields in the order of levels, which supports custom output formats (such as Chinese full pinyin, English abbreviation, etc.).
[0041] In this embodiment, the historical security information is the security logs (such as attack types, vulnerability records) of the transmission interface and the receiving interface within a specified time period. For example, an interface has suffered 3 SQL injection attacks and 2 identity authentication failures in the past 7 days. The interface security events are collected through a log collection system (such as ELK Stack), and the time series data is stored using a time series database (such as InfluxDB).
[0042] The beneficial effects of the above technical solution are as follows: through multi-level matching, authenticity judgment, and intelligent correction, problems such as fuzzy addresses and incorrect expressions can be effectively solved. Based on historical security data and real-time risk assessment, the dynamic adjustment of encryption policies is realized to ensure the security of data.
[0043] A method for address standardization and dynamic management of the present invention inputs the collected address information into a hierarchical entry framework to perform hierarchical division on the address information, including: Establish a hierarchical entry framework, where each level contains corresponding standard entries and a thesaurus. The hierarchical entry framework includes at least country, province, city, district, county, street, community / village, building, unit, floor, and house number; Use natural language processing technology to perform word segmentation on the collected address information to obtain a number of independent words; Match each independent word with the standard entries and thesaurus in the hierarchical entry framework, and perform the matching in sequence from the high level to the low level to obtain an independent word group at the level where the matching is successful; Determine the remaining words that have not been successfully matched among the number of independent words, and respectively semantically associate each remaining word with the independent word group at the level where the matching is successful, and combine the relevant data of the geographic information database to determine the possible levels to which they belong; Temporarily store the independent word group at the level where the matching is successful in the second storage block; Temporarily store the independent words at the possible levels to which they belong in the first storage block.
[0044] In this embodiment, a knowledge graph technology is used to construct the logical structure of the hierarchical entry framework. Taking the level as a node, the standard entry and thesaurus as node attributes, the association relationship between nodes is established. The standard entries are extracted from authoritative address data sources. Using word vector models in natural language processing, such as Word2Vec, BERT, etc., the semantic similarity between words is calculated to generate a thesaurus. The entry framework is updated and maintained regularly. According to administrative division adjustments, the emergence of new place names, etc., the standard entries and synonyms are supplemented and modified in a timely manner.
[0045] In this embodiment, during the matching process, when an independent word successfully matches a standard word at a certain level in the hierarchical word structure or a word in the synonym library, a set composed of the independent word and other independent words that are successfully matched at the same level subsequently is formed. The method combines sequential traversal and the bidirectional longest matching algorithm. Starting from the national level, each independent word is sequentially subjected to string matching and semantic similarity calculation with the standard words and the synonym library at this level. If the similarity reaches the set threshold (such as 0.8), it is considered a successful match, and the matching of the next independent word at this level continues; if all words fail to match at a certain level, the next level is entered for continued matching. The caching technology is used to store the matched words and results, reducing repeated calculations and improving the matching efficiency.
[0046] In this embodiment, the remaining words that fail to match successfully are independent words that do not successfully match any standard words and the synonym library at any level during the matching process with the hierarchical word structure. These words may fail to match due to reasons such as non-standard expressions, typos, newly emerged place names, etc.
[0047] Semantic association is to analyze the semantic connection between the remaining words that fail to match successfully and the group of independent words at the level that has been successfully matched, and judge whether the remaining words may be more detailed address information at the level that has been matched, or have semantic relationships such as inclusion or adjacency with the level that has been matched.
[0048] The geographic information database is a database that stores geographic spatial data and related attribute information, including administrative division boundaries, road networks, building locations, etc. By querying the geographic information database, the spatial location relationship and geographic features of address information can be obtained to assist in judging the level to which the unmatched words belong.
[0049] The beneficial effects of the above technical solution are as follows: By constructing a hierarchical word structure and applying natural language processing technology, efficient hierarchical division of address information is achieved. It can accurately identify the hierarchical attribution of most standard address information, store the successfully matched part of the address in the second storage block, providing a clear data basis for subsequent address standardization and structuring processing. For the part of the address that fails to match successfully, through semantic association analysis and the assistance of the geographic information database, the possible level to which it belongs is determined and stored in the first storage block, providing a direction for further address correction and improvement. Overall, the accuracy and efficiency of address information processing are improved, and the problem of unclear hierarchical division of address information is effectively solved.
[0050] A method for address standardization and dynamic management of the present invention constructs a true / false vector and a true / false attribute vector, including: Determine the first quantity of the split structure data and the combined quantity of the expression combinations at the corresponding divided level to obtain the first matching coefficient; According to the second matching coefficient of each split structure data in the corresponding division level and each expression standard under the expression combination, and determine whether the largest second matching coefficient is 1; If it is 1, determine that the corresponding split structure data is true; Otherwise, determine that the corresponding split structure data is false, analyze all the involved second matching coefficients, and combine the historical matching failure event set based on each expression standard at the corresponding division level to determine the true and false attribute of the corresponding split structure data, where the true and false attribute is related to the possible reasons for the corresponding split structure data being false; Construct a true and false vector based on the largest second matching coefficient corresponding to each split structure data. At the same time, construct a true and false attribute vector from the true and false attributes corresponding to each split structure data.
[0051] In this embodiment, the split structure data is a data segment obtained by further splitting the address information that has been hierarchically divided and converted into initial structure data according to the expression combination rule of the corresponding division level.
[0052] The first quantity refers to the total number of split structure data at the corresponding division level.
[0053] The expression combination is a set formed by combining various possible expression ways in each address division level to cover the diversity of address expressions at this level.
[0054] The combination quantity is the total number of expression ways included in the expression combination at the corresponding division level.
[0055] In this embodiment, the first matching coefficient = the first quantity / the combination quantity.
[0056] In this embodiment, the expression standard is a specific expression form that conforms to the specification, is correct, and can accurately express the address information in each expression combination. The second matching coefficient is a quantitative index used to measure the matching degree between each split structure data and each expression standard under the expression combination, and its value range is usually between 0 and 1. The closer the value is to 1, the higher the matching degree; a value of 1 indicates a perfect match.
[0057] The historical matching failure event set is a set recording all the matching failure cases in the process of matching split structure data with expression standards in the past at the corresponding division level. Each event contains detailed information such as the failed split structure data, the corresponding expression standard, and the reason for failure.
[0058] In this embodiment, for example, the constructed true and false vector is [0.9, 0.3, 1], and the true and false attribute vector is [incomplete expression, incorrect administrative division name, no error].
[0059] The beneficial effects of the above technical solution are as follows: By introducing the first matching coefficient and the second matching coefficient, the split structure data is analyzed from two dimensions of quantity and similarity. Combining with the historical matching failure event set, the accurate judgment of the authenticity of the split structure data of the address information is realized, and the specific reasons for the false data are determined. The constructed authenticity vector and authenticity attribute vector provide quantitative and structured data support for the subsequent comprehensive evaluation, error correction and data optimization of the address information, can effectively identify errors and non-standard expressions in the address information, and provide a strong guarantee for the address standardization process.
[0060] A method for address standardization and dynamic management of the present invention determines the authenticity attribute of the corresponding split structure data, including: Input each historical matching failure event in the historical matching failure event set into an event analysis model to obtain a first failure factor. At the same time, retrieve the matching process log of each historical matching failure event from the historical diary record database, and input it into a log analysis model to obtain a second failure factor, and establish a failure factor table corresponding to the historical matching failure event; Perform aggregation analysis on all failure factor tables to establish a complete description framework; Perform a unique matching judgment on all the second matching coefficients corresponding to the false split structure data, that is, judge whether the second quantity of the second matching coefficient with a value of 0 is N - 1. If so, regard the second matching coefficient that is not 0 as the third matching coefficient, and combine the first matching coefficient and the expression standard of the third matching coefficient to construct a search index, where N represents the quantity of all the second matching coefficients under the corresponding split structure data; If not, determine the coefficient ratio of the second matching coefficient that is not 0, and combine the expression standard under the second matching coefficient that is not 0 and the first matching coefficient to construct a search index; Based on the search index, search and locate the possible causes in the complete description framework to obtain the authenticity attribute of the corresponding split structure data.
[0061] In this embodiment, the event analysis model is a model constructed based on machine learning or data analysis algorithms, which is used to analyze and extract features from historical matching failure events, and dig out the key factors leading to the matching failure from various information of the events, that is, the first failure factor. And the first failure factor is the key failure factor obtained by the event analysis model processing the historical matching failure events. For example, incorrect expression of administrative division names lacks key information, etc. These factors are a high-level summary of the failure events.
[0062] The historical diary record database is a database specifically used to store the detailed logs of the address information matching process. The log content includes the algorithm steps adopted during the matching, the intermediate results of data processing, the details of the matching attempts with various expression standards, etc. The matching process log is a detailed record of the specific process of each address information matching operation, which completely presents the whole process information from the input data to the obtained matching result.
[0063] The log analysis model is a model for parsing and analyzing the matching process log. By sorting out and mining the data in the log, the deep-seated reasons for the matching failure are found, and the second failure factor is obtained.
[0064] The second failure factor is the failure reason extracted by the log analysis model from the matching process log. For example, when using the cosine similarity algorithm for calculation, the similarity calculation error is caused by the mismatch of the word vector dimensions.
[0065] The failure factor table is a table formed by integrating the first failure factor and the second failure factor corresponding to each historical matching failure event. Taking the event as the unit, it clearly presents the key failure factor combinations of each failure event.
[0066] Aggregate analysis is to summarize, classify, statistically analyze and analyze the data in multiple failure factor tables, find the commonalities and laws between different failure events, mine the common patterns and key factor sets leading to the matching failure. The complete description framework is established through aggregate analysis, which is a structured framework that can comprehensively and systematically describe various matching failure situations and their reasons. It classifies and integrates the failure reasons to form a hierarchical and logically clear system, which is convenient for subsequent query and use. The clustering algorithm in data mining (such as K-means clustering) is used to cluster the failure factors in the failure factor table, and the similar failure reasons are grouped into one category. Then, combined with statistical analysis methods, calculate indicators such as the frequency and proportion of each type of failure reason. Based on the clustering and statistical results, use a tree structure or a graph structure to construct a complete description framework, and clearly present various failure reasons and their relationships. For example, using knowledge graph technology, taking the failure reason category as the node and the association relationship between categories as the edge, construct a visual complete description framework. For example, when performing aggregate analysis on a large number of failure factor tables, it is found that reasons such as misspelling of administrative division names and non-standard abbreviations of street names appear frequently, and they can be aggregated into the category of incorrect address name expressions; reasons such as improper handling of special characters by the word segmentation algorithm and incorrect setting of semantic similarity calculation parameters can be aggregated into the category of algorithm processing problems. Through such classification and integration, a complete description framework is established, including major categories such as incorrect address name expressions and algorithm processing problems, as well as various sub-categories of failure reasons.
[0067] The third matching coefficient is in the unique matching judgment. When the second quantity of the second matching coefficient with a value of 0 is N - 1, the non-zero second matching coefficient is regarded as the third matching coefficient alone, which represents the highest matching degree of the split structure data with a certain expression standard.
[0068] The search index is constructed based on the correlation coefficient and the expression standard, and is used to quickly locate and find the possible failure reasons of the corresponding split structure data in the complete description framework. It is similar to the table of contents of a book and can help the system efficiently find the target information. For example, for the split structure data of Fuzhou District (district and county level) which is false, the second matching coefficients obtained by matching with the expression standards such as Fuzhou City, Xiamen City, and Quanzhou City are 0, 0, and 0.2 respectively. At this time, N = 3, and the second quantity is 2 (two 0-value coefficients), meeting the N - 1 condition, and 0.2 is the third matching coefficient. Combine the first matching coefficient and the expression standard of Fuzhou City to construct the search index. If the second matching coefficients are 0.1, 0.3, and 0.2 respectively, the unique matching condition is not met. Calculate the proportion of each non-zero coefficient (such as 0.1 accounts for 20%, 0.3 accounts for 60%, and 0.2 accounts for 20%), select the highest proportion of 0.3 and its corresponding expression standard, and combine the first matching coefficient to construct the search index. Use the query statement of the database (such as the SELECT statement in SQL) to query in the database storing the complete description framework according to the keywords of the search index to obtain the corresponding authenticity attribute.
[0069] The beneficial effects of the above technical solution are: through in-depth analysis of historical matching failure events and multi-dimensional data mining, a complete and systematic description framework of failure reasons is established, and a precise search index is constructed in combination with the matching coefficient characteristics of the split structure data, which can efficiently and accurately determine the authenticity attribute of the split structure data, providing a detailed and reliable basis for the subsequent correction, optimization of address data, and address standardization processing.
[0070] A method for address standardization and dynamic management of the present invention analyzes the existence reliability of the same corresponding elements in the authenticity vector and the authenticity attribute vector, and filters unreliable elements depending on the existence reliability, including: Determine the error factor of each authenticity attribute;
[0071] Among them, is the error factor of the wth authenticity attribute; n represents the number of categories of possible occurrence reasons corresponding to the wth authenticity attribute; represents the total number of categories of all possible occurrence reasons; represents the category weight of the ith category reason under the wth authenticity attribute; Denote the error probability of converting unstructured data to structured data caused by the $i$-th category reason under the $w$-th authenticity attribute; Determine the existence reliability of the authenticity value and the error factor of the same corresponding element;
[0072] where, is the existence reliability of the element corresponding to the $w$-th authenticity attribute; Denote the authenticity value of the element corresponding to the $w$-th authenticity attribute; If the existence reliability is less than the preset reliability, the corresponding split structured data is regarded as an unreliable element.
[0073] In this embodiment, the error factor is comprehensively calculated from two dimensions: the proportion of the number of cause categories and the weighted error probability. reflects the proportion of the number of possible cause categories corresponding to the $w$-th authenticity attribute in the total number of categories, and reflects the coverage range of the reasons related to this authenticity attribute; is the average of the weighted error probabilities of different category reasons under this authenticity attribute. By the category weight distinguish the importance of different reasons, and then combine the error probability reasonably quantifies the possibility of errors caused by this authenticity attribute.
[0074] In this embodiment, the sigmoid function can and map the values of to a reasonable interval, avoid extreme values dominating the result, and then take the smaller value of the two through the min function, which reflects a conservative evaluation of authenticity judgment and error influence, and more rigorously measures the existence reliability of the element.
[0075] In this embodiment, the error factor is the probability weight of the error caused by the authenticity attribute. For example, the administrative division name error attribute contains 2 types of reasons: ① spelling error (weight 0.6, error probability 0.8); ② abbreviation usage (weight 0.4, error probability 0.3), then the error factor = (0.6×0.8 + 0.4×0.3) / (sum of total weights) = 0.6. Use Bayesian network to model the association of error reasons, and statistically calculate the prior probability and conditional probability of each factor through historical error data.
[0076] In this embodiment, the unreliable element is the split structured data (the error field that needs to be corrected with emphasis) whose existence reliability is lower than the threshold. For example, the matching coefficient of the house number field is 0.2, the error factor is 0.8, and the existence reliability is 0.16 < 0.5, which is determined as an unreliable element. Automatically determine the screening threshold through a threshold filtering algorithm (such as Otsu threshold method).
[0077] The beneficial effects of the above technical solution are as follows: By obtaining a quantitative result through the reliability formula, unreliable elements are finally screened based on this result, forming a complete closed-loop for the quality control of address data, making the technical process more logical and complete.
[0078] A method for address standardization and dynamic management according to the present invention corrects the expressions corresponding to the unreliable elements, including: Extract independent words consistent with the corresponding level from the unpartitioned part of the first storage block according to the level to which the unreliable element belongs, and regard them as the first words; Obtain the association relationship according to the semantic relevance, grammatical structure between the description of the unreliable element and each first word, and the position and association of the unreliable element and the first word in the address logic framework judged according to the overall information logic relationship set for the corresponding partition level; Based on the historical standard level description, determine the historical relevant correction rules, and combine the association relationship to correct the description corresponding to the unreliable element.
[0079] In this embodiment, the semantic relevance is an index measuring the similarity degree between the description of the unreliable element and the first word at the semantic level. The closer the semantics, the higher the relevance. The grammatical structure is the composition form of the description of the unreliable element and the first word in grammar, including the part of speech of the word, phrase structure, etc. The overall information logic relationship is the overall logic rule of the address information at the corresponding level. For example, the address segment at the district / county level needs to conform to the administrative division specification and be reasonably associated with the upper-level city, lower-level street and other levels in the address logic framework. The address logic framework is constructed according to the address levels and is a structured framework containing the logical association between address information at each level, clearly defining the inclusion and subordination relationships between levels such as country, province, city, district / county, etc.
[0080] In this embodiment, the association relationship is the connection determined between the unreliable element and the first word based on semantic relevance, grammatical structure, and overall information logical relationship, including semantic matching relationship, grammatical adaptation relationship, logical subordination relationship, etc., which is used to guide the correction of the unreliable element. For example, the unreliable element is Jinghai District (described as Jinghai District), and the first words are Haidian District and Dongcheng District. Calculate the semantic relevance: the cosine similarity of the word vectors between Jinghai District and Haidian District is 0.8 (assumed), and that with Dongcheng District is 0.3. Analyze the grammatical structure: all three are in the structure of noun + district, and the grammatical structure matching degrees are all 1. Judge the overall information logical relationship: Jinghai District (assumed to be incorrect), Haidian District, and Dongcheng District should all belong to the district and county level of Beijing, and their positions are the same in the address logical framework. Generally speaking, the association relationship between Jinghai District and Haidian District is stronger. The association relationship can be described as a high-matching relationship with highly similar semantics (relevance 0.8), consistent grammatical structure, and belonging to the same district and county level of Beijing; the association relationship with Dongcheng District is a general-matching relationship with relatively low semantic relevance (0.3), consistent grammatical structure, and belonging to the same district and county level of Beijing.
[0081] In this embodiment, a historical standard level description database is constructed and stored using an SQL database or a document database (such as MongoDB). Using the level to which the unreliable element belongs, error characteristics (such as semantic errors, grammatical errors), etc. as retrieval keywords, use database query statements to retrieve relevant historical standard level descriptions and correction rules, and analyze the retrieved historical standard level descriptions to extract correction rules closely associated with the current unreliable element. Through text mining techniques (such as keyword matching, rule template matching), applicable correction rules can be screened out from historical descriptions. Combining the association relationship and historical relevant correction rules, a correction strategy is determined. For example, if the association relationship indicates that a certain first word has the most semantic match with the unreliable element, and the historical correction rules support preferentially selecting such words for correction, then the unreliable element is replaced with the corresponding first word. Specific replacement and adjustment operations are implemented through programming code (such as Python scripts) to complete the correction of the description of the unreliable element. For example, for the unreliable element Jinghai District, it is retrieved from the historical standard level description that in the district and county level of Beijing, there was a case where Jinghai District was corrected to Haidian District due to a spelling mistake. The corresponding historical relevant correction rule is that for district and county names in the same city, if the error is caused by similar pronunciation or glyphs, select the correct name with the highest semantic relevance for correction. Combining the association relationship obtained previously, Jinghai District and Haidian District have the highest semantic relevance, consistent grammatical structure, and belong to the same district and county level of Beijing, so the correction operation is performed to correct Jinghai District to Haidian District.
[0082] The beneficial effects of the above technical solution are as follows: By accurately extracting candidate correction words (the first words), analyzing the association relationships in multiple dimensions (semantics, grammar, logic), and combining historical correction experience (historical standard level descriptions and rules), a scientific and efficient unreliable element correction process is constructed, continuously improving the address standardization processing ability.
[0083] A method for address standardization and dynamic management of the present invention establishes a security protection mechanism, including: Analyze the historical security information of each transmission interface and receiving interface within a specified time period to determine the set of risk types and the risk probability distribution of each risk type within the corresponding time period, and obtain the information risk combination at each time point; Extract the input behavior sequence of the receiving interface and each transmission interface at each time point within the specified time period for the login end of the corresponding transmission interface, and obtain the behavior risk combination at each time point; Generate a first security evaluation value, a second security evaluation value, and a comprehensive security evaluation value according to the information risk combination and behavior risk combination at the same time point, and screen and retain the maximum security evaluation value; Draw a curve for the retained maximum security evaluation value in chronological order of time points and divide the value levels according to the security protection standard to obtain several curve segments; Generate a security encryption factor based on the current security configuration information of the login end of the transmission interface involved in each curve segment and in combination with the value level, and set the encryption method for the corresponding continuous sub-time period according to the security encryption factor, where the encryption method is to perform information coding scrambling on the input address information based on the chaos level of the security encryption factor; Sort all the encryption methods in chronological order of continuous sub-time periods to obtain the initial protection mechanism.
[0084] In this embodiment, the security encryption factor is a dynamic encryption parameter generated by combining the security evaluation value and the current security configuration. For example, a high chaos level factor (such as a 256-bit random key) is generated when the security evaluation value is high, and a low level factor (such as a 128-bit key) is generated when the evaluation value is low. Use an encryption algorithm (such as AES) in combination with a random number generator (such as SHA-256) to dynamically generate the factor, and the factor strength is positively correlated with the security level.
[0085] In this embodiment, the chaos level is an index for measuring the degree of scrambling of the address information by the encryption method (the higher the level, the stronger the anti-attack ability). For example, chaos level 1 is simple substitution encryption, and level 5 is dynamic confusion encryption based on a neural network. Different levels of chaos encryption are implemented by combining different encryption algorithms (such as symmetric encryption + asymmetric encryption + hash algorithm), and the level is adjusted through parameters such as key complexity and encryption period.
[0086] In this embodiment, the resource consumption is the system resources consumed when changing the encryption method (such as CPU occupancy rate, memory bandwidth). For example, when changing from AES-128 to AES-256, the CPU occupancy rate increases from 20% to 35%, which is regarded as an increase in resource consumption. The resource usage data is collected in real time through a system monitoring tool (such as Prometheus), and the resource difference before and after the encryption method switch is calculated.
[0087] In this embodiment, the transmission interface is the outlet for sending address information, the receiving interface is the inlet for receiving address information, and the specified time period is a pre-set time interval for analyzing security information. For example, analyzing the security situation in the past 7 days or 1 month.
[0088] Historical security information is the security-related records of the transmission interface and the receiving interface within the specified time period, including the types of attacks suffered (such as SQL injection, malicious code injection), data leakage events, access anomalies (such as unauthorized access attempts), etc. The set of dangerous types is the set of all dangerous types identified from the historical security information, such as {SQL injection attack, unauthorized access, data tampering}.
[0089] The dangerous probability distribution is the probability of each dangerous type occurring at different time points within the specified time period.
[0090] The information dangerous combination is the combination of various dangerous types and their corresponding occurrence probabilities at each time point.
[0091] In this embodiment, the login end of the transmission interface is the terminal for logging in to the transmission interface to operate on the address information, such as the computer terminal used by logistics staff and the handheld device terminal of delivery personnel.
[0092] The input behavior sequence is the sequence of input operations on the address information by the login end of the transmission interface at each time point, including the frequency of inputting addresses, the number of addresses input at one time, and whether the format of the input addresses is abnormal (such as continuous input of a large number of addresses with incorrect formats), etc.
[0093] The behavior dangerous combination is the combination of the dangerous behaviors identified in the input behavior sequence and their occurrence situations at each time point. For example, the behavior dangerous combination at a certain time point is (high-frequency address input, 10 pieces / minute; incorrect format address input, 3 pieces).
[0094] In this embodiment, the first security evaluation value = Σ (weight of dangerous type × dangerous probability); The second security evaluation value = basic score - Σ (deduction for dangerous behaviors), where the value of the score ranges from 0 to 1.
[0095] The comprehensive security evaluation value = 0.6 × the first security evaluation value + 0.4 × the second security evaluation value.
[0096] The safety protection standard is a pre-established standard for classifying safety assessment values. For example, a safety assessment value of <0.2 is set as low risk, 0.2-0.6 is set as medium risk, and >0.6 is set as high risk.
[0097] The value classification is based on the security protection standards, and the maximum security assessment value at each time point is classified into the corresponding risk level (such as low, medium, and high risk).
[0098] The curve is drawn with time as the horizontal axis and the maximum safety assessment value as the vertical axis. The maximum safety assessment values at each time point are connected into a curve to intuitively show the time change trend of security risks.
[0099] A curve segment is a continuous part of the curve with the same safety assessment value level. Each curve segment corresponds to a continuous period of time and the same risk level.
[0100] The current security configuration information is the current security settings of the transmission interface login end, including firewall rules (such as whether to enable intrusion detection), encryption protocol version (such as TLS1.2, TLS1.3), identity authentication method (such as password + verification code, biometrics), etc.
[0101] Information coding scrambling is to encode and convert the input address information in an encrypted manner, disrupt the arrangement and storage form of the original information, and achieve encryption protection, such as converting XX Road in Chaoyang District, Beijing to X2# District Chaoyang North$X Road.
[0102] The continuous sub-time period is a continuous time interval corresponding to the curve segment, and each curve segment corresponds to a continuous sub-time period.
[0103] In this embodiment, a security encryption factor generation model is established, and the security encryption factor is generated through weighted calculation according to the curve segment level (such as a high risk level corresponds to a basic factor of 1.5) and the current security configuration information (such as turning on an advanced firewall and multiplying the factor by 1.2).
[0104] Determine the chaos level according to the security encryption factor (e.g. factor > 3 corresponds to chaos level 3), select the corresponding encryption algorithm (e.g. chaos level 3 uses AES algorithm combined with random permutation) to encode and scramble the address information, and write an encryption program (e.g. Python combined with pycryptodome library) to implement it.
[0105] A corresponding encryption method is associated with each continuous sub-time period, and the encryption strategy is stored.
[0106] The beneficial effects of the above technical solution are as follows: From the two dimensions of information risk and behavior risk, combining historical data and real-time behavior, comprehensively evaluate the interface security risk, making the security evaluation more accurate, being able to detect different types of security threats in a timely manner, dynamically adjust the encryption method according to the security evaluation results, adopting high-strength encryption when the security risk is high, and adapting to relatively lightweight encryption when the risk is low. While ensuring security, the system performance can be optimized. By drawing a security evaluation curve and dividing the curve segments, the changing trend of security risk can be visually presented.
[0107] For a method for address standardization and dynamic management according to the present invention, after obtaining the initial protection mechanism, it further includes: Determine the segment length of each continuous sub-time period and the resource consumption for encrypting method transformation under adjacent continuous sub-time periods; If each resource consumption is less than the preset consumption, regard the initial protection mechanism as a security protection mechanism; Otherwise, lock the first sub-time period whose resource consumption is greater than or equal to the preset consumption. If the security encryption factor of the first sub-time period is greater than that of the previous sub-time period, continue to keep the corresponding encryption method unchanged; Otherwise, change the encryption method of the first sub-time period to that of the previous sub-time period, thereby obtaining a security protection mechanism.
[0108] In this embodiment, the preset consumption is a preset and acceptable resource consumption threshold for encrypting method transformation. This threshold is determined according to factors such as the hardware performance of the system and the resource tolerance of the service. For example, set the CPU consumption threshold to 15% and the memory consumption threshold to 120MB.
[0109] With the help of system monitoring tools (such as the top command in the Linux system, the task manager in Windows, or professional APM tools such as Prometheus + Grafana), before and after the encrypting method transformation, collect data such as CPU occupancy rate, memory usage, and process execution time. Calculate the difference in resource data before and after the transformation to obtain the resource consumption for encrypting method transformation.
[0110] In this embodiment, associate and store the segment length of each continuous sub-time period and the resource consumption for adjacent transformations. A data table can be established using a database (such as MySQL), and the fields include sub-time period ID, segment length, previous encryption method, next encryption method, resource consumption (CPU), resource consumption (memory), etc., which is convenient for subsequent query and analysis.
[0111] The beneficial effects of the above technical solution are as follows: By analyzing resource consumption and judging encryption factors, the initial protection mechanism is optimized to create a security protection mechanism that better adapts to system resources and security requirements. While ensuring the security of address information transmission, it improves the operation efficiency and stability of the system, and builds a dynamically adjustable protection barrier for the secure transmission of address data.
[0112] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and its equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. An address standardization and dynamic management method, characterized in that Including: Step 1: Input the collected address information into the hierarchical entry framework to perform hierarchical division on the address information. Temporarily store the un-successfully divided part in the first storage block, and temporarily store the successfully divided part in the second storage block; Step 2: Convert the unstructured data in the second storage block into initial structured data according to the conversion standard of each division level; Step 3: Split the initial structured data according to the expression combination of the corresponding division level, and respectively judge the authenticity of each split structured data with the corresponding expression standard under the expression combination to construct a authenticity vector and a authenticity attribute vector, where the elements of the authenticity vector and the authenticity attribute vector correspond one by one; Step 4: Analyze the existence reliability of the same corresponding elements in the authenticity vector and the authenticity attribute vector, and screen out unreliable elements depending on the existence reliability; Step 5: Establish an association relationship between the unreliable elements and the un-successfully divided part in the first storage block, and retrieve the historical standard level description depending on the level to which the unreliable elements belong, and correct the expressions corresponding to the unreliable elements to obtain a standardized address, where the standardized address is obtained by sequentially structuring and describing in the order from the high level to the low level, and the corresponding last level is the lowest level in the collected address information; Step 6: Determine the historical security information of each transmission interface for transmitting the standardized address and the receiving interface of the big data platform during a specified time period, and establish a security protection mechanism and dynamically manage it.
2. The address standardization and dynamic management method according to claim 1, wherein Inputting the collected address information into the hierarchical entry framework to perform hierarchical division on the address information includes: Establish a hierarchical entry framework, where each level contains corresponding standard entries and a synonym library, and the hierarchical entry framework includes at least country, province, city, district, county, street, community / village, building, unit, floor and house number; Use natural language processing technology to perform word segmentation on the collected address information to obtain a number of independent words; Match each independent word with the standard entries and the synonym library in the hierarchical entry framework, and perform the matching in the order from the high level to the low level to obtain an independent word group at the level where the matching is successful; Determine the remaining words that are not successfully matched among the number of independent words, and respectively semantically associate each remaining word with the independent word group at the level where the matching is successful, and combine the relevant data in the geographic information database to determine the possible level to which it belongs; Temporarily store the independent word group at the level where the matching is successful in the second storage block; Temporarily store the independent words at the possible level to which they belong in the first storage block.
3. The address standardization and dynamic management method according to claim 1, characterized in that Constructing a authenticity vector and a authenticity attribute vector includes: Determine the first quantity of the split structured data and the combination quantity of the expression combination under the corresponding division level to obtain a first matching coefficient; According to the second matching coefficient of each split structured data in the corresponding division level and each expression standard under the expression combination, and judge whether the largest second matching coefficient is 1; If it is 1, determine that the corresponding split structured data is true; Otherwise, it is determined that the corresponding split structure data is false, and all the second matching coefficients involved are analyzed. Combining with the historical matching failure event set based on each expression standard at the corresponding division level, the authenticity attribute of the corresponding split structure data is determined, where the authenticity attribute is related to the possible reasons for the corresponding split structure data being false; A authenticity vector is constructed based on the largest second matching coefficient corresponding to each split structure data. At the same time, an authenticity attribute vector is constructed from the authenticity attributes corresponding to each split structure data.
4. The address standardization and dynamic management method according to claim 3, characterized in that, Determining the authenticity attribute of the corresponding split structure data includes: Each historical matching failure event in the historical matching failure event set is input into the event analysis model to obtain the first failure factor. At the same time, the matching process log of each historical matching failure event is retrieved from the historical diary record database and input into the log analysis model to obtain the second failure factor, and a failure factor table for the corresponding historical matching failure event is established; Aggregate analysis is performed on all failure factor tables to establish a complete description framework; Perform a unique matching judgment on all the second matching coefficients corresponding to the false split structure data, that is, judge whether the second quantity of the second matching coefficient with a value of 0 is N - 1. If so, regard the second matching coefficient that is not 0 as the third matching coefficient, and combine the expression standards of the first matching coefficient and the third matching coefficient to construct a search index, where N represents the number of all second matching coefficients under the corresponding split structure data; If not, determine the coefficient ratio of the second matching coefficient that is not 0, and combine the expression standard under the second matching coefficient that is not 0 and the first matching coefficient to construct a search index; Based on the search index, search and locate the possible reasons in the complete description framework to obtain the authenticity attribute of the corresponding split structure data.
5. The address standardization and dynamic management method according to claim 1, characterized in that Analyze the existence reliability of the same corresponding elements in the authenticity vector and the authenticity attribute vector, and filter out unreliable elements depending on the existence reliability, including: Determine the error factor of each authenticity attribute; ; Among them, is the failure factor of the w-th authenticity attribute; n represents the number of categories of possible causes corresponding to the w-th authenticity attribute; represents the total number of categories of all possible causes; represents the category weight of the i-th category cause under the w-th authenticity attribute; represents the error probability of converting unstructured data to structured data caused by the i-th category cause under the w-th authenticity attribute; Determine the existence reliability of the authenticity value and the error factor of the same corresponding element; ; Among them, is the existence reliability of the element corresponding to the w-th authenticity attribute; represents the authenticity value of the element corresponding to the w-th authenticity attribute; If the existence reliability is less than the preset reliability, regard the corresponding split structure data as an unreliable element.
6. The address standardization and dynamic management method according to claim 1, characterized in that Correct the description corresponding to the unreliable element, including: Extract independent words consistent with the corresponding level from the unpartitioned part of the first storage block according to the level to which the unreliable element belongs, and regard them as the first words; Obtain the association relationship based on the semantic relevance, grammatical structure of the description of the unreliable element and each first word, and the position and association of the unreliable element and the first word in the address logic framework judged according to the overall information logic relationship set by the corresponding division level; Based on the historical standard level description, determine the historical relevant correction rules, and combine the association relationship to correct the description corresponding to the unreliable element.
7. The address standardization and dynamic management method according to claim 1, characterized in that Establish a security protection mechanism, including: Analyze the historical security information of each transmission interface and receiving interface in a specified time period, determine the set of dangerous types and the dangerous probability distribution of each dangerous type in the corresponding time period, and obtain the information danger combination at each time point; Extract the input behavior sequences of the login ends of the corresponding transmission interfaces at each time point of the receiving interface and each transmission interface within a specified time period, and obtain the behavior risk combinations at each time point; Generate a first security evaluation value, a second security evaluation value, and a comprehensive security evaluation value according to the information risk combinations and behavior risk combinations at the same time point, and screen and retain the maximum security evaluation value; Draw a curve for the retained maximum security evaluation value in chronological order of time points and divide the value levels according to the security protection standard to obtain several curve segments; Generate a security encryption factor based on the current security configuration information of the login ends of the transmission interfaces involved in each curve segment and in combination with the value level, and set an encryption method for the corresponding continuous sub-time period corresponding to the curve segment according to the security encryption factor, wherein the encryption method is to perform information coding scrambling on the input address information based on the chaos level of the security encryption factor; Sort all the encryption methods in sequence according to the continuous sub-time period order to obtain an initial protection mechanism.
8. The address standardization and dynamic management method according to claim 7, characterized in that After obtaining the initial protection mechanism, it further includes: Determine the segment length of each continuous sub-time period and the resource consumption for changing the encryption method under adjacent continuous sub-time periods; If each resource consumption is less than the preset consumption, regard the initial protection mechanism as a security protection mechanism; Otherwise, lock the first sub-time period with a resource consumption greater than or equal to the preset consumption. If the security encryption factor of the first sub-time period is greater than the security encryption factor of the previous sub-time period, continue to retain the corresponding encryption method unchanged; Otherwise, change the encryption method of the first sub-time period to the encryption method of the previous sub-time period to obtain a security protection mechanism.
Citation Information
Patent Citations
Method and device for standardizing geographic addresses
CN110019575A
Address error correction method and device, computer equipment and readable medium
CN116432633A
Mail misclassification risk determination method and device, and medium
CN119647938A
Address normalization using deep learning and address feature vectors
US10839156B1
Method for displaying entity-associated information based on electronic book and electronic device
US20220343077A1
Cited By
Address text correlation analysis method, system and equipment and storage medium
CN120850955A