Address identification method and device, electronic equipment and storage medium
By acquiring and analyzing multiple geographical data of target customers and building and comparing map matrices, the problems of address deviation and relocation detection in traditional methods are solved, and more accurate address recognition and business expansion are achieved.
Patent Information
- Application Number
- CN202510193309.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-10
AI Technical Summary
Traditional address positioning methods cannot accurately determine whether the group’s customer address is relocated or there are deviations, resulting in inaccuracy in business expansion and service provision.
By obtaining at least two geographical data of the target customer, grid marking is performed on the digitized map, a map matrix is constructed, and the similarity function value between different map matrices is calculated based on the preset coordinate relationship recognition model. If the similarity function value is greater than the preset threshold, the geographical data is determined to be the correct address.
Accurate identification of addresses is achieved, business expansion and service provision inaccuracy caused by address deviations is avoided, and the accuracy of address determination is improved.
Smart Images

Figure CN120123445A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of big data technology, and in particular, to a method and apparatus for identifying an address, an electronic device, and a storage medium. Background Art
[0002] Traditional location verification for the geographical locations of corporate customers is performed by analyzing and predicting based on addresses in various public platforms or business information, or manually verifying the geographical locations offline. However, for problems such as address changes of corporate customers, through ordinary conventional distance calculations, it is impossible to accurately infer whether the corporate customer's address has been relocated or whether there are deviations in the specific address.
[0003] Deviations in the address will affect the subsequent business expansion of front-line business personnel and prevent service providers from accurately finding the customer's address for services. Summary of the Invention
[0004] The present disclosure provides a method, apparatus, electronic device, and storage medium for identifying an address. Its main purpose is to solve the problem of address deviation.
[0005] According to a first aspect of the present disclosure, there is provided a method for identifying an address, including:
[0006] Obtaining at least two geographical data of a target customer; wherein, the data sources of different geographical data are different;
[0007] Respectively performing grid marking of the geographical data on a digital map to construct a map matrix; wherein, different geographical data correspond to different map matrices;
[0008] Calculating a similarity function value between different map matrices based on a preset coordinate relationship recognition model, and determining that the geographical data is a correct address when the similarity function value is greater than a preset threshold.
[0009] Optionally, calculating the similarity function value between different map matrices based on the preset coordinate relationship recognition model includes:
[0010] Respectively calculating the covariance matrices between different map matrices;
[0011] Calculating the similarity function value according to the covariance matrix and a geographical data position comprehensive function;
[0012] Wherein, the geographical data position comprehensive function is the sum of the position information matrices of each geographical data; the position information matrix includes an address code.
[0013] Optionally, calculating the similarity function value between different map matrices according to the covariance matrix includes:
[0014] Calculate a similarity function based on the covariance matrix and the geographical data location synthesis function;
[0015] Accumulate all elements within the similarity function to obtain the similarity function value.
[0016] Optionally, the determining that the geographical data is the correct address when the similarity function value is greater than a preset threshold includes:
[0017] Determine the rating of the similarity function value based on a preset scoring threshold; wherein, a high rating corresponds to a high degree of correctness of the address.
[0018] Optionally, the respectively grid-marking the geographical data on a digital map and constructing a map matrix includes:
[0019] Perform grid division based on a preset scale and the coordinate information in each piece of geographical data;
[0020] Generate a map matrix respectively according to each of the divided grids, and mark the coordinate information in the map matrix.
[0021] Optionally, after calculating the similarity function value between different map matrices based on a preset coordinate relationship recognition model, the method further includes:
[0022] When the geographical data is not the correct address, obtain on-site inspection geographical data and construct a digital matrix;
[0023] Use the on-site inspection geographical data as the correct address of the target customer, and train the preset coordinate relationship recognition model according to the digital matrix.
[0024] Optionally, the geographical data includes at least two of registration information, industrial and commercial registration information, and station information.
[0025] According to a second aspect of the present disclosure, there is provided an address recognition device, including:
[0026] A first acquisition unit, configured to acquire at least two pieces of geographical data of a target customer; wherein, the data sources of different geographical data are different;
[0027] A construction unit, configured to respectively grid-mark the geographical data on a digital map and construct a map matrix; wherein, different geographical data correspond to different map matrices;
[0028] A determination unit, configured to calculate the similarity function value between different map matrices based on a preset coordinate relationship recognition model, and determine that the geographical data is the correct address when the similarity function value is greater than a preset threshold.
[0029] Optionally, the determining unit is further configured to:
[0030] Calculate the covariance matrices between different ones of the map matrices respectively;
[0031] Calculate the similarity function value according to the covariance matrix and the geographical data location synthesis function;
[0032] Wherein, the geographical data location synthesis function is the sum of the location information matrices of each geographical data; the location information matrix includes address codes.
[0033] Optionally, the determining unit is further configured to:
[0034] Calculate a similarity function according to the covariance matrix and the geographical data location synthesis function;
[0035] Accumulate all elements in the similarity function to obtain the similarity function value.
[0036] Optionally, the determining unit is further configured to:
[0037] Determine the rating of the similarity function value based on a preset scoring threshold; wherein, a high rating corresponds to a high correctness of the address.
[0038] Optionally, the constructing unit is further configured to:
[0039] Perform grid division based on a preset scale and the coordinate information in each geographical data;
[0040] Generate map matrices respectively according to each of the divided grids, and mark the coordinate information in the map matrices.
[0041] Optionally, the apparatus further includes:
[0042] A second obtaining unit, configured to, after the determining unit calculates the similarity function value between different map matrices based on a preset coordinate relationship recognition model, when the geographical data is not the correct address, obtain on-site inspection geographical data and construct a digital matrix;
[0043] A training unit, configured to use the on-site inspection geographical data as the correct address of the target customer, and train the preset coordinate relationship recognition model according to the digital matrix.
[0044] Optionally, the geographical data includes at least two of registration information, industrial and commercial registration information, and station information.
[0045] According to a third aspect of the present disclosure, there is provided an electronic device, including:
[0046] At least one processor; and
[0047] A memory communicatively connected to the at least one processor; wherein,
[0048] The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method described in the foregoing first aspect.
[0049] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method described in the foregoing first aspect.
[0050] According to a fifth aspect of the present disclosure, there is provided a computer program product including a computer program, and when the computer program is executed by a processor, the method described in the foregoing first aspect is implemented.
[0051] The address recognition method, device, electronic device and storage medium provided by the present disclosure, the main technical solutions include: obtaining at least two geographical data of a target customer; wherein, the data sources of different geographical data are different; respectively performing grid marking on the geographical data on a digital map to construct a map matrix; wherein, different geographical data correspond to different map matrices; calculating a similarity function value between different said map matrices based on a preset coordinate relationship recognition model, and when the similarity function value is greater than a preset threshold, determining that the geographical data is a correct address. Compared with the related art, in the embodiment of the present application, by calculating the covariance matrix between matrixed maps, the statistical correlation between geographical information from different sources is comprehensively measured, and whether different geographical information matches is accurately discriminated from a statistical perspective. Different from the traditional distance-based calculation method, the address can be recognized more accurately.
[0052] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0054] Figure 1 is a schematic flowchart of a method for recognizing an address provided by an embodiment of the present disclosure;
[0055] Figure 2 is a schematic flowchart of a method for recognizing an address provided by an embodiment of the present disclosure;
[0056] Figure 3A schematic diagram of grid division provided by an embodiment of the present disclosure;
[0057] Figure 4 A schematic diagram of the identification process of an address provided by an embodiment of the present disclosure;
[0058] Figure 5 A schematic structural diagram of an address identification device provided by an embodiment of the present disclosure;
[0059] Figure 6 A schematic structural diagram of an address identification device provided by an embodiment of the present disclosure;
[0060] Figure 7 A schematic block diagram of an exemplary electronic device provided by an embodiment of the present disclosure. Detailed implementation manners
[0061] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted below.
[0062] Next, the address identification method, device, electronic device, and storage medium of the embodiments of the present disclosure are described with reference to the accompanying drawings.
[0063] Figure 1 A schematic flowchart of an address identification method provided by an embodiment of the present disclosure.
[0064] As Figure 1 shown, the method includes the following steps:
[0065] Step 101: Obtain at least two geographical data of the target customer; wherein, the data sources of different geographical data are different.
[0066] In some embodiments, the geographical data can be obtained through satellite remote sensing data, ground survey data, and publicly accessible paths, and the embodiments of the present application do not limit this.
[0067] Step 102: Respectively perform grid marking on the geographical data on the digital map to construct a map matrix; wherein, different geographical data correspond to different map matrices.
[0068] Performing grid marking on the digital map and constructing a map matrix are used to integrate and analyze different types of geographical data.
[0069] First, determine the size of the grid. In some embodiments, the grid size can be determined based on the accuracy requirements. For example, when the accuracy requirements are high, small grids are used; when the accuracy requirements are low, large grids are used. Specifically, it can be set according to actual needs, and the embodiments of the present application do not limit this.
[0070] Assign attribute values to each grid cell. For example, obtain the location coordinates of the target customer in the geographical data, and assign attribute values to the grid cells corresponding to the location coordinates for differentiation. Finally, create a matrix where the rows and columns represent the grid cells, and the values in the matrix represent the attribute values of the corresponding grid cells.
[0071] Different types of geographical data can be analyzed independently in their respective map matrices or superimposed together for comprehensive analysis, thereby providing a more comprehensive view of the spatial information.
[0072] Step 103: Calculate the similarity function value between different said map matrices based on a preset coordinate relationship recognition model, and determine that the geographical data is the correct address when the similarity function value is greater than a preset threshold.
[0073] The similarity function is used to quantify the similarity degree between two map matrices. In some embodiments, it can be calculated based on cosine similarity, Euclidean distance, or Jaccard similarity coefficient. Specifically, the embodiments of the present application do not limit this.
[0074] For each pair of map matrices, use the similarity function to calculate their similarity values. Specifically, the overall similarity score between the two matrices can be obtained through methods such as averaging, weighted averaging, or other aggregation methods. In some embodiments, a similarity threshold is set according to the specific application scenario and requirements. Determine when two map matrices are considered to represent the same geographical data through the threshold. If the similarity function value is greater than or equal to the preset threshold, it is considered that the geographical data represented by the two map matrices is matched, that is, the geographical data is the correct address; if the similarity function value is lower than the preset threshold, it is considered that the geographical data represented by the two map matrices is not matched, that is, the geographical data is the wrong address.
[0075] The address recognition method provided by the present disclosure mainly includes the following technical solutions: obtaining at least two geographical data of a target customer; wherein, the data sources of different geographical data are different; respectively performing grid marking on the geographical data on a digital map to construct a map matrix; wherein, different geographical data correspond to different map matrices; calculating the similarity function value between different map matrices based on a preset coordinate relationship recognition model, and determining that the geographical data is a correct address when the similarity function value is greater than a preset threshold. Compared with the related art, the embodiment of the present application comprehensively measures the statistical correlation between geographical information from different sources by calculating the covariance matrix between matrixized maps, accurately discriminates whether different geographical information matches from a statistical perspective, and is different from the traditional distance-based calculation method, and can more accurately identify the address.
[0076] In some embodiments, the geographical data includes at least two of registration information, industrial and commercial registration information, and station information. The calculating the similarity function value between different map matrices based on a preset coordinate relationship recognition model includes:
[0077] Calculating the covariance matrix between different map matrices respectively;
[0078] Calculating the similarity function value according to the covariance matrix and the geographical data position comprehensive function;
[0079] Wherein, the geographical data position comprehensive function is the sum of the position information matrices of each geographical data; the position information matrix contains address codes.
[0080] Based on the abstraction processing of position information, abstract the covariance matrix of registration information, industrial and commercial information, and station information to construct a digital matrix.
[0081] Covariance and covariance matrix: According to the definition of variance, given d random variables, x k , k = 1, 2,..., d, the variances of these random variables are:
[0082]
[0083] Wherein, x ki represents the i-th observation sample in the random variable , n represents the sample size, and the number of observation samples corresponding to each random variable is n.
[0084] (2) Construction of covariance matrix
[0085] For these random variables, calculate the covariance between each pair, that is:
[0086]
[0087] Therefore, the covariance matrix is as follows:
[0088]
[0089] Among them, the elements on the diagonal are the variances of each random variable, and the elements off the diagonal are the covariances between pairwise random variables. According to the definition of covariance, the covariance matrix is recognized as a symmetric matrix.
[0090] In this solution, the covariance matrix is as follows:
[0091]
[0092] Among them, M kk′ (p) represents the covariance matrix between geographical information positions, represents the registered geographical location information matrix, M k′ represents the registered industrial and commercial location information matrix, M k represents the digital matrix of the information of personnel stationed at the station. Through calculation, the covariance matrix M kk′ (p) of all coordinate elements can be obtained.
[0093] The covariance matrix contains all the information of the coordinate data set. This proposal constructs a coordinate relationship judgment based on the covariance matrix, which is more comprehensive and abstract than the traditional distance judgment. Compared with the traditional method that only relies on a single distance index, it can more accurately describe the matching relationship between coordinates and avoid judgment errors. Based on the covariance matrix, this proposal establishes a coordinate relationship model from the perspectives of statistics and matrix operations, realizing a comprehensive digital abstract description of the determination of the target geographical location.
[0094] Please refer to Figure 2 , Figure 2 which is the flowchart of a method for identifying an address provided by an embodiment of the present disclosure, including:
[0095] Step 201, calculate a similarity function according to the covariance matrix and the geographical data position synthesis function.
[0096] According to the quantum state evolution mode defined by the Schrödinger equation:
[0097] H|ψ> = E|ψ>
[0098] Quantum state probability calculation formula:
[0099] Calculate the similarity function between different matrixed maps:
[0100] Dr = <M kk′ (p)|M>
[0101] Among them, M kk′(p) represents each element in the covariance matrix, M represents the geographical information location integration function, D r represents the similarity function, D r represents calculating the correlation between each type of geographical location information and the geographical information location integration function. Based on the similarity function, the correlation between different geographical information can be calculated, thereby inferring whether the user's geographical location has changed or is incorrect.
[0102] M k′ 、M k The three matrices respectively have corresponding multi-level text addresses in geographical information. For example, the existing nine-level address: No. 1501, Unit 3, 5th Floor, Building 5, Fenghua Community, Maple Leaf Road, Chenghua Avenue, Chenghua District, Chengdu City. Code Chengdu City as 1, Chenghua District as 2, Chenghua Avenue as 3, Maple Leaf Road No. 1 as 1, Fenghua Community as 5, Building 5 as 5, Unit 3 as 3, 5th Floor as 5, and 1501 as 1. This proposal can abstract these elements into a matrix vector: [1, 2, 3, 1, 5, 5, 3, 5, 1], forming a 9-dimensional vector.
[0103] This proposal can judge whether the address is repeated based on this coding rule, ensure that each address is non-repeated and unique, and abstract the matrix vector into a 3×3 matrix. For example, the vector [1, 2, 3, 1, 5, 5, 3, 5, 1] is abstracted into the following 3Δ3 matrix:
[0104]
[0105] It should be noted that this description method is only an exemplary illustration and not a specific limitation on the specific content. The embodiments of this application do not limit this.
[0106] Step 202, accumulate all elements in the similarity function to obtain the similarity function value.
[0107] By summing all elements of the D r vector to calculate the final score of the similarity function.
[0108] To clearly illustrate the calculation of the similarity function value in the embodiments of this application, the following is an example for illustration:
[0109] The covariance matrix M kk′ (p) and the geographical information location integration function M are respectively:
[0110]
[0111] Calculate the matrix product of M kk′ (p) and M to obtain the similarity function D rVector form:
[0112]
[0113] By summing all elements of the D r vector to calculate the final score of the similarity function:
[0114] Dr score = ∑Dr = 5.0 + 9.5 = 14.5
[0115] The final score Dr of the similarity function D r is obtained by calculating the sum of the elements of the vector obtained from the matrix product of M score (p) and M, and the result is 14.5. kk′ (p) and M, and the result is 14.5.
[0116] It should be noted that this description method is only an exemplary illustration and not a specific limitation on specific values.
[0117] Step 203, determining the rating of the similarity function value based on a preset scoring threshold; wherein, a high rating corresponds to a high address correctness.
[0118] Based on D r to determine whether to adjust the conclusion of the user's geographical location. When it is necessary to verify the new geographical location, a range of similarity values is provided based on a large number of previous sample sets. For example: for a large number of government and enterprise customers with different geographical location information coordinates, it can be obtained that different geographical information coordinates actually correspond to the same location coordinate in the Internet map, and it can also be obtained that different geographical information coordinates actually correspond to different location coordinates in the Internet map. Then, in both cases, there will be a range of similarity values. Using this range of the sample set, the geographical location information of new government and enterprise customers to be identified is identified. According to the similarity function value obtained by the algorithm, it is judged whether the address is real, whether it represents an address, and whether it corresponds to the same geographical location information on the Internet map.
[0119] Scoring based on the magnitude of the similarity function value can more intuitively identify the accuracy of address discrimination.
[0120] Score according to the magnitude of the similarity function value. The smaller the function value, the more similar the different coordinates are, the more accurate the address discrimination is, and the higher the score. Set function value segments, and different scoring strategies are adopted in different intervals. The interval setting is determined according to the actual situation. Considering the differences between automatic discrimination and manual discrimination, the score for automatic discrimination is high, and the score for manual discrimination is low.
[0121] A. When the similarity function value is less than 0.1, give 100 points, indicating that the address discrimination is highly accurate.
[0122] B. The similarity function value is between 0.1 and 0.3, and the score is 80. The address is basically accurate.
[0123] C. The similarity function value is between 0.3 and 0.5, and the score is 60. The address is roughly correct but there are slight differences.
[0124] D. The similarity function value is between 0.5 and 0.7, and the score is 40. The address is not very accurate and there is a certain degree of deviation.
[0125] E. The similarity function value is greater than 0.7, and the score is 20. The address is basically inaccurate.
[0126] F. The result of manual discrimination is scored 80. The result of manual discrimination is not as accurate as automatic discrimination.
[0127] According to the score corresponding to the similarity function value, the manual discrimination score is used to obtain the comprehensive score S: S = S1 + S2 × C
[0128] Where S1 is the score corresponding to the similarity function value, S2 is the manual discrimination score, and C is the discrimination type coefficient. When automatically discriminating, C = 1, and when manually discriminating, C = 0.8.
[0129] If the score is very high, the address similarity is high, indicating that the positions of several address information are very close. When offline business needs to be handled, the customer address can be accurately found according to the entered address, the station address, and the industrial and commercial information address.
[0130] In some embodiments, the grid marking of the geographical data on the digital map respectively to construct the map matrix includes:
[0131] Based on a preset scale and the coordinate information in each of the geographical data, grid division is performed;
[0132] Map matrices are respectively generated according to each of the divided grids, and the coordinate information is marked in the map matrix.
[0133] Please refer to Figure 3 , Figure 3 which is a schematic diagram of grid division provided by an embodiment of the present application. As Figure 3 shown, for the same group customer, there are registration information, industrial and commercial information, personnel station information, and Internet map information (this map information is all the doorplate points with existing and registered positions in the regional map) for this customer, and the corresponding geographical coordinate information is obtained, and on the map, the grid is divided well. For example Figure 2As shown in the figure, on the coordinate map of each piece of information, a 16×16 coordinate grid is drawn. The grid position where the registration location is located is marked as the number 1; the geographical location of the in-station information is accumulated according to the frequency of personnel appearance. If it appears n times, it is marked as the number n. The following information network is abstracted into four 16×16 digital matrices. For example:
[0134]
[0135] Establish a matrix map for data from different sources, compare the markings of the same grid in different matrices, determine whether the target objects such as group customers have changed their locations, realize the automatic and accurate identification of the target geographical locations, and avoid judgment deviation. The matrix map can comprehensively integrate different information sources in all directions, avoid the limitations of a single data source, and improve the accuracy of location determination.
[0136] In some embodiments, after calculating the similarity function value between different said map matrices based on a preset coordinate relationship recognition model, the method further includes:
[0137] In the case where the geographical data is not the correct address, obtain on-site inspection geographical data and construct a digital matrix;
[0138] Use the on-site inspection geographical data as the correct address of the target customer, and train the preset coordinate relationship recognition model according to the digital matrix.
[0139] Introduce the on-site inspection information of personnel, generate a separate matrix for the location information of personnel inspection, mark the position points as the number 1, and abstract it into a digital matrix:
[0140] This step is a preparation for the final location determination. If there is a conflict between the marked information on the map and the on-site inspection information of the manual later, the location information of this inspection is used as the most powerful evidence.
[0141] Next, according to Figure 4 , take an application scenario as an example for illustration. Please refer to Figure 4 , Figure 4 which is a schematic flowchart of an address recognition method provided by an embodiment of the present application:
[0142] For a government and enterprise customer, there may be multiple address information from different sources in the system, such as registration information, industrial and commercial information, personnel in-station information, etc. This proposal aims to automatically determine whether these address information from different sources correspond to the same real address through an algorithm, so as to determine the accurate address of the customer.
[0143] 1. The two data for calculating similarity are:
[0144] a) A series of matrices generated by different source addresses such as registration address information, industrial and commercial registration information, and personnel station information through the "map matrix module", which reflect the position distribution of each address in the standard map grid.
[0145] b) The covariance matrix obtained by operating on the above address distribution matrix through the "covariance matrix calculation module", which reflects the statistical correlation between different address distributions.
[0146] These two types of matrix data are used for the calculation of the similarity function to determine the similarity degree between different addresses.
[0147] 2. For the two situations where different geographical information coordinates correspond to positions on the Internet map, there is no difference in the process of obtaining the similarity value range. Specifically:
[0148] a) Whether multiple address information corresponds to the same real address or different real addresses, they will all first pass through the "map matrix module" to generate a position distribution matrix in the standard map grid.
[0149] b) Then these position distribution matrices will be input into the "covariance matrix calculation module" and converted into a covariance matrix reflecting the correlation between addresses through calculation.
[0150] c) Finally, the covariance matrices of different addresses will be substituted into the similarity function for operation to obtain a value representing the similarity degree.
[0151] d) The difference is that if multiple addresses correspond to the same real location, their position distributions are closer, the covariance is larger, and the similarity value calculated finally will be very high. Otherwise, the similarity value is lower.
[0152] e) Through the training of a large number of sample data of known addresses, the discrimination threshold of the similarity value can be obtained. Using this threshold interval, the actual corresponding relationship of multiple addresses of new customers can be determined.
[0153] In summary, the calculation process of the similarity value of multiple address information is the same, but the similarity values of the same address and different addresses will fall in different intervals. It is precisely by using this rule that the system realizes the automatic determination of address authenticity.
[0154] In the initial stage of the algorithm, for the new samples to be identified, the system inputs their respective source geographical coordinates and calculates the similarity function between different coordinate data through the trained model to determine whether the coordinates of the new samples correspond to the same geographical location. If there are obvious differences in the calculation results of the similarity function, it indicates that the different coordinates of the new samples correspond to different geographical locations and further verification is required.
[0155] This judgment criterion and range can be determined through the analysis and statistics of a large number of known sample data.
[0156] Specifically, the following steps can be taken:
[0157] Collect a large number of known address samples, which include multi-source coordinate data corresponding to the same geographical location and different geographical locations. For each sample, calculate the similarity function values between its different-source coordinates. For samples corresponding to the same geographical location, statistically analyze the distribution range of their similarity function values, and calculate a threshold range, such as [0.8, 1]. This means that when the similarity function value of a new sample falls within this range, it can be considered that its different-source coordinates are very likely to correspond to the same geographical location.
[0158] For samples corresponding to different geographical locations, also statistically analyze the distribution range of their similarity function values, and calculate another threshold range, such as [0, 0.2]. This means that when the similarity function value of a new sample falls within this range, it can be considered that its different-source coordinates are very likely to correspond to different geographical locations.
[0159] The similarity function values between the two threshold ranges, such as (0.2, 0.8), can be regarded as the "fuzzy interval". When the similarity function value of a new sample falls within this interval, it is difficult to directly judge its coordinate correspondence, and further verification is required.
[0160] Therefore, the judgment criterion for the so-called "obvious difference" is that the similarity function value falls within the threshold range determined in step 4. When the similarity function value of a new sample meets this condition, it is considered that there is an obvious difference between its different-source coordinates, corresponding to different geographical locations.
[0161] The system marks the coordinates of the new sample on the Internet map, and directly calculates the matching degree between the different-source coordinates and the standard map coordinates, as a further basis for judging the correctness of the new sample coordinates.
[0162] If the system cannot automatically judge the accurate correspondence of the new sample coordinates, the result of on-site manual investigation is used as the golden standard for judging the accuracy of the new sample coordinates. The result of the manual investigation is fed back to the system to continuously optimize the judgment ability of the model.
[0163] By training a large number of known samples and constructing a learning model for coordinate data matching, the present invention realizes the intelligent and automatic judgment of the geographical location information of government and enterprise customers, avoiding the address determination deviation existing in the traditional method. If the Dr value is very small, it is reasonable to consider that the address manually entered in the system is accurate, and the subsequent business processing can be carried out based on the information registered in the system. The trust priority of the address can be defined as the system-entered address, the on-site information address, and the industrial and commercial registration address from high to low.
[0164] Corresponding to the above address recognition method, the present invention also provides an address recognition device. Since the device embodiments of the present invention correspond to the above method embodiments, for details not disclosed in the device embodiments, reference may be made to the above method embodiments, and details will not be elaborated herein.
[0165] Figure 5 As shown in the structural schematic diagram of an address recognition device provided by an embodiment of the present disclosure, Figure 5 it includes:
[0166] A first acquisition unit 31, configured to acquire at least two geographical data of a target customer; wherein, the data sources of different geographical data are different;
[0167] A construction unit 32, configured to respectively perform grid marking on the geographical data on a digital map to construct a map matrix; wherein, different geographical data correspond to different map matrices;
[0168] A determination unit 33, configured to calculate a similarity function value between different map matrices based on a preset coordinate relationship recognition model, and determine that the geographical data is a correct address when the similarity function value is greater than a preset threshold.
[0169] The address recognition device provided by the present disclosure mainly includes the following technical solutions: acquiring at least two geographical data of a target customer; wherein, the data sources of different geographical data are different; respectively performing grid marking on the geographical data on a digital map to construct a map matrix; wherein, different geographical data correspond to different map matrices; calculating a similarity function value between different map matrices based on a preset coordinate relationship recognition model, and determining that the geographical data is a correct address when the similarity function value is greater than a preset threshold. Compared with the related art, the embodiments of the present application comprehensively measure the statistical correlation between geographical information from different sources by calculating the covariance matrix between matrix maps, accurately distinguish whether different geographical information matches from a statistical perspective, and is different from the traditional distance-based calculation method, and can more accurately identify addresses.
[0170] Further, in a possible implementation manner of an embodiment of the present disclosure, the determination unit 33 is further configured to:
[0171] Calculate the covariance matrix between different map matrices respectively;
[0172] Calculate the similarity function value according to the covariance matrix and the geographical data position comprehensive function;
[0173] Wherein, the geographical data position comprehensive function is the sum of the position information matrices of each geographical data; the position information matrix includes address codes.
[0174] Further, in a possible implementation manner of the embodiments of the present disclosure, the determining unit 33 is further configured to:
[0175] Calculate a similarity function according to the covariance matrix and the geographic data location synthesis function;
[0176] Accumulate all elements in the similarity function to obtain the similarity function value.
[0177] Further, in a possible implementation manner of the embodiments of the present disclosure, the determining unit 33 is further configured to:
[0178] Determine a rating of the similarity function value based on a preset scoring threshold; wherein, a high rating corresponds to a high correctness of the address.
[0179] Further, in a possible implementation manner of the embodiments of the present disclosure, the constructing unit 32 is further configured to:
[0180] Perform grid division based on a preset scale and the coordinate information in each piece of the geographic data;
[0181] Generate a map matrix according to each of the divided grids respectively, and mark the coordinate information in the map matrix.
[0182] Further, in a possible implementation manner of the embodiments of the present disclosure, as Figure 6 shown, the apparatus further includes:
[0183] A second obtaining unit 34, configured to obtain on-site inspection geographic data and construct a digital matrix when the geographic data is not the correct address after the determining unit 33 calculates the similarity function value between different map matrices based on a preset coordinate relationship recognition model;
[0184] A training unit 35, configured to use the on-site inspection geographic data as the correct address of the target customer, and train the preset coordinate relationship recognition model according to the digital matrix.
[0185] Further, in a possible implementation manner of the embodiments of the present disclosure, the geographic data includes at least two of registration information, industrial and commercial registration information, and station information.
[0186] It should be noted that the foregoing explanations of the method embodiments are also applicable to the apparatus of the embodiments of the present disclosure, with the same principle, and are not limited in the embodiments of the present disclosure.
[0187] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0188] Figure 7 FIG. 1 is a schematic block diagram of an example electronic device 400 that may be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as, for example, personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0189] As Figure 7 shown, the device 400 includes a computing unit 401 that can perform various appropriate actions and processes in accordance with a computer program stored in a ROM (Read-Only Memory) 402 or a computer program loaded from a storage unit 408 into a RAM (Random Access Memory) 403. In the RAM 403, various programs and data required for the operation of the device 400 can also be stored. The computing unit 401, the ROM 402, and the RAM 403 are connected to each other via a bus 404. An I / O (Input / Output) interface 405 is also connected to the bus 404.
[0190] A plurality of components in the device 400 are connected to the I / O interface 405, including: an input unit 406, such as a keyboard, a mouse, etc.; an output unit 407, such as various types of displays, speakers, etc.; a storage unit 408, such as a magnetic disk, an optical disk, etc.; and a communication unit 409, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 409 allows the device 400 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0191] The computing unit 401 can be various general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 401 include, but are not limited to, a CPU (Central Processing Unit), a GPU (Graphic Processing Units), various dedicated AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, a DSP (Digital Signal Processor), and any suitable processor, controller, microcontroller, etc. The computing unit 401 executes the various methods and processes described above, such as the method for identifying an address. For example, in some embodiments, the method for identifying an address can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 408. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 400 via the ROM 402 and / or the communication unit 409. When the computer program is loaded into the RAM 403 and executed by the computing unit 401, one or more steps of the methods described above can be performed. Alternatively, in other embodiments, the computing unit 401 can be configured to execute the aforementioned method for identifying an address in any other suitable manner (e.g., by means of firmware).
[0192] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application Specific Standard Products), SoCs (System On Chip), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0193] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.
[0194] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a RAM, a ROM, an EPROM (Electrically Programmable Read-Only-Memory), or a flash memory, an optical fiber, a CD-ROM (Compact Disc Read-Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0195] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (Cathode-Ray Tube) or an LCD (Liquid Crystal Display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball), through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and the input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0196] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: LAN (Local Area Network), WAN (Wide Area Network), the Internet, and blockchain networks.
[0197] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS"). The server can also be a server of a distributed system, or a server combined with blockchain.
[0198] Among them, it should be noted that artificial intelligence is a discipline that studies to make a computer simulate certain thinking processes and intelligent behaviors of humans (such as learning, reasoning, thinking, planning, etc.), and there are both hardware-level technologies and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech recognition technology, natural language processing technology, and machine learning / deep learning, big data processing technology, and knowledge graph technology.
[0199] It should be understood that various forms of processes shown above can be used, steps can be reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is made herein.
[0200] The above specific embodiments do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. A method for identifying an address, characterized in that: include: Obtain at least two geographic data of target customers; wherein different geographic data have different data sources; Marking the geographic data in grids on the digitized map respectively to construct a map matrix; wherein different geographic data corresponds to different map matrices; The similarity function values between the different map matrices are calculated based on a preset coordinate relationship recognition model, and when the similarity function value is greater than a preset threshold, the geographic data is determined to be a correct address.
2. The method according to claim 1, characterized in that The calculating of similarity function values between different map matrices based on the preset coordinate relationship recognition model includes: respectively calculating the covariance matrices between the different map matrices; Calculating the similarity function value according to the covariance matrix and the geographic data location comprehensive function; The geographic data location comprehensive function is the sum of the location information matrices of each geographic data; the location information matrix includes the address code.
3. The method according to claim 2, characterized in that The calculating of similarity function values between different map matrices according to the covariance matrix comprises: Calculating a similarity function based on the covariance matrix and the geographic data location comprehensive function; All elements in the similarity function are accumulated to obtain the similarity function value.
4. The method according to claim 3, characterized in that When the similarity function value is greater than a preset threshold, determining that the geographic data is a correct address includes: The rating of the similarity function value is determined based on a preset scoring threshold; wherein a high rating corresponds to a high accuracy of the address.
5. The method according to claim 1, characterized in that The step of marking the geographic data in grids on the digitized map to construct a map matrix comprises: Performing grid division based on a preset scale and coordinate information in each of the geographic data; A map matrix is generated according to each of the divided grids, and the coordinate information is marked in the map matrix.
6. The method according to claim 5, characterized in that After calculating the similarity function values between different map matrices based on the preset coordinate relationship recognition model, the method further includes: In the case where the geographic data is not a correct address, obtaining field survey geographic data and constructing a digital matrix; The field survey geographical data is used as the correct address of the target customer, and the preset coordinate relationship recognition model is trained according to the digital matrix.
7. The method according to any one of claims 1 to 6, characterized in that The geographic data includes at least two of registration information, business registration information, and station information.
8. An address recognition device, characterized in that: include: A first acquisition unit is used to acquire at least two geographic data of target customers; wherein different geographic data have different data sources; A construction unit, used to grid-mark the geographic data on the digitized map to construct a map matrix; wherein different geographic data corresponds to different map matrices; The determination unit is used to calculate the similarity function value between different map matrices based on a preset coordinate relationship recognition model, and determine that the geographic data is a correct address when the similarity function value is greater than a preset threshold.
9. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-7.
11. A computer program product, characterized in that The invention comprises a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 7.