Operator standard address data matching method, system and device based on CTF-ICF-TP improved SIMHASH algorithm and medium
By improving the SIMHASH algorithm and fine-tuning large model to correct the standard address of user input, the problem of inaccurate address matching in the prior art is solved, and the accuracy of standard address matching and search intelligence are improved.
Patent Information
- Application Number
- CN202411833901.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-13
- Publication Date
- 2025-05-13
AI Technical Summary
The existing address matching methods are inaccurate or inaccurate when matching the standard address results, which affects the development of front-end services.
The SIMHASH algorithm improved based on CTF-ICF-TP is used to correct the standard address of user input by fine-tuning the large model, and sort the recall results in similarity to improve the accuracy of standard address matching.
The accuracy of standard address matching and the ability of application systems to adapt to the environment are improved. Through the introduction of large models, the intelligence of search is improved, and the search time and error rate are significantly reduced.
Smart Images

Figure CN119988990A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and in particular to a method, system, device and medium for matching operator standard address data based on a CTF-ICF-TP improved SIMHASH algorithm. Background Art
[0002] Operators have accumulated considerable address data due to the popularity of broadband services. Through address standardization technology, addresses are standardized and managed for easy use. Standardized addresses are called standard addresses of operators. Standard addresses are divided into several levels, including province, city, district, county, street, community, building, unit, floor and household number, etc. When accepting broadband services, the front-end staff of the business hall conducts address matching search based on the community information provided by the user. The current method is to drill down layer by layer through the tree structure and directly call the database statement query or enter the accurate community name to use Solr for query matching. Solr uses Tokenizer to perform word segmentation and divides the query text into a series of vocabulary units, which are called tokens. Tokens are processed by using TokenFilter. For example, LowercaseFilter is used to convert tokens to lowercase, StopFilter is used to remove stop words, and StemmingFilter is used to extract words from stems. When querying, Solr parses the query entered by the user and splits the query string into a series of query items. Each query item contains a field name, an operator and query content. Such as Boolean query, phrase query, wildcard query, etc. However, in some cases, incorrect characters, incomplete characters, homophones or synonyms are entered. Since many abbreviations or aliases provided by community names or business customers do not have specific meanings, synonyms cannot be fully understood and matched in a natural language manner. This results in inaccurate or unavailable results when matching standard addresses, affecting the development of front-end business. Summary of the invention
[0003] Aiming at the problem that the existing address matching methods are inaccurate or cannot find the results when matching the standard address, which affects the development of the front-end business, the present invention proposes an operator standard address data matching method, system, device and medium based on the CTF-ICF-TP improved SIMHASH algorithm; the method corrects the standard address input by the user through fine-tuning the large model, and calculates the similarity ranking of the recall results according to the improved SIMHASH algorithm, thereby improving the accuracy of standard address matching and the ability of the application system to adapt to the environment, and through the introduction of the large model, the intelligence of the search is improved.
[0004] The specific implementation contents of the present invention are as follows:
[0005] A method for matching operator standard address data based on CTF-ICF-TP improved SIMHASH algorithm, specifically comprising the following steps:
[0006] Step S1: pre-processing the operator's standard address data and saving it to the vector database;
[0007] Step S2: Based on the pre-processed operator standard address data, call the Llama-Factory framework to supervise and fine-tune the large model to obtain the standard address large model;
[0008] Step S3: Calling the API interface to input the user query data into the standard address large model, and recalling the standard address according to the vector database;
[0009] Step S4: calling the SIMHASH algorithm to calculate the similarity ranking between the recalled standard addresses and the user query data, and obtaining the address matching result according to the ranking result.
[0010] In order to better implement the present invention, further, the step S4 specifically includes the following steps:
[0011] Step S41: segmenting the recalled standard address and user query data content and calculating weights to obtain a plurality of standard address feature segmentations and a plurality of user query data feature segmentations;
[0012] Step S42: Obtain a standard address feature vector and a user query data feature vector according to the standard address feature segmentation and the user query data feature segmentation, and assign weights;
[0013] Step S43: according to the standard address feature vector and the user query data feature vector, a hash function is called to calculate the hash value of the standard address feature vector and the hash value of the user query data feature vector;
[0014] Step S44: Calculate the standard address weight vector according to the standard address feature weight and the standard address feature vector hash value; calculate the user query data weight vector according to the user query data feature vector weight and the user query data feature vector hash value;
[0015] Step S45: merging the standard address weight vectors to obtain a new standard address weight vector; merging the user query data weight vectors to obtain a new user query data weight vector;
[0016] Step S46: Reduce the dimension of the new standard address weight vector and the new user query data weight vector to obtain the standard address SimHash value and the user query data SimHash value;
[0017] Step S47: Calculate the Hamming distance between the standard address SimHash value and the user query data SimHash value, sort them according to the Hamming distance, and obtain the address matching result.
[0018] In order to better implement the present invention, further, the specific operation of step S42 is: according to the address information word frequency of the current text of the current category, the number of categories, the position score of the current feature participle in the standard address, and the median position of the current feature participle in the sentence, calculate the standard address feature vector weight and the user query data feature vector weight.
[0019] In order to better implement the present invention, further, the specific operation process of step S42 is:
[0020]
[0021] in, represents the address information word frequency of the i-th text of category a, represents the frequency of words containing term t, represents the frequency of term t in category a, M represents the total category, PW(t) represents the position score of the term in the standard address, and Median(Sent) represents the median position of the term in all sentences.
[0022] In order to better implement the present invention, further, the step S1 specifically includes the following steps:
[0023] Step S11: vectorization preprocessing obtains operator standard address data from the relational database and saves it to the vectorization database;
[0024] Step S12: convert the standard address cell level data obtained by Pinyin to full Pinyin and save it to the vectorized database.
[0025] In order to better implement the present invention, further, the specific operation of step S2 is: first, the preprocessed operator standard address data is loaded according to the set json array data format, and then according to the Llama-Factory framework, the LORA fine-tuning algorithm is called, and the supervised fine-tuning SFT is called to train the Qwen2.5-32B large model to obtain the standard address large model.
[0026] In order to better implement the present invention, further, the step S3 specifically includes the following steps:
[0027] Step S31: inputting user query data into the standard address big model according to the API interface of the standard address big model;
[0028] Step S32: calling the nearest neighbor search algorithm to correct the query content of the user query data and generate an error correction result;
[0029] Step S33: Recall the standard address from the vector database according to the error correction result.
[0030] Based on the above-mentioned operator standard address data matching method based on CTF-ICF-TP improved SIMHASH algorithm, in order to better realize the present invention, further, a operator standard address data matching system based on CTF-ICF-TP improved SIMHASH algorithm is proposed, which is used to execute the above-mentioned operator standard address data matching method based on CTF-ICF-TP improved SIMHASH algorithm; including a preprocessing unit, a supervision fine-tuning unit, a recall unit, and a matching unit;
[0031] The preprocessing unit is used to preprocess the operator standard address data and save it to the vector database;
[0032] The supervision and fine-tuning unit is used to call the Llama-Factory framework to supervise and fine-tune the large model according to the pre-processed operator standard address data to obtain the standard address large model;
[0033] The recall unit is used to call the API interface to input the user query data into the standard address large model, and recall the standard address according to the vector database;
[0034] The matching unit is used to call the SIMHASH algorithm to sort the recalled standard addresses according to the user query data, and obtain the address matching result according to the similarity value.
[0035] Based on the above-mentioned operator standard address data matching method based on the CTF-ICF-TP improved SIMHASH algorithm, in order to better realize the present invention, further, an electronic device is proposed, including a memory and a processor; a computer program is stored on the memory; when the computer program is executed on the processor, the above-mentioned operator standard address data matching method based on the CTF-ICF-TP improved SIMHASH algorithm is implemented.
[0036] Based on the above-mentioned operator standard address data matching method based on the CTF-ICF-TP improved SIMHASH algorithm, in order to better realize the present invention, further, a computer-readable storage medium is proposed, on which computer instructions are stored; when the computer instructions are executed on the above-mentioned electronic device, the above-mentioned operator standard address data matching method based on the CTF-ICF-TP improved SIMHASH algorithm is implemented.
[0037] The present invention has the following beneficial effects:
[0038] (1) The present invention corrects the standard address input by the user by fine-tuning the large model, and uses the improved SIMHASH algorithm to sort the recall results by similarity, thereby improving the accuracy of standard address matching and the ability of the application system to adapt to the environment. The introduction of the large model improves the intelligence of the search.
[0039] (2) The present invention can find the target address faster, reduce the search time and error rate, and significantly improve the search efficiency.
[0040] (3) The present invention uses the search to match the closest standard address, which can ensure that even if the user inputs an address in a non-standard format, the most relevant correct address can be found, thereby improving the search accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 A schematic diagram of the SIMHASH algorithm process provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0042] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. It should be understood that the described embodiments are only part of the embodiments of the present invention, not all of the embodiments, and therefore should not be regarded as limiting the scope of protection. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technical personnel in this field without making creative work are within the scope of protection of the present invention.
[0043] In the description of the present invention, it should be noted that, unless otherwise clearly specified and limited, the terms "disposed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be an indirect connection through an intermediate medium, or it can be the internal communication of two elements. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0044] Embodiment 1:
[0045] This embodiment proposes a method for matching operator standard address data based on the CTF-ICF-TP improved SIMHASH algorithm, which specifically includes the following steps:
[0046] Step S1: pre-processing the operator's standard address data and saving it to the vector database;
[0047] Furthermore, the step S1 specifically includes the following steps:
[0048] Step S11: vectorization preprocessing obtains operator standard address data from the relational database and saves it to the vectorization database;
[0049] Step S12: convert the standard address cell level data obtained by Pinyin to full Pinyin and save it to the vectorized database.
[0050] Step S2: Based on the pre-processed operator standard address data, call the Llama-Factory framework to supervise and fine-tune the large model to obtain the standard address large model;
[0051] Furthermore, the specific operation of step S2 is: first, the preprocessed operator standard address data is loaded according to the set json array data format, and then according to the Llama-Factory framework, the LORA fine-tuning algorithm is called, and the supervised fine-tuning SFT is called to train the Qwen2.5-32B large model to obtain the standard address large model.
[0052] Step S3: Calling the API interface to input the user query data into the standard address large model, and recalling the standard address according to the vector database;
[0053] Furthermore, the step S3 specifically includes the following steps:
[0054] Step S31: According to the API interface of the standard address big model, the user query data is input into the standard address big model.
[0055] Step S32: calling the nearest neighbor search algorithm to correct the query content of the user query data and generate an error correction result;
[0056] Step S33: Recall the standard address from the vector database according to the error correction result.
[0057] Step S4: calling the SIMHASH algorithm to calculate the similarity ranking between the recalled standard addresses and the user query data, and obtaining the address matching result according to the ranking result.
[0058] Furthermore, the step S4 specifically includes the following steps:
[0059] Step S41: segmenting the recalled standard address and user query data content and calculating weights to obtain a plurality of standard address feature segmentations and a plurality of user query data feature segmentations;
[0060] Step S42: Obtain a standard address feature vector and a user query data feature vector according to the standard address feature segmentation and the user query data feature segmentation, and assign weights;
[0061] Furthermore, the specific operation of step S42 is: according to the address information word frequency of the current text of the current category, the number of categories, the position score of the current feature word in the standard address, and the median position of the current feature word in the sentence, calculate the standard address feature vector weight and the user query data feature vector weight.
[0062] The specific operation process of step S42 is as follows:
[0063]
[0064] in, represents the address information word frequency of the i-th text of category a, represents the frequency of words containing term t, represents the frequency of term t in category a, M represents the total category, PW(t) represents the position score of the term in the standard address, and Median(Sent) represents the median position of the term in all sentences.
[0065] Step S43: according to the standard address feature vector and the user query data feature vector, a hash function is called to calculate the hash value of the standard address feature vector and the hash value of the user query data feature vector;
[0066] Step S44: Calculate the standard address weight vector according to the standard address feature weight and the standard address feature vector hash value; calculate the user query data weight vector according to the user query data feature vector weight and the user query data feature vector hash value;
[0067] Step S45: merging the standard address weight vectors to obtain a new standard address weight vector; merging the user query data weight vectors to obtain a new user query data weight vector;
[0068] Step S46: Reduce the dimension of the new standard address weight vector and the new user query data weight vector to obtain the standard address SimHash value and the user query data SimHash value;
[0069] Step S47: Calculate the Hamming distance between the standard address SimHash value and the user query data SimHash value, sort them according to the Hamming distance, and obtain the address matching result.
[0070] Working principle: This embodiment corrects the standard address of the user input by fine-tuning the large model, and calculates the similarity ranking of the recall results according to the improved SIMHASH algorithm, thereby improving the accuracy of standard address matching and the ability of the application system to adapt to the environment. The introduction of the large model improves the intelligence of the search.
[0071] Embodiment 2:
[0072] This embodiment is based on the above embodiment 1. Figure 1 As shown, a specific embodiment is described in detail.
[0073] This embodiment integrates large-model intelligent error correction, vector recall, pinyin recall and other technologies to achieve intelligent search, association and error correction of user-entered addresses, which can effectively improve the accuracy of standard address retrieval.
[0074] Step S1: The standard address data stored in the relational database must be vectorized and stored in batches. The standard address data must be vectorized using TF-IDF (Term Frequency-Inverse Document Frequency). The standard address data must be preprocessed and saved in the vector database for later use in vector recall.
[0075] The data of the standard address community level is converted into full pinyin, and vectorized preprocessing is also performed and saved in the vector database, so that the vector of pinyin can be recalled in the subsequent steps.
[0076] Step S2: Perform a supervised fine-tuning on the Qwen2.5-32B large model based on Llama-Factory, and load the standard address dataset in the agreed json array data format. Select Lora for the fine-tuning method and SFT for the training phase. After the training is completed, the output is a trained standard address large model with optimized parameters. Perform error correction on the user input information based on the large model.
[0077] Step S3: Smart matching First, use the full pinyin method to approximate the recall of homophones for standard addresses. There are a large number of standard addresses, so the computational overhead of accurate nearest neighbor search is very high on large-scale data sets. Therefore, we use the approximate nearest neighbor search algorithm HNSW to improve the retrieval efficiency. Its principle is the multi-layer graph structure of the vector library, that is, to build a small world graph navigation search containing multiple levels: starting from the high level, searching down layer by layer until the nearest neighbor is found. [Hierarchical Navigable Small World]. This improves the high query efficiency and accuracy of pinyin retrieval, and can support dynamic updates. The disadvantage is that the memory consumption will be relatively large. Recall several standard address data with high similarity.
[0078] The query content input by the user is corrected for standard addresses through a large model for standard address fine-tuning, and the content input by the user is corrected through the API interface of the large model. At the same time, to ensure that the large model stably returns content recognizable by the program, the information we need is returned through prompt engineering. For example:
【We want to create a function for address retrieval, and the addresses involved are in xx Province. The input addresses may have typos, such as homophones. As a vector retrieval assistant, your task is to give the corrected suggested address in JSON format based on the input address information, and the return field name is address_suggestion;
[0079] Example: Input "亲亲家元", after address correction, return {"address_suggestion": "亲亲家园"}】;
[0080] Then, based on the correction result, the standard address is recalled again in the vector library.
[0081] Based on the original content input by the user, the standard address is directly recalled in the vector library.
[0082] Step S4: Aggregate the standard address data recalled in different ways and perform data deduplication processing.
[0083] Then, according to the content input by the user, the list data of the recalled standard addresses is sorted by similarity. Here, the sorting is implemented using an improved SIMHASH algorithm. Since the text of the standard address is short and important terms basically do not repeat in a standard address, in order to improve the recall accuracy, the weight calculation method is optimized and the SIMHASH algorithm is improved. The similarity between the information input by the user and each standard address in the recalled list is calculated to obtain a similarity value, and they are arranged in descending order of the similarity value to return the final intelligent query result. The specific logic is as follows:
[0084] The SIMHASH algorithm mainly consists of five processes: word segmentation, hashing, weighting, merging, and dimensionality reduction; the example is as Figure 1 shown.
[0085] 1. Word segmentation and weight calculation.
[0086] Segment the standard address or the content input by the user and calculate the weights, generating n feature words and assigning each feature word a weight. The weight calculation is done by aggregating the standard address texts within the same district and county of the same city into a large text, and then calculating the frequency of the feature segmentation, which is called CTF. The higher the frequency of the feature segmentation in all texts of this category, the greater the CTF score. Additionally, calculate the importance score based on the position of the segmentation in the sentence, and the weight of the segmentation with a specific label will also increase accordingly. Use the inverse class frequency ICF score to calculate the discrimination in different categories. term represents the feature segmentation, and the specific calculation formula is as follows:
[0087]
[0088] Represents the word frequency of the address information of the i-th text in category a, is the word frequency containing term t, is the word frequency of term t in category a. M represents the total number of categories (i.e., the total number of regions); category a is the category of the regional division of the standard address. For example, Yuhua District, Shijiazhuang City is one of the categories; Represents the weight of term t in text d.
[0089] In addition, PW(t) is the position score of the segmentation in the standard address. When users use it, the demand for querying the names of communities / villages is relatively large. The position of the community or village in the standard address is more important than other segmentations. Median(Sen t ) is the median position of this word in all sentences.
[0090] For example, given a standard address: "Qingyuan Street, Qinqin Jiayuan, Yuhua District, Shijiazhuang, Hebei", after segmentation, the result is: Hebei Shijiazhuang Yuhua District Qingyuan Street Qinqin Jiayuan. Then assign weights to each feature vector: Hebei(2) Shijiazhuang(3) Yuhua District(4) Qingyuan Street(2) Qinqin(5) Jiayuan(5), where the numbers in the brackets represent the importance of this word in the whole sentence. The larger the number, the more important it is.
[0091] 2.Hash.
[0092] Calculate the hash value of each feature vector through the hash function. The hash value is an n-bit signature composed of binary numbers 0 and 1. For example, the hash value of "Qinqin", Hash(Qinqin), is 110101, and the hash value of "Jiayuan", Hash(Jiayuan), is "101001". In this way, the string becomes a series of numbers.
[0093] 3.Weighting.
[0094] After the previous calculations, the Hash strings of each word vector and the weights corresponding to the word vectors have been obtained. In the third step, calculate the weight vector W = hash * weight. When encountering 1, multiply the hash value and the weight value positively; when encountering 0, multiply the hash value and the weight value negatively. For example, weighting the hash value "110101" of "亲亲" gives: W(亲亲) = 110101 * 5 = 5 5 -5 5 -5 5, and weighting the hash value "101001" of "家园" gives: W(家园) = 101001 * 2 = 2 -2 2 -2 -2 2. Similar operations are performed on the remaining feature vectors.
[0095] 4. Merge.
[0096] For a text, the weight vectors of each feature word after text segmentation have been calculated. In this merging stage, add up the weight vectors of all the word vectors in the text to obtain a new weight vector. Taking the first two feature vectors as an example, for instance, adding "5 5 -5 5 -5 5" of "亲亲" and "2 -2 2 -2 -2 2" of "家园" gives "5+2, 5-2, -5+2, 5-2, -5-2, 5+2,", resulting in "7 3 -3 3 -7 7".
[0097] 5. Dimensionality reduction.
[0098] For the weight vector of the text obtained after the previous merging, set 1 for positions greater than 0 and 0 for positions less than or equal to 0, and the SimHash value of the text can be obtained.
[0099] Finally, calculate the similarity between the user input content and the standard address based on the Hamming distance of the SimHash of different statements. For example, reducing the calculated "9 -9 1 -1 1 9" above (recording 1 for a certain bit greater than 0 and 0 for less than 0), the resulting 01 string is: "110101", thus forming their SimHash signature.
[0100] After calculating the signature for each text according to SimHash, calculate the Hamming distance between the two signatures (the number of 1s after the two binary numbers are XORed). That is, the number of different binary (01 string) values corresponding to the two SimHashes is called the Hamming distance between the two SimHashes.
[0101] For example: For 10101 and 00110, starting from the first bit, the first, fourth, and fifth bits are different in turn, so the Hamming distance is 3. For binary strings a and b, the Hamming distance is equal to the number of 1s in the result of the a XOR b operation (general algorithm).
[0102] By using Redis query, divide the 64-bit SimHash into 4 pieces of content and store them as keys in Redis;
[0103] Accurately query the first 16 digits;
[0104] The SIMHASH distance between the calculation and the user input content is verified by the standard address data, and it is set to be similar if it is less than 3. Finally, it is sorted by distance to complete the standard address intelligent matching logic calculation.
[0105] Working principle: By optimizing the standard address matching algorithm and sorting method, the search efficiency is significantly improved, and users can find the target address faster, reducing search time and error rate. The search accuracy is improved. By searching and matching the closest standard address, it can ensure that even if the user enters an address in a non-standard format, the most relevant correct address can be found. The user experience is improved: in terms of user interaction, user perception and satisfaction are improved.
[0106] This embodiment selects several hundred test data in batches to test the search algorithm, and the final result can achieve an accuracy rate of more than 99%. Table 1 captures a part of the test data.
[0107] Table 1 Test data schematic
[0108]
[0109] This embodiment uses vector library, word segmentation, large model and simhash algorithm to solve the accuracy problem of standard address search. The core innovation is to fine-tune the large model to correct the standard address input by the user, and the calculation method of similarity sorting of the recall results based on the improved SIMHASH algorithm. Ultimately, the accuracy of standard address matching is improved, the ability of the application system to adapt to the environment is improved, and the intelligence of the search is improved through the introduction of the large model.
[0110] The other parts of this embodiment are the same as those of the above-mentioned embodiment 1, and thus will not be described in detail.
[0111] Embodiment 3:
[0112] This embodiment, based on any one of the above-mentioned embodiments 1-2, proposes an operator standard address data matching system based on the CTF-ICF-TP improved SIMHASH algorithm, which is used to execute the above-mentioned operator standard address data matching method based on the CTF-ICF-TP improved SIMHASH algorithm; it includes a preprocessing unit, a supervision fine-tuning unit, a recall unit, and a matching unit;
[0113] The preprocessing unit is used to preprocess the operator standard address data and save it to the vector database;
[0114] The supervision and fine-tuning unit is used to call the Llama-Factory framework to supervise and fine-tune the large model according to the pre-processed operator standard address data to obtain the standard address large model;
[0115] The recall unit is used to call the API interface to input the user query data into the standard address large model, and recall the standard address according to the vector database;
[0116] The matching unit is used to call the SIMHASH algorithm to sort the recalled standard addresses according to the user query data, and obtain the address matching result according to the similarity value.
[0117] This embodiment also proposes an electronic device, including a memory and a processor; a computer program is stored on the memory; when the computer program is executed on the processor, the above-mentioned operator standard address data matching method based on the CTF-ICF-TP improved SIMHASH algorithm is implemented.
[0118] This embodiment also proposes a computer-readable storage medium, on which computer instructions are stored; when the computer instructions are executed on the above-mentioned electronic device, the above-mentioned operator standard address data matching method based on the CTF-ICF-TP improved SIMHASH algorithm is implemented.
[0119] The other parts of this embodiment are the same as any one of the above-mentioned embodiments 1-2, so they will not be repeated here.
[0120] The above description is only a preferred embodiment of the present invention and does not limit the present invention in any form. Any simple modification or equivalent change made to the above embodiment based on the technical essence of the present invention shall fall within the protection scope of the present invention.
Claims
1. A method for matching operator standard address data based on CTF-ICF-TP improved SIMHASH algorithm, characterized in that: The specific steps include: Step S1: pre-processing the operator's standard address data and saving it to the vector database; Step S2: Based on the pre-processed operator standard address data, call the Llama-Factory framework to supervise and fine-tune the large model to obtain the standard address large model; Step S3: Calling the API interface to input the user query data into the standard address large model, and recalling the standard address according to the vector database; Step S4: calling the SIMHASH algorithm to calculate the similarity ranking between the recalled standard addresses and the user query data, and obtaining the address matching result according to the ranking result.
2. According to claim 1, a method for matching operator standard address data based on CTF-ICF-TP improved SIMHASH algorithm is characterized in that: The step S4 specifically comprises the following steps: Step S41: segmenting the recalled standard address and user query data content and calculating weights to obtain a plurality of standard address feature segmentations and a plurality of user query data feature segmentations; Step S42: Obtain a standard address feature vector and a user query data feature vector according to the standard address feature segmentation and the user query data feature segmentation, and assign weights; Step S43: according to the standard address feature vector and the user query data feature vector, a hash function is called to calculate the hash value of the standard address feature vector and the hash value of the user query data feature vector; Step S44: Calculate the standard address weight vector according to the standard address feature weight and the standard address feature vector hash value; calculate the user query data weight vector according to the user query data feature vector weight and the user query data feature vector hash value; Step S45: merging the standard address weight vectors to obtain a new standard address weight vector; merging the user query data weight vectors to obtain a new user query data weight vector; Step S46: Reduce the dimension of the new standard address weight vector and the new user query data weight vector to obtain the standard address SimHash value and the user query data SimHash value; Step S47: Calculate the Hamming distance between the standard address SimHash value and the user query data SimHash value, sort them according to the Hamming distance, and obtain the address matching result.
3. The method for matching operator standard address data based on CTF-ICF-TP improved SIMHASH algorithm according to claim 2, characterized in that: The specific operation of step S42 is: according to the address information word frequency of the current text of the current category, the number of categories, the position score of the current feature word in the standard address, and the median position of the current feature word in the sentence, calculate the standard address feature vector weight and the user query data feature vector weight.
4. The method for matching operator standard address data based on the CTF-ICF-TP improved SIMHASH algorithm according to claim 3, characterized in that: The specific operation process of step S42 is as follows: in, represents the address information word frequency of the i-th text of category a, represents the frequency of words containing term t, represents the frequency of term t in category a, M represents the total number of categories, PW(t) represents the position score of the segmented word in the standard address, Median(Sen t ) represents the median position of the word in all sentences.
5. The method for matching operator standard address data based on CTF-ICF-TP improved SIMHASH algorithm according to claim 1, characterized in that: The step S1 specifically includes the following steps: Step S11: vectorization preprocessing obtains operator standard address data from the relational database and saves it to the vectorization database; Step S12: convert the standard address cell level data obtained by Pinyin to full Pinyin and save it to the vectorized database.
6. The method for matching operator standard address data based on CTF-ICF-TP improved SIMHASH algorithm according to claim 1, characterized in that: The specific operation of step S2 is: first, the pre-processed operator standard address data is loaded according to the set json array data format, and then according to the Llama-Factory framework, the LORA fine-tuning algorithm is called, and the supervised fine-tuning SFT is called to train the Qwen2.5-32B large model to obtain the standard address large model.
7. The method for matching operator standard address data based on CTF-ICF-TP improved SIMHASH algorithm according to claim 1, characterized in that: The step S3 specifically comprises the following steps: Step S31: inputting user query data into the standard address large model according to the API interface of the standard address large model; Step S32: calling the nearest neighbor search algorithm to correct the query content of the user query data and generate an error correction result; Step S33: Recall the standard address from the vector database according to the error correction result.
8. An operator standard address data matching system based on CTF-ICF-TP improved SIMHASH algorithm, used to execute the operator standard address data matching method based on CTF-ICF-TP improved SIMHASH algorithm as claimed in claim 1; characterized in that: It includes pre-processing unit, supervised fine-tuning unit, recall unit and matching unit; The preprocessing unit is used to preprocess the operator standard address data and save it to the vector database; The supervision and fine-tuning unit is used to call the Llama-Factory framework to supervise and fine-tune the large model according to the pre-processed operator standard address data to obtain the standard address large model; The recall unit is used to call the API interface to input the user query data into the standard address large model, and recall the standard address according to the vector database; The matching unit is used to call the SIMHASH algorithm to sort the recalled standard addresses according to the user query data, and obtain the address matching result according to the similarity value.
9. An electronic device, characterized in that: It comprises a memory and a processor; a computer program is stored in the memory; when the computer program is executed on the processor, the operator standard address data matching method based on the CTF-ICF-TP improved SIMHASH algorithm as described in any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that: The computer readable storage medium has computer instructions stored thereon; When the computer instructions are executed on the electronic device as claimed in claim 9, the operator standard address data matching method based on the CTF-ICF-TP improved SIMHASH algorithm as claimed in any one of claims 1 to 7 is implemented.