Address book user searching method and system
By performing word segmentation and TF-IDF calculations on the query statements and user information in the address book user search, query and user vectors are constructed, and cosine similarity is used to match and sort, the problem of inefficiency in traditional address book search is solved, and more efficient and accurate search results are achieved.
Patent Information
- Application Number
- CN202510030062.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-05-09
AI Technical Summary
Traditional address book user search methods are inefficient, making it difficult to quickly and accurately find the required contacts in large-scale contact information, and lacks intelligent search suggestions and result sorting, resulting in poor user experience.
By performing word segmentation processing on the input query statement, the word frequency-inverse document frequency (TF-IDF) of the query word segmentation and user information in the address book is calculated, the query vector and user vector are constructed, and the cosine similarity is used to match, and the most relevant user information is sorted and returned.
It significantly improves the accuracy and efficiency of search results, saves user time, improves the overall efficiency of search, and is suitable for address book user search in different fields and locale environments.
Smart Images

Figure CN119961503A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a method and system for searching address book users, belonging to the technical field of data retrieval. Background Art
[0002] With the rapid development of information technology, people's demand for address book user search is increasing. Both individual users and corporate users hope to find the required contacts quickly and accurately. Traditional technologies usually use simple keyword matching to search, requiring users to manually enter keywords and browse search results. This method is not only time-consuming and labor-intensive, but also prone to missed searches or mis-searches due to inaccurate keywords. When faced with a large amount of contact information, traditional search methods are inefficient and users need to spend a long time to find the required contact information. Due to the lack of intelligent search suggestions and result sorting, users have a poor experience during the search process and need to manually filter and judge the relevance of search results.
[0003] The patent document with the patent number "CN115687870A" has developed a place name matching method based on matrix operations. The problem with this method is that although the matrix operation and vector inner product algorithm can achieve accurate matching of place names, it may face the problems of large amount of calculation and low search efficiency when processing large-scale place name databases. Summary of the invention
[0004] In order to solve the above problems existing in the prior art, the present invention proposes a method and system for searching address book users.
[0005] The technical solution of the present invention is as follows:
[0006] In one aspect, the present invention provides a method for searching for users in a contact list, comprising the following steps:
[0007] Perform word segmentation processing on the input query statement according to the preset character length to obtain a query word segmentation set;
[0008] Calculate the term frequency-inverse document frequency of each query word in the query word set and each user information in the address book to obtain a two-dimensional query vector;
[0009] Perform word segmentation on each user information in the address book to obtain several user word segmentation sets, calculate the word frequency-inverse document frequency of each user word in the user word segmentation set and each user information in the address book to obtain a 3D user vector;
[0010] Decomposing the 3D user vector into a number of 2D user vectors, calculating the cosine similarity between the 2D query vector and the number of 2D user vectors to obtain a number of similarity values, and storing the number of similarity values in a similarity array;
[0011] The similarity values corresponding to the subscripts of the similarity array are sorted from high to low, and the user information corresponding to the first z similarity values is obtained as the address book search result according to the preset return number z.
[0012] As a preferred implementation, the word segmentation method is:
[0013] Determine the preset character length and segment the query statement according to the preset length:
[0014] PQ=P(Q);
[0015] P(Q)={w1,w i ,...,w n};
[0016] Among them, PQ represents the query word set, P represents the word processing function, Q represents the query statement, and w i Represents the i-th query word, and n represents the preset maximum number.
[0017] As a preferred implementation, the query vector is calculated as follows:
[0018] S1, let i = 1, initialize the query vector array cx[][];
[0019] S2, let j = 1, where i and j represent numeric variables;
[0020] S3. Calculate w i And the term frequency-inverse document frequency of the j-th user information d:
[0021] TF_IDF(w i ,d)=TF(w i ,d)*IDF(w i );
[0022]
[0023] Among them, TF_IDF(w i ,d) represents the i-th query word w i The term frequency-inverse document frequency, TF(w i ,d) represents the i-th query word w i The word frequency of the jth user information d, IDF(w i ) represents the i-th query word w i and the inverse document frequency of the jth user information d, θ represents the i-th query word w i The number of times it appears in the j-th user information d, μ represents the total number of word segments in the j-th user information d, N represents the number of user information, and σ represents the number of words that contain the i-th query word w iThe number of j-th user information d;
[0024] The obtained TF_IDF(w i ,d) Store in cx[i][j];
[0025] S4. If j<n, execute step S3 and j=j+1, otherwise execute step S5;
[0026] S5. If i<n, execute step S2 and i=i+1, otherwise execute step S6;
[0027] S6. Use the query vector array cx[][] as the query vector;
[0028] The calculation method of user vector is:
[0029] F1. Let j = 1, initialize the user vector array yh[][][];
[0030] F2. Perform word segmentation processing on the j-th user information according to a preset character length to obtain a user word segmentation set g of the j-th user information;
[0031] F3, let k = 1, where k represents a numeric variable;
[0032] F4, let i = 1;
[0033] F5. Calculate the term frequency-inverse document frequency of the k-th user segmentation f of g and the i-th user information d:
[0034] TF_IDF(f,d)=TF(f,d)*IDF(f);
[0035]
[0036] Wherein, TF_IDF(f,d) represents the term frequency-inverse document frequency of the kth user segmentation f of the user segmentation set g of the jth user information and the i-th user information d, TF(f,d) represents the term frequency of the kth user segmentation f of the user segmentation set g of the jth user information and the i-th user information d, IDF(f) represents the inverse document frequency of the kth user segmentation f of the user segmentation set g of the jth user information and the i-th user information d, ε represents the number of times the kth user segmentation f of the user segmentation set g of the jth user information appears in the i-th user information d, γ represents the total number of segmentations of the i-th user information d, N represents the number of user information, and σ represents the number of i-th user information d of the kth user segmentation f of the user segmentation set g containing the j-th user information;
[0037] Store the obtained TF_IDF(f,d) into yh[j][k][i];
[0038] F6. If i<n, execute step F5 and i=i+1, otherwise execute step F7;
[0039] F7. If k<n, execute step F4 and k=k+1, otherwise execute step F8;
[0040] F8. If j<n, execute step F2 and j=j+1, otherwise execute step F9;
[0041] F9. Use the user vector array yh[][][] as the user vector.
[0042] As a preferred implementation, the similarity value is calculated as follows:
[0043] H1, let i=1, initialize the similarity array xsd[];
[0044] H2. Calculate the cosine similarity CS(cx[][],yh[i][][]) between the 2D query vector cx[][] and the i-th 2D user vector yh[i][][]:
[0045]
[0046]
[0047] Among them, · represents the dot product, ||cx[][]|| represents the modulus of cx[][], and cx[1][n*n] represents changing the cx[][] with n rows and n columns into a one-dimensional array cx[1][n*n] with 1 row and n*n columns.
[0048] Store CS(cx[][],yh[i][][]) into the similarity array xsd[i];
[0049] H3. If i<n, execute step H2 and i=i+1, otherwise execute step H4;
[0050] H4, step ends.
[0051] As a preferred implementation, the sorting method is:
[0052]
[0053] Among them, TopK[z] represents an array for storing z user information, sort(xsd[],i) means sorting xsd[] and taking out the similarity value of the i-th subscript, and top(sort(xsd[],i)) means obtaining the user information corresponding to sort(xsd[],i).
[0054] In another aspect, the present invention further provides a system for searching a user in an address book, comprising:
[0055] Query preprocessing module: performs word segmentation processing on the input query statement according to the preset character length to obtain a query word segmentation set;
[0056] Query vector construction module: calculate the word frequency-inverse document frequency of each query word in the query word set and each user information in the address book to obtain a two-dimensional query vector;
[0057] User vector construction module: Segment each user information in the address book to obtain several user segmentation sets, calculate the word frequency of each user segmentation in the user segmentation set and each user information in the address book - inverse document frequency to obtain a 3D user vector;
[0058] Similarity array construction module: decompose the 3D user vector into several 2D user vectors, calculate the cosine similarity between the 2D query vector and several 2D user vectors to obtain several similarity values, and store the several similarity values in the similarity array;
[0059] Address book search module: Sort the similarity values corresponding to each subscript of the similarity array from high to low, and obtain the user information corresponding to the first z similarity values as the address book search result according to the preset return number z.
[0060] The present invention has the following beneficial effects:
[0061] The present invention calculates the term frequency-inverse document frequency (TF-IDF) of query segmentation and user information, constructs query vector and user vector, and uses cosine similarity for matching, which can more accurately capture the correlation between query and user information, thereby significantly improving the accuracy of search results. By calculating the similarity value and sorting the results, the most relevant z user information can be quickly returned. This not only saves the user's time, but also improves the overall efficiency of the search. The segmentation processing and vector construction process can be adjusted and optimized according to different language habits and field characteristics. This enables the method to be widely used in address book user searches in different fields and different language environments, showing strong adaptability. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] Figure 1 The present invention is a flowchart for implementing the method. DETAILED DESCRIPTION
[0063] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0064] It should be understood that the step numbers used in this document are only for convenience of description and are not intended to limit the order in which the steps are executed.
[0065] It should be understood that the terms used in the present specification are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the present specification and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.
[0066] The terms “include” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or combinations thereof.
[0067] The term "and / or" means and includes any and all possible combinations of one or more of the associated listed items.
[0068] Embodiment 1:
[0069] See also Figure 1 The present invention provides a method for searching for address book users, comprising the following steps:
[0070] Perform word segmentation processing on the input query statement according to the preset character length to obtain a query word segmentation set;
[0071] Calculate the term frequency-inverse document frequency of each query word in the query word set and each user information in the address book to obtain a two-dimensional query vector;
[0072] The user information in the address book is segmented to obtain several user segmentation sets, and the word frequency-inverse document frequency of each user segmentation in the user segmentation set and each user information in the address book are calculated to obtain a three-dimensional user vector; the user information segmentation in this embodiment can be implemented by referring to the segmentation method of the query statement and will not be repeated here.
[0073] Decompose the 3D user vector into several 2D user vectors (decompose the first dimension subscript of the 3D array into n: decompose yh[][][] into yh[1][][], yh[2][][], ..., yh[n][][], then just take out yh[1][][], yh[2][][], ..., yh[n][][] in sequence to get n 2D arrays), calculate the cosine similarity between the 2D query vector and several 2D user vectors to get several similarity values, and store the several similarity values in the similarity array;
[0074] The similarity values corresponding to the subscripts of the similarity array are sorted from high to low, and the user information corresponding to the first z similarity values is obtained as the address book search result according to the preset return number z.
[0075] As a preferred implementation, the word segmentation method is:
[0076] Determine the preset character length and segment the query statement according to the preset length:
[0077] PQ=P(Q);
[0078] P(Q)={w1,w i ,...,w n};
[0079] Among them, PQ represents the query word set, P represents the word processing function, Q represents the query statement, and w i Represents the i-th query word, and n represents the preset maximum number.
[0080] As a preferred implementation, the query vector is calculated as follows:
[0081] S1, let i = 1, initialize the query vector array cx[][];
[0082] S2, let j = 1, where i and j represent numeric variables;
[0083] S3. Calculate w i And the term frequency-inverse document frequency of the j-th user information d:
[0084] TF_IDF(w i ,d)=TF(w i ,d)*IDF(w i );
[0085]
[0086] Among them, TF_IDF(w i ,d) represents the i-th query word w i The term frequency-inverse document frequency, TF(w i ,d) represents the i-th query word w i The word frequency of the jth user information d, IDF(w i ) represents the i-th query word w i and the inverse document frequency of the jth user information d, θ represents the i-th query word w i The number of times it appears in the j-th user information d, μ represents the total number of word segments in the j-th user information d, N represents the number of user information, and σ represents the number of words that contain the i-th query word w iThe number of j-th user information d;
[0087] The obtained TF_IDF(w i ,d) Store in cx[i][j];
[0088] S4. If j<n, execute step S3 and j=j+1, otherwise execute step S5;
[0089] S5. If i<n, execute step S2 and i=i+1, otherwise execute step S6;
[0090] S6. Use the query vector array cx[][] as the query vector;
[0091] The calculation method of user vector is:
[0092] F1. Let j = 1, initialize the user vector array yh[][][];
[0093] F2. Perform word segmentation processing on the j-th user information according to a preset character length to obtain a user word segmentation set g of the j-th user information;
[0094] F3, let k = 1, where k represents a numeric variable;
[0095] F4, let i = 1;
[0096] F5. Calculate the term frequency-inverse document frequency of the k-th user segmentation f of g and the i-th user information d:
[0097] TF_IDF(f,d)=TF(f,d)*IDF(f);
[0098]
[0099] Wherein, TF_IDF(f,d) represents the term frequency-inverse document frequency of the kth user segmentation f of the user segmentation set g of the jth user information and the i-th user information d, TF(f,d) represents the term frequency of the kth user segmentation f of the user segmentation set g of the jth user information and the i-th user information d, IDF(f) represents the inverse document frequency of the kth user segmentation f of the user segmentation set g of the jth user information and the i-th user information d, ε represents the number of times the kth user segmentation f of the user segmentation set g of the jth user information appears in the i-th user information d, γ represents the total number of segmentations of the i-th user information d, N represents the number of user information, and σ represents the number of i-th user information d of the kth user segmentation f of the user segmentation set g containing the j-th user information;
[0100] Store the obtained TF_IDF(f,d) into yh[j][k][i];
[0101] F6. If i<n, execute step F5 and i=i+1, otherwise execute step F7;
[0102] F7. If k<n, execute step F4 and k=k+1, otherwise execute step F8;
[0103] F8. If j<n, execute step F2 and j=j+1, otherwise execute step F9;
[0104] F9. Use the user vector array yh[][][] as the user vector.
[0105] As a preferred implementation, the similarity value is calculated as follows:
[0106] H1, let i=1, initialize the similarity array xsd[];
[0107] H2. Calculate the cosine similarity CS(cx[][],yh[i][][]) between the 2D query vector cx[][] and the i-th 2D user vector yh[i][][]:
[0108]
[0109] Among them, · represents the dot product, ||cx[][]|| represents the modulus of cx[][], and cx[1][n*n] represents changing the cx[][] with n rows and n columns into a one-dimensional array cx[1][n*n] with 1 row and n*n columns.
[0110] Let arr[n][n] be a two-dimensional array with n rows and n columns. At this time, we only need to sequentially concatenate the n-column elements of each row starting from the second row after the elements of the first row to obtain a one-dimensional array arr[1][n*n] with 1 row and n*n columns. The modulus calculation can calculate the transformed one-dimensional array ||cx[1][n*n]||. The modulus is the square of the elements corresponding to each subscript of the one-dimensional array and the sum of them. The square root of the sum is the modulus of the one-dimensional array. The dot product is the multiplication of the elements of the corresponding subscripts of the two one-dimensional arrays and then the sum.
[0111] In this embodiment, n represents a preset maximum number and is not limited here.
[0112] Store CS(cx[][],yh[i][][]) into the similarity array xsd[i];
[0113] H3. If i<n, execute step H2 and i=i+1, otherwise execute step H4;
[0114] H4, step ends.
[0115] As a preferred implementation, the sorting method is:
[0116]
[0117] Among them, TopK[z] represents an array for storing z user information, sort(xsd[],i) means sorting xsd[] and taking out the similarity value of the i-th subscript, and top(sort(xsd[],i)) means obtaining the user information corresponding to sort(xsd[],i).
[0118] Embodiment 2:
[0119] The present invention also provides a system for searching address book users, comprising:
[0120] Query preprocessing module: performs word segmentation processing on the input query statement according to the preset character length to obtain a query word segmentation set;
[0121] Query vector construction module: calculate the word frequency-inverse document frequency of each query word in the query word set and each user information in the address book to obtain a two-dimensional query vector;
[0122] User vector construction module: Segment each user information in the address book to obtain several user segmentation sets, calculate the word frequency of each user segmentation in the user segmentation set and each user information in the address book - inverse document frequency to obtain a 3D user vector;
[0123] Similarity array construction module: decompose the 3D user vector into several 2D user vectors, calculate the cosine similarity between the 2D query vector and several 2D user vectors to obtain several similarity values, and store the several similarity values in the similarity array;
[0124] Address book search module: Sort the similarity values corresponding to each subscript of the similarity array from high to low, and obtain the user information corresponding to the first z similarity values as the address book search result according to the preset return number z.
[0125] The system is used to implement the method in Example 1, which will not be described in detail here.
[0126] In the embodiments of the present application, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent the existence of A alone, the existence of A and B at the same time, and the existence of B alone. Among them, A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following" and similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b and c can be represented by: a, b, c, a and b, a and c, b and c, or a and b and c, where a, b, c can be single or multiple.
[0127] Those of ordinary skill in the art will appreciate that the various units and algorithm steps described in the embodiments disclosed herein can be implemented in a combination of electronic hardware, computer software, and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0128] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0129] In several embodiments provided in the present application, if any function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory; hereinafter referred to as: ROM), random access memory (Random Access Memory; hereinafter referred to as: RAM), disk or optical disk, and other media that can store program codes.
[0130] The above descriptions are merely embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A method for searching for users in a contact list, characterized in that: The following steps are involved: Perform word segmentation processing on the input query statement according to the preset character length to obtain a query word segmentation set; Calculate the term frequency-inverse document frequency of each query word in the query word set and each user information in the address book to obtain a two-dimensional query vector; Perform word segmentation on each user information in the address book to obtain several user word segmentation sets, calculate the word frequency-inverse document frequency of each user word in the user word segmentation set and each user information in the address book to obtain a 3D user vector; Decomposing the 3D user vector into a number of 2D user vectors, calculating the cosine similarity between the 2D query vector and the number of 2D user vectors to obtain a number of similarity values, and storing the number of similarity values in a similarity array; The similarity values corresponding to the subscripts of the similarity array are sorted from high to low, and the user information corresponding to the first z similarity values is obtained as the address book search result according to the preset return number z.
2. The method for searching for address book users according to claim 1, characterized in that: The word segmentation method is: Determine the preset character length and segment the query statement according to the preset length: PQ=P(Q); P(Q)={w1,w i ,...,w n }; Among them, PQ represents the query word set, P represents the word processing function, Q represents the query statement, and w i Represents the i-th query word, and n represents the preset maximum number.
3. The method for searching for address book users according to claim 2, characterized in that: The query vector is calculated as: S1, let i = 1, initialize the query vector array cx[][]; S2, let j = 1, where i and j represent numeric variables; S3. Calculate the i-th query word w i And the term frequency-inverse document frequency of the j-th user information d: TF_IDF(w i ,d)=TF(w i ,d)*IDF(w i ); Among them, TF_IDF(w i ,d) represents the i-th query word w i The term frequency-inverse document frequency, TF(w i ,d) represents the i-th query word w i The word frequency of the jth user information d, IDF(w i ) represents the i-th query word w i and the inverse document frequency of the jth user information d, θ represents the i-th query word w i The number of times it appears in the jth user information d, μ represents the total number of word segments in the jth user information d, N represents the number of user information, and σ represents the number of words that contain the i-th query word w i The number of j-th user information d; The obtained TF_IDF(w i ,d) Store in cx[i][j]; S4. If j<n, execute step S3 and j=j+1, otherwise execute step S5; S5. If i<n, execute step S2 and i=i+1, otherwise execute step S6; S6. Use the query vector array cx[][] as the query vector; The calculation method of user vector is: F1. Let j = 1, initialize the user vector array yh[][][]; F2. Perform word segmentation processing on the j-th user information according to a preset character length to obtain a user word segmentation set g of the j-th user information; F3, let k = 1, where k represents a numeric variable; F4, let i = 1; F5. Calculate the term frequency-inverse document frequency of the kth user segmentation f of the user segmentation set g of the jth user information and the i-th user information d: TF_IDF(f,d)=TF(f,d)*IDF(f); Wherein, TF_IDF(f,d) represents the term frequency-inverse document frequency of the kth user segmentation f of the user segmentation set g of the jth user information and the i-th user information d, TF(f,d) represents the term frequency of the kth user segmentation f of the user segmentation set g of the jth user information and the i-th user information d, IDF(f) represents the inverse document frequency of the kth user segmentation f of the user segmentation set g of the jth user information and the i-th user information d, ε represents the number of times the kth user segmentation f of the user segmentation set g of the jth user information appears in the i-th user information d, γ represents the total number of segmentations of the i-th user information d, N represents the number of user information, and σ represents the number of i-th user information d of the kth user segmentation f of the user segmentation set g containing the j-th user information; Store the obtained TF_IDF(f,d) into yh[j][k][i]; F6. If i<n, execute step F5 and i=i+1, otherwise execute step F7; F7. If k<n, execute step F4 and k=k+1, otherwise execute step F8; F8. If j<n, execute step F2 and j=j+1, otherwise execute step F9; F9. Use the user vector array yh[][][] as the user vector.
4. The method for searching for address book users according to claim 3, characterized in that: The similarity value is calculated as: H1, let i=1, initialize the similarity array xsd[]; H2. Calculate the cosine similarity CS(cx[][],yh[i][][]) between the 2D query vector cx[][] and the i-th 2D user vector yh[i][][]: Among them, · represents the dot product, ||cx[][]|| represents the modulus of cx[][], and cx[1][n*n] represents changing the cx[][] with n rows and n columns into a one-dimensional array cx[1][n*n] with 1 row and n*n columns. Store CS(cx[][],yh[i][][]) into the similarity array xsd[i]; H3. If i<n, execute step H2 and i=i+1, otherwise execute step H4; H4, step ends.
5. The method for searching for address book users according to claim 4, characterized in that: The sorting method is: Among them, TopK[z] represents an array for storing z user information, sort(xsd[],i) means sorting xsd[] and taking out the similarity value of the i-th subscript, and top(sort(xsd[],i)) means obtaining the user information corresponding to sort(xsd[],i).
6. A system for searching users in an address book, characterized in that: include: Query preprocessing module: performs word segmentation processing on the input query statement according to the preset character length to obtain a query word segmentation set; Query vector construction module: calculate the word frequency-inverse document frequency of each query word in the query word set and each user information in the address book to obtain a two-dimensional query vector; User vector construction module: Segment each user information in the address book to obtain several user segmentation sets, calculate the word frequency of each user segmentation in the user segmentation set and each user information in the address book - inverse document frequency to obtain a 3D user vector; Similarity array construction module: decompose the 3D user vector into several 2D user vectors, calculate the cosine similarity between the 2D query vector and several 2D user vectors to obtain several similarity values, and store the several similarity values in the similarity array; Address book search module: Sort the similarity values corresponding to each subscript of the similarity array from high to low, and obtain the user information corresponding to the first z similarity values as the address book search result according to the preset return number z.
7. The address book user search system according to claim 6, characterized in that: The query preprocessing module uses the following word segmentation method: Determine the preset character length and segment the query statement according to the preset length: PQ=P(Q); P(Q)={w1,w i ,...,w n }; Among them, PQ represents the query word set, P represents the word processing function, Q represents the query statement, and w i Represents the i-th query word, and n represents the preset maximum number.
8. The address book user search system according to claim 7, characterized in that: The query vector construction module calculates the query vector as follows: S1, let i = 1, initialize the query vector array cx[][]; S2, let j = 1, where i and j represent numeric variables; S3. Calculate w i And the term frequency-inverse document frequency of the j-th user information d: TF_IDF(w i ,d)=TF(w i ,d)*IDF(w i ); Among them, TF_IDF(w i ,d) represents the i-th query word w i The term frequency-inverse document frequency, TF(w i ,d) represents the i-th query word w i The word frequency of the jth user information d, IDF(w i ) represents the i-th query word w i and the inverse document frequency of the jth user information d, θ represents the i-th query word w i The number of times it appears in the j-th user information d, μ represents the total number of word segments in the j-th user information d, N represents the number of user information, and σ represents the number of words that contain the i-th query word w i The number of j-th user information d; The obtained TF_IDF(w i ,d) Store in cx[i][j]; S4. If j<n, execute step S3 and j=j+1, otherwise execute step S5; S5. If i<n, execute step S2 and i=i+1, otherwise execute step S6; S6. Use the query vector array cx[][] as the query vector; The user vector construction module, the calculation method of the user vector is: F1. Let j = 1, initialize the user vector array yh[][][]; F2. Perform word segmentation processing on the j-th user information according to a preset character length to obtain a user word segmentation set g of the j-th user information; F3, let k = 1, where k represents a numeric variable; F4, let i = 1; F5. Calculate the term frequency-inverse document frequency of the k-th user segmentation f of g and the i-th user information d: TF_IDF(f,d)=TF(f,d)*IDF(f); Wherein, TF_IDF(f,d) represents the term frequency-inverse document frequency of the kth user segmentation f of the user segmentation set g of the jth user information and the i-th user information d, TF(f,d) represents the term frequency of the kth user segmentation f of the user segmentation set g of the jth user information and the i-th user information d, IDF(f) represents the inverse document frequency of the kth user segmentation f of the user segmentation set g of the jth user information and the i-th user information d, ε represents the number of times the kth user segmentation f of the user segmentation set g of the jth user information appears in the i-th user information d, γ represents the total number of segmentations of the i-th user information d, N represents the number of user information, and σ represents the number of i-th user information d of the kth user segmentation f of the user segmentation set g containing the j-th user information; Store the obtained TF_IDF(f,d) into yh[j][k][i]; F6. If i<n, execute step F5 and i=i+1, otherwise execute step F7; F7. If k<n, execute step F4 and k=k+1, otherwise execute step F8; F8. If j<n, execute step F2 and j=j+1, otherwise execute step F9; F9. Use the user vector array yh[][][] as the user vector.
9. The address book user search system according to claim 8, characterized in that: The similarity array construction module, the similarity value is calculated as follows: H1, let i=1, initialize the similarity array xsd[]; H2. Calculate the cosine similarity CS(cx[][],yh[i][][]) between the 2D query vector cx[][] and the i-th 2D user vector yh[i][][]: Among them, · represents the dot product, ||cx[][]|| represents the modulus of cx[][], and cx[1][n*n] represents changing the cx[][] with n rows and n columns into a one-dimensional array cx[1][n*n] with 1 row and n*n columns. Store CS(cx[][],yh[i][][]) into the similarity array xsd[i]; H3. If i<n, execute step H2 and i=i+1, otherwise execute step H4; H4, step ends.
10. The address book user search system according to claim 9, characterized in that: The address book search module sorts by: Among them, TopK[z] represents an array for storing z user information, sort(xsd[],i) means sorting xsd[] and taking out the similarity value of the i-th subscript, and top(sort(xsd[],i)) means obtaining the user information corresponding to sort(xsd[],i).
Citation Information
Patent Citations
Place name matching method based on matrix operation
CN115687870A
Cited By
Smart phone book searching method based on TF-IDF pinyin vector model
CN120743976A