A user feature query method based on multi-dimensional calculation

By matching institutional users with scholar databases using multidimensional calculation methods, the problems of incomplete, non-standard, and outdated user information have been solved, resulting in richer and more real-time user information and improved accuracy in review and auditing.

CN115952218BActive Publication Date: 2026-01-13TONGFANG KNOWLEDGE DIGITAL PUBLISHING TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310023685.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-09
Publication Date
2026-01-13
Estimated Expiration
2043-01-09

AI Technical Summary

Technical Problem

Incomplete, non-standard, outdated, or insufficient user information registration makes it difficult for organizations to obtain accurate and detailed user information, affecting the review and audit process.

Method used

A user feature query method based on multidimensional computation is adopted, which matches institutional user information with scholar information in the scholar database through multiple dimensions, including name, institution, email, research direction, etc., and updates user information using precise and fuzzy matching algorithms.

Benefits of technology

Enrich, standardize, and update user information in real time to help organizations obtain more accurate and detailed user information, supporting a more efficient review and approval process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115952218B_ABST
    Figure CN115952218B_ABST
Patent Text Reader

Abstract

The application discloses a user feature query algorithm based on multi-dimensional calculation, which comprises the following steps: obtaining original information of an institutional unit user; calculating matching degrees of the user and scholars in a scholar database; and updating user information by using information of the scholar with the highest matching degree. The application can not only help the institutional unit enrich the information dimension of the user, but also correct, standardize and update the user information in real time, so that the institutional unit can obtain more abundant, accurate and standard user information in time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of information acquisition technology, and in particular to a user feature query method based on multidimensional computation. Background Technology

[0002] Institutions such as research institutes, companies, laboratories, newspapers, journals, universities, and book publishers often need to refer to user information to make judgments or formulate strategies during project review, research team building, content review, thesis defense, and professional title evaluation. For example, this includes selecting review experts during project review, considering the applicant's academic achievements during professional title evaluation, selecting experts / scholars during research team building, selecting authors' academic achievements and review experts during content review, and conducting regular assessments of journal editors, editors-in-chief, and editorial board members. Users register basic personal information when joining an institution, such as name, work unit / university, mailing address, email address, research direction, research field, and role, so that the institution can have a more comprehensive and in-depth understanding of the user. When authors submit manuscripts, the institution can understand the author's research field and direction based on their personal information; when expert review is required, the institution can select a more suitable expert based on the expert's personal information (research direction, research field, work unit). Furthermore, obtaining richer user information would be more helpful to institutions, such as understanding the author's or expert's publication information, including the number of publications, citation frequency, download frequency, G-index, and H-index, etc. However, basic user information presents some problems:

[0003] 1) Incomplete information registration; for example, a large number of users only left their name and email address, leaving the rest of the information blank.

[0004] 2) Information registration is not standardized, for example: the work unit is too simple or redundant;

[0005] 3) The information is outdated, for example: the work unit registered 5 years ago has long since changed;

[0006] 4) The information is not rich enough. For example, it is difficult for users to compile and register their academic achievements themselves. Summary of the Invention

[0007] To address the aforementioned technical problems, the purpose of this invention is to provide a user feature query method based on multidimensional computation.

[0008] The objective of this invention is achieved through the following technical solution:

[0009] A user feature query method based on multidimensional computation includes the following steps:

[0010] Step A: Obtain the original information of the institutional users;

[0011] Step B calculates the matching degree between users and scholars in the scholar database from multiple dimensions;

[0012] Step C: Update the user information by retrieving the information of the scholar with the highest matching degree.

[0013] Compared with the prior art, one or more embodiments of the present invention may have the following advantages:

[0014] It can not only help organizations enrich the information dimensions of users, but also correct, standardize and update user information in real time, so that organizations can obtain richer, more accurate and standardized user information in a timely manner. Attached Figure Description

[0015] Figure 1 This is a flowchart of a user feature query algorithm based on multidimensional computation.

[0016] Figure 2 This is a flowchart illustrating the multi-dimensional calculation of the matching degree between users and scholars in the scholar database;

[0017] Figures 3-6 This is a flowchart of the process for extracting institution names and matching scholar information;

[0018] Figure 7 , Figure 8 This is a flowchart of extracting information from a scholar database and matching scholar information. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in further detail below with reference to the embodiments and accompanying drawings.

[0020] like Figure 1 The diagram shows the process flow of a user feature query algorithm based on multidimensional computation, including:

[0021] 1. Obtaining Original Information of Institutional Users: This module mainly obtains information about institutional users, including: name, work unit / institution, mailing address, email address, research direction, research field, role, and the sponsoring unit of the institution. Roles include: expert role, author role (submitting manuscripts or applying for professional titles), and "editor" role (e.g., editor, chief editor, editorial board member, etc.). Often, a user has multiple roles.

[0022] 2. Multi-dimensional calculation of the matching degree between users and scholars in the scholar database: This module mainly calculates the matching degree between institutional users and scholars in the scholar database from multiple dimensions. The scholar database information mainly includes: scholar name, gender, scholar ID, standard primary institution, previous institution, previous institution, scholar's institution, original institution, research direction, research field, scholar's G index, scholar's H index, number of publications, average citation frequency, download frequency, etc. The specific calculation process for each dimension is as follows:

[0023] 2.1 Match scholars sequentially by extracting institution names (e.g., Figure 2 (As shown)

[0024] This module executes four sub-modules sequentially: matching scholars based on author / expert's institution / university, matching scholars based on author / expert's mailing address, matching scholars based on author / expert's email address, and matching scholars based on the author / expert's "editor" status and the journal's sponsoring institution. Figures 3-6 (As shown). The four sub-modules use the same matching method, as follows:

[0025] 1) Name of the extraction institution:

[0026] Based on the user role and the source of the organization information, all possible organization names are extracted sequentially from the four sub-modules. The organization names extracted from each sub-module are arranged in reverse order of sequence length. If the organization information source sequence is not among the extracted organizations, the organization information sequence is added as the first organization name.

[0027] 2) Extract scholar information from the scholar database

[0028] This module queries the scholar database by the name of the user from the institution / unit, and obtains information about scholars with the same name, including: scholar's name, gender, scholar ID, standard first-level institution, previous institution, previous institution of the scholar, scholar's institution, and original institution of the paper.

[0029] 3) Matching scholar information

[0030] The search retrieves the institution names extracted from 1) among the standard first-level institutions, previous institutions, previous institutions of scholars, institutions of scholars, and institutions of the original text. Matching methods include exact matching and fuzzy search; if exact matching yields no results, fuzzy matching is performed.

[0031] Exact match:

[0032] a. Completely match the extracted institution names from the standard first-level institution, the institution where the scholar previously worked, the institution where the scholar worked, the institution of the scholar, and the institution of the original text;

[0033] b. Returns all standard Level 1 organizations corresponding to the hit organization name;

[0034] c. Select the scholar code with the highest G-index from all standard Tier 1 institutions corresponding to the matched institution name, and use it as the scholarID of the matched scholar;

[0035] d. Return the scholar information corresponding to the scholarID.

[0036] Fuzzy search:

[0037] a. Calculate the shortest edit distance (ED) and the percentage of shortest edit distance (EDR) for all institution names in the institution name, standard first-level institution, former institution, former institution of the scholar, institution of the scholar, and institution of the original text;

[0038] The calculation method for a.1ED is as follows:

[0039] ED(i,j)=min{ED(i-1,j)+1,ED(i,j-1)+1,ED(i-1,j-1)+c(i,j)}

[0040] in:

[0041]

[0042] Where: S is the original sequence, and T is the target sequence.

[0043] The calculation method for a.2EDR is as follows:

[0044] S and T are the two sequences used to compute ED;

[0045] b. Sort ED and EDR in ascending order, and take the name of the institution with the smallest ED and EDR;

[0046] c. If EDR < threshold (preset 0.3), then the matching of the extracted organization name fails, and the matching continues to the next extracted organization name;

[0047] d. Returns all standard Level 1 organizations corresponding to the hit organization name;

[0048] e. Select the scholar code with the highest scholar G index from all standard first-level institutions corresponding to the hit institution name, and use it as the scholar ID of the hit scholar;

[0049] f. Returns the scholar information corresponding to the scholarID.

[0050] 2.2 Matching scholars based on the author's / expert's research direction and field

[0051] like Figure 7 As shown, 1) Extract scholar information from the scholar database.

[0052] This module searches the scholar database by institutional user name to obtain information about scholars with the same name, including: scholar's name, gender, scholar ID, research direction, and research field;

[0053] 2) Matching scholar information

[0054] Based on the input author / expert's research direction and research field, match the research direction and research field of scholars returned by the scholar database, and select the scholar with the closest similarity as the matched scholar.

[0055] Matching method:

[0056] a. Take the set of single words S1 that constitutes the author's / expert's research direction and research field, and remove stop words;

[0057] b. Take the set of single words S2 of the research directions and research fields of scholars returned by the scholar database, and remove the stop words;

[0058] c. Calculate the proportion of single-character words in S1 to words in S2:

[0059] d. Take the scholar ID with the highest percentage as the scholar ID of the matched scholar;

[0060] e. Returns the posting information corresponding to the scholarID.

[0061] 2.3 Matching scholars based on author / expert name and scholar index (G index)

[0062] like Figure 8 As shown, 1) Extract scholar information from the scholar database.

[0063] By searching the scholar database by institutional user name, information about scholars with the same name can be obtained, including: scholar name, gender, scholar ID, and scholar G-index.

[0064] 2) Matching scholar information

[0065] Based on the authors' G-index in the returned results, scholars are matched, and the scholar with the highest G-index is taken as the scholarID of the matched scholar. The publication information corresponding to the scholarID is then returned.

[0066] 3. Update and enrich user information by retrieving the scholar information with the highest matching degree: This module mainly uses the scholar information with the highest matching degree to supplement the information dimensions of institutional users.

[0067] The specific implementation method is as follows:

[0068] Task 1. An organization / unit needs to update and enrich the information of user A.

[0069] The processing flow is as follows:

[0070] 1. Obtain the original information of the institutional user: The user's information is as follows: "Name": "Huang Xihua", "Gender": "Male", "Work Unit / Institution": "School of Materials Science and Engineering, XXX University", "Research Field": "", "Research Direction": "Preparation and performance research of functionalized polymer nanomaterials: preparation and performance research of semiconductor polymer photocatalyst materials; preparation of carbon materials based on semiconductor polymers and research on their electrical properties; design and synthesis of biofunctional aliphatic branched polymers and their applications in drug sustained release, etc.", "Correspondence Address": "", "Role": "Expert, Author", "Organizing Institution": "Chinese XXX Society; XX Technology Group Co., Ltd.", "Email Address": "***@aust.edu.cn".

[0071] 2. Calculate the matching degree between users and scholars in the scholar database from multiple dimensions:

[0072] 2.1 Matching scholars by extracting institution names in sequence

[0073] 1) Extraction of organization name

[0074] Based on the user's expert / author role, this list is arranged in order from "work unit / institution" to "mailing address".

[0075] "and "email address" extracted institution name orgs = ["XX University of Technology, School of Materials Science and Engineering"]

[0076] [School of Materials Science and Engineering, XX University of Technology]

[0077] 2) Extract scholar information from the scholar database

[0078] This finds information on 48 scholars with the same name, constructing an institutional dictionary orgs_dic = {"Institution 1":{"Scholar ID":[Scholar ID 1, Scholar ID 2],"otherinfo":{"Standard First-Level Unit":"","

[0079] Publication count: "5",......}},"XX University of Science and Technology, School of Materials Science and Engineering":{"Scholar ID"

[0080] ":["000037366****"],"otherinfo":{"Standard Level 1 Unit":"XX University of Science and Technology","Number of Publications":"15",......},......}

[0081] 3) Matching scholar information

[0082] Here, the scholar with scholar ID "000037366****" was precisely found, which is the scholar with the highest match degree with user A.

[0083] 3. Update and enrich user information based on the scholar information with the highest matching degree: Update and enrich user information based on the scholar information with scholar number "000037366****".

[0084] Task 2. An organization / unit needs to update and enrich the information of user B.

[0085] The processing flow is as follows:

[0086] 1. Obtain the original information of the institutional user: The information obtained for this user is as follows: "Name":"An X","Gender":"","Work unit / institution":"","Research field":"","Research direction":"","Mailing address":"","Role":"Expert","Organizing unit of the institution":"XX University","Email address":"****@sjtu.edu.cn".

[0087] 2. Calculate the matching degree between users and scholars in the scholar database from multiple dimensions:

[0088] 2.1 Matching scholars by extracting institution names in sequence

[0089] 1) Extraction of organization name

[0090] Based on the user's expert / author role, this list is arranged in order from "work unit / institution" to "mailing address".

[0091] The institution name extracted from the "email address" is orgs = ["XX Jiaotong University"].

[0092] 2) Extract scholar information from the scholar database

[0093] Here we found information on 7 scholars with the same name, and constructed an institutional dictionary orgs_dic = {"Institution 1":{"Scholar ID":[Scholar ID 1, Scholar ID 2],"otherinfo":{"Standard Level 1 Unit":"","

[0094] Publication count: "3",......}},"School of Materials Science and Engineering, XX University of Science and Technology":{"Scholar ID"}

[0095] ":["000036790****"],"otherinfo":{"Standard First-Level Unit":"XX Jiaotong University","Number of Documents Published":"99",......},......}

[0096] 3) Matching scholar information

[0097] Here, the scholar with scholar ID "000036790****" was precisely found, which is the scholar with the highest match degree with user B.

[0098] 3. Update and enrich user information based on the scholar information with the highest matching degree: Update and enrich user information based on the scholar information with scholar number "000036790****".

[0099] Task 3. An organization / unit needs to update and enrich the information of user C.

[0100] The processing flow is as follows:

[0101] 1. Obtain the original information of the institutional user: The information obtained for this user is as follows: "Name":"Cao XX","Gender":"Female","Work Unit / Institution":"","Research Field":"","Research Direction":"Application Research of Modern Detection Technology; Application Research of Control Theory","Mailing Address":"No. 5, Dongfeng Road, XX","Role":"Author, Expert","Organizing Institution":"XX Coal Society","Email Address":"****@zzuli.edu.cn".

[0102] 2. Calculate the matching degree between users and scholars in the scholar database from multiple dimensions:

[0103] 2.1 Matching scholars by extracting institution names in sequence

[0104] 1) Extraction of organization name

[0105] Based on the user's expert / author role, the organization name orgs = [] is extracted sequentially from "work unit / institution", "mailing address" and "email address". If no organization name is extracted, the search for this module ends and proceeds to the next module.

[0106] 2.2 Matching scholars based on the author's / expert's research direction and field

[0107] 1) Extract scholar information from the scholar database

[0108] This finds information on 6 scholars with the same name, and constructs a dictionary of research directions and fields: RDF_dic = {"Research Direction and Field": {"Scholar ID": [Scholar ID 1, Scholar ID 2], ...}

[0109] "otherinfo":{"Standard First-Level Unit":"","Number of Documents Issued":"6",......}},"Research on Testing Technology and Virtual Instruments, Electric Power Industry; Automation Technology; Radio Electronics":{"Scholar Number"}

[0110] ":["000006361****"],"otherinfo":{"Standard primary unit":"XX University of Light Industry","

[0111] Publication volume":"133",......},......}

[0112] 2) Match scholar information

[0113] Match the research directions and fields of the scholars returned by the scholar database according to the research directions and fields of the input author / expert, and select the scholar with the closest match as the hit scholar.

[0114] Matching method:

[0115] a. Take the set S1 of single Chinese characters that make up the research directions and fields of the author / expert, and remove the stop single Chinese characters. S1 = ['generation','manufacture','response','technology','control','technique','detection','current','theory','application','

[0116] research','study','discussion'];

[0117] b. Take the set S2 of single Chinese characters in the research directions and fields of the scholars returned by the scholar database, and remove the single Chinese characters. S2 = ['industry','instrument','operation','force','motion','automation','and','instrument','electronics','science','engineering','

[0118] technology','simulation','none','technique','detection','electricity','research','line','self','virtual'];

[0119] c. Calculate the proportion of the single Chinese characters in S1 in S2: After calculation, the highest proportion is 0.26, which is the research direction and field corresponding to the scholar with the ID 000006361****.

[0120] d. Return the publication information corresponding to this scholarID = "000006361****".

[0121] 3. Update and enrich the user information with the information of the scholar with the highest matching degree: Update and enrich the user information according to the information of the scholar with the ID 000006361****. [[ID=3⑥]]

[0122] While the embodiments disclosed in this invention are as described above, the content is merely for the purpose of facilitating understanding of the invention and is not intended to limit the invention. Any person skilled in the art to which this invention pertains may make any modifications and variations in form and detail of the implementation without departing from the spirit and scope disclosed herein; however, the scope of patent protection for this invention shall still be determined by the scope defined in the appended claims.

Claims

1. A method for user feature query based on multi-dimensional computation, characterized in that, It comprises the following steps: Step A: obtaining original information of the institutional user; Step B: calculating the matching degree between the user and the scholars in the scholar database in multiple dimensions; Step C: updating the user information with the information of the scholar with the highest matching degree; The scholar information matching comprises precise matching and fuzzy matching; wherein the precise matching comprises: completely matching the extracted institution name from the standard first-level institutions, the former unit, the scholar's former unit, the scholar's unit and the original text institution; returning all the standard first-level institutions corresponding to the hit institution name; selecting the scholar code with the highest G index from all the standard first-level institutions corresponding to the hit institution name as the scholar ID of the matched scholar; returning the scholar information corresponding to the scholar ID; The fuzzy matching comprises: calculating the shortest edit distance ED and the shortest edit distance ratio EDR of the institution name and all the institution names in the standard first-level institutions, the former unit, the scholar's former unit, the scholar's unit and the original text institution; sorting the ED and EDR in ascending order, and taking the hit institution name with the smallest ED and EDR; if EDR < threshold value 0.3, the extracted institution name fails to match, and the next extracted institution name is matched; returning all the standard first-level institutions corresponding to the hit institution name; selecting the scholar code with the highest G index from all the standard first-level institutions corresponding to the hit institution name as the scholar ID of the matched scholar; returning the scholar information corresponding to the scholar ID; The calculation formula of the shortest edit distance ED is: ED(i,j) = min{ED(i-1,j)+1, ED(i,j-1)+1, ED(i-1,j-1)+c(i,j)} wherein: wherein: S is the original sequence, and T is the target sequence; The calculation formula of the edit distance ratio EDR is: S, T are the original sequence and target sequence for calculating ED; In step C, the standard first-level institutions, research directions and scholar indexes of the matched scholars are used to dynamically cover the original user information.

2. The multi-dimensional computation based user characteristic query method of claim 1, wherein, Step B is to calculate the matching degree between the institutional user and the scholars in the scholar database in multiple dimensions, wherein the calculation process of each dimension comprises: B1: matching the scholars by extracting the institution name one by one; B2: matching the scholars according to the author / expert research direction and research field; B3: matching the scholars according to the author / expert name and scholar index.

3. The method for user profiling based on multidimensional computation as claimed in claim 2 wherein, The B1 comprises: extracting the institution name from the sub-module according to the user role and the source of the institution information; querying the scholar database by the name of the institutional user to obtain the information of the scholars with the same name; extracting the institution name and matching the scholar information.

4. The multi-dimensional computation based user feature query method of claim 2, wherein, The B2 specifically comprises: B2.1: extracting the scholar information by querying the scholar database by the name of the institutional user; B2.2: matching the research direction and research field of the scholars returned by the scholar database according to the input author / expert research direction and research field, and taking the nearest scholar as the matched scholar.

5. The method for user profiling based on multidimensional computation as claimed in claim 4 wherein, The matching method of B2.2 specifically comprises: taking the single-word set S1 constituting the author / expert research direction and research field, and removing the stop single-word; taking the single-word set S2 of the research direction and research field of the scholars returned by the scholar database, and removing the stop single-word; Calculate the proportion of single words in S1 in S2: taking the scholar number with the highest proportion as the scholar ID of the matched scholar; Return the corresponding publication information of the scholarID.

6. The multi-dimensional computation based user feature query method of claim 2, wherein, The B3 specifically includes: Query the scholar database through the name of the user of the institutional unit, and extract scholar information from the scholar database; Match the scholar from the G index, take the scholar with the highest G index as the scholarID of the matched scholar, and return the corresponding publication information of the scholarID.

Citation Information

Patent Citations

  • Author disambiguation method and device based on subject tree clustering

    CN111221968A