User matching methods, devices, electronic equipment, and computer-readable media

By comprehensively considering data such as user name, code, location, and contact information during the user matching process, the success rate of user matching is improved, solving the problem of low success rate in existing technologies that rely solely on user name for matching, and achieving more accurate user matching.

CN117195146BActive Publication Date: 2026-04-03中国移动通信集团云南有限公司 +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-29
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

In existing technologies, user matching mainly relies on comparing user names, failing to effectively utilize data other than user names, resulting in a low matching success rate.

Method used

By inputting at least two of the first user's username, user code, user location, and user contact information into the matching model, the output indicates the second user data that matches the first user among the archived users, considering matching multiple different types of user data in different dimensions.

Benefits of technology

It improves the success rate of user matching by accurately determining user matching relationships through multi-dimensional data analysis, reducing false matches and missed matches.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117195146B_ABST
    Figure CN117195146B_ABST
Patent Text Reader

Abstract

This application provides a user matching method, apparatus, electronic device, and computer-readable medium, relating to the field of big data technology. The method includes: inputting first user data of a first user into a matching model, the first user data including at least two of a first user name, a first user code, a first user location, and first user contact information; and outputting second user data through the matching model, the second user data indicating a second user among archived users who matches the first user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of big data technology, and in particular to a method, apparatus, electronic device, and computer-readable medium for user matching. Background Technology

[0002] When managing users within an organization, it's common practice to determine if a target user is already an archived user to avoid wasting resources due to duplicate archiving. Currently, this is primarily based on usernames to determine the match between the target user and archived users. For example, this can be done by directly comparing the target user's name with the names of archived users, or by using data mining algorithms to compare the target user's name with the names of archived users.

[0003] However, the above method does not take into account data other than usernames, resulting in a low success rate for user matching. Summary of the Invention

[0004] The purpose of this application is to provide a user matching method, apparatus, electronic device, and computer-readable medium that can improve the success rate of user matching.

[0005] To solve the above-mentioned technical problems, the embodiments of this application are implemented through the following aspects.

[0006] In a first aspect, embodiments of this application provide a user matching method, comprising: inputting first user data of a first user into a matching model, the first user data including at least two of a first user name, a first user code, a first user location, and a first user contact information; and outputting second user data through the matching model, the second user data being used to indicate a second user among archived users who matches the first user.

[0007] Secondly, embodiments of this application provide a user matching apparatus, comprising: an input module for inputting first user data of a first user into a matching model, the first user data including at least two of a first user name, a first user code, a first user location, and a first user contact information; and an output module for outputting second user data through the matching model, the second user data indicating a second user among archived users who matches the first user.

[0008] Thirdly, embodiments of this application provide an electronic device, including: a memory, a processor, and computer-executable instructions stored in the memory and executable on the processor, wherein the computer-executable instructions, when executed by the processor, implement the user matching method described in the first aspect above.

[0009] Fourthly, embodiments of this application provide a computer-readable storage medium for storing computer-executable instructions, which, when executed by a processor, implement the user matching method described in the first aspect above.

[0010] In this embodiment, by inputting the first user's first user data into the matching model, the first user data includes at least two of the first user name, first user code, first user location, and first user contact information; through the matching model, the second user data is output, which is used to indicate the second user among the archived users who match the first user. This approach can consider multiple different types of user data and perform user matching in different dimensions, thereby improving the success rate of user matching. Attached Figure Description

[0011] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This illustration shows a flowchart of a user matching method provided in an embodiment of this application;

[0013] Figure 2 This illustration shows another flowchart of a user matching method provided in an embodiment of this application;

[0014] Figure 3 This illustration shows a schematic diagram of a user matching method provided in an embodiment of this application;

[0015] Figure 4 This illustration shows a structural schematic diagram of a user matching device provided in an embodiment of this application;

[0016] Figure 5 A schematic diagram of the hardware structure of an electronic device for implementing a user matching method provided in an embodiment of this application. Detailed Implementation

[0017] To enable those skilled in the art to better understand the technical solutions in this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this application.

[0018] Figure 1 This diagram illustrates a user matching method provided in an embodiment of this application. This method can be executed by an electronic device, such as a terminal device or a server device. In other words, the method can be executed by software or hardware installed on the terminal device or server device. The server includes, but is not limited to, a single server, a server cluster, a cloud server, or a cloud server cluster. As shown, the method may include the following steps.

[0019] Step S110: Input the first user's first user data into the matching model.

[0020] The first user is the user of the pre-defined application or software system, which can be an individual user or an organizational user, such as a corporate user, company user, or group user. The first user data is data associated with the first user, used to indicate or characterize the first user, including at least two of the following: first user name, first user code, first user location, and first user contact information. The matching model is used to perform user matching based on the input first user data.

[0021] Optionally, the first user is a user who is using the application or software system for the first time. The matching model is used to match the first user with archived users based on at least two of the user name, user code, user location and user contact information to determine whether the first user belongs to an archived user or whether the first user can be associated with an archived user.

[0022] Optionally, the first user data is data fused from different data sources. These different sources include, for example, data obtained from a data model, data obtained from a webpage, and data provided by an application. Data fusion includes cleaning the data from different sources to remove duplicate, redundant, dirty, and corrupted data. Data fusion also includes assigning a unique identifier to data belonging to the same user. For example, assigning a first ID to the first user data. Data fusion further includes classifying the data from different sources.

[0023] This step provides the first set of user data, which includes data from different dimensions, and can provide a more comprehensive data foundation for user matching.

[0024] Step S120: Output the second user data.

[0025] The matching model outputs second user data, which indicates a second user among the archived users that matches the first user. In other words, there is a correspondence between the second user data and the second user, and the corresponding second user can be determined based on the second user data. The number of second users can be one or more, and there is no limitation here.

[0026] Since the first user data includes at least two of the following: first user name, first user ID, first user location, and first user contact information, the second user determined based on the first user data matches the first user in at least two of the following aspects: user name, user ID, user location, and user contact information. Matching is achieved, for example, by the two users being completely identical, or by the similarity between them exceeding a specified threshold.

[0027] This step allows for accurate user matching based on user data from multiple dimensions.

[0028] In this embodiment, by inputting the first user's first user data into the matching model, the first user data includes at least two of the first user name, first user code, first user location, and first user contact information; through the matching model, the second user data is output, which is used to indicate the second user among the archived users who match the first user. This approach can consider multiple different types of user data and perform user matching in different dimensions, thereby improving the success rate of user matching.

[0029] In one possible implementation, outputting the second user data through the matching model includes: outputting the second user data based on the probability value of each of the archived users matching the first user through the matching model.

[0030] The matching model determines probability values. These probability values ​​represent the degree of match between the archived user and the first user. A higher probability value indicates a stronger match between the archived user and the first user. Optionally, the second user is the archived user with the highest probability value; in other words, the second user is the archived user with the highest degree of match with the first user.

[0031] The second user, determined by probability value, has a high degree of matching with the first user, which can further improve the success rate of user matching.

[0032] In one possible implementation, the user location includes a home location and location coordinates. Before outputting the second user data, the method further includes: determining that the probability of the second user matching the first user is 0 if at least one of the following conditions is met: if the second home location of the second user is known, the second home location of the second user is different from the first home location of the first user; if the second home location is known, the similarity between the second user's second username and the first user's username is lower than a first threshold; if the second home location is unknown, the distance between the second location coordinates of the second user and the first location coordinates of the first user is higher than a second threshold; if the second home location is unknown, the second location coordinates do not meet a predetermined location condition, and the similarity between the second user's username and the first user's username is lower than a third threshold.

[0033] The location field indicates the region to which the user belongs, such as the city, district, county, or grid to which the user belongs. The location coordinates are the user's coordinates in a coordinate system such as the world coordinate system (WGS84), including longitude and latitude.

[0034] Combination Figure 2 Based on the archived user data, step S201 determines whether the location of the second user is known. For users whose location is known, during user matching, users whose location is inconsistent with that of the first user are removed. In other words, it is determined that the user does not match the first user, and the probability of the user matching the first user is 0. For example, if it is known that the location of an archived user is City B, and the location of the first user is City A, then the archived user is removed. For users whose location is known, it is also necessary to remove users whose usernames have a similarity to the first user's name that is lower than a first threshold, such as users whose usernames have a similarity of 0 to the first user's name.

[0035] For users whose location is unknown, step S202 further determines whether the user's location coordinates are valid, i.e., whether the user's location coordinates meet predetermined location conditions. For users with valid location coordinates, users whose location coordinates are more than a specified threshold away from the first location coordinates need to be removed. For example, if the distance between the location coordinates of an archived user and the first location coordinates is greater than the threshold of 5km, in other words, the location coordinates of the archived user are outside the 5km range of the first location coordinates, then the archived user is removed during user matching. For users with valid location coordinates, users whose usernames have a similarity to the first username that is less than a first threshold also need to be removed, for example, users whose usernames have a similarity of 0 to the first username. In addition, for users with invalid user coordinates, users whose usernames have a similarity to the first username that is less than a third threshold need to be removed, for example, the third threshold is 20%.

[0036] Optionally, calculating the similarity between usernames includes removing phrases from usernames that are meaningless for matching, such as "limited," "responsibility," "company," or place names containing "province," "city," or "district"; defining the edit distance as the minimum number of insertions, deletions, and replacements required to transform the original string (s) into the target string (t); and calculating the similarity according to the following formula (1). The original string (s) is, for example, the first username, and the target string (t) is, for example, the second username.

[0037]

[0038] The steps in this implementation can be performed before or after the probability value is determined. Taking the step of performing the step before determining the probability value as an example, this step performs a preliminary screening of user data based on location, coordinate values, and username similarity, which can eliminate mismatched user data in advance and improve the efficiency of data processing.

[0039] In one possible implementation, the second position coordinates not satisfying a predetermined position condition includes: the second position coordinates being located outside the intersection area of ​​multiple target circle center areas, wherein the target circle center area is an area determined based on the third position coordinates of each core member of the second user and a predetermined radius, and the third position coordinates satisfying a first predetermined condition.

[0040] Determining whether the second position coordinates meet the predetermined position conditions includes determining the optimal radius r, and determining whether the second position coordinates are valid based on the optimal radius r.

[0041] The process of determining the optimal radius r includes, for example, selecting N group users with valid latitude and longitude coordinates and M group users with invalid latitude and longitude coordinates. The information for each group user includes the group's latitude and longitude coordinates and the individual latitude and longitude coordinates of multiple group members. For each group user, a circle is drawn with the latitude and longitude coordinates of each group member as the center and a preset neighborhood radius eps (e.g., 500m) as the radius. It is determined whether the group members falling within this circle exceed a preset percentage (e.g., 75%) of the total number of group members. If they do, the group member is identified as a core member, and their third position coordinates satisfy a first predetermined condition. Then, for each group, the optimal radius r is initialized to a minimum value (e.g., 100m). A circle is drawn with the latitude and longitude coordinates of each core member as the center and the optimal radius r as the radius, obtaining the target center area corresponding to each core member. The intersection area (i.e., the common area) of multiple target center areas is determined, and it is determined whether the group's latitude and longitude coordinates are located within the intersection area. Calculate the percentage, p, of N valid latitude and longitude groups falling into their corresponding intersecting regions; and calculate the percentage, q, of M invalid latitude and longitude groups falling into their corresponding intersecting regions. If p is not greater than a first preset ratio (e.g., 90%) and q is not less than a second preset ratio (e.g., 5%), then increment the optimal radius r according to a preset gradient, and repeat the above operation until p and q meet their respective requirements; finally, fix the determined r as the optimal radius r. If p or q cannot simultaneously meet the requirements, then fix the r that first meets the requirement of p as the optimal radius r, or fix the r that first meets the requirement of q as the optimal radius r.

[0042] The process of determining the validity of the second position coordinates based on the optimal radius r includes, for example, identifying the core members of the group user to be judged, drawing circles with the latitude and longitude of each core member as the center and the predetermined optimal radius r as the radius, obtaining the center area corresponding to each core member, determining the intersection area (i.e., the common area) of multiple center areas, and judging whether the group's latitude and longitude are located within the intersection area; if the group's latitude and longitude are located within the intersection area, the group's latitude and longitude are determined to be valid; otherwise, the group's latitude and longitude are determined to be invalid.

[0043] Table 1

[0044]

[0045] As shown in Table 1 above, the dense clustering algorithm (DBscan) includes: selecting points that meet a threshold within a given radius as core points; and including all points reachable by the core point density to form density-connectable clusters. Since most group members' daytime activities are concentrated near the group, and most group members tend to cluster within a certain area, density clustering can delineate the activity area with the highest density, thus determining whether the group's latitude and longitude are also included within that area. Adjusting the density clustering parameters mainly involves changing the size of the delineated area to select a reasonable range for determining the validity of the group's latitude and longitude. For model training, judging the validity of latitude and longitude serves as an important indicator for model prediction, improving the accuracy of latitude and longitude, and acting as an important criterion in the selection of positive samples. The determination of the validity of group latitude and longitude is mainly based on the latitude and longitude of group members, judging whether the group's latitude and longitude are valid and whether they can be used in subsequent models.

[0046] In this embodiment, by inputting the first user's first user data into the matching model, the first user data includes at least two of the first user name, first user code, first user location, and first user contact information; through the matching model, the second user data is output, which is used to indicate the second user among the archived users who match the first user. This approach can consider multiple different types of user data and perform user matching in different dimensions, thereby improving the success rate of user matching.

[0047] Figure 3 The diagram illustrates another flowchart of a user matching method provided in this application, including:

[0048] Step S301: Determine the user data sample whose number of the second user is the first value as the positive user data sample.

[0049] The positive user data sample is a data sample in which the first user is highly likely to match the second user, where the first user can be directly associated with the archived user data sample. For example, the first user can be directly associated with a unique second user if at least one of the following conditions is met: the number of second users whose first user code matches the second user code is 1; the number of second users whose first user code matches the first user code and whose first and second location matches is 1; the number of second users whose first and second location match and whose first and second user names have a 100% similarity is 1; the number of second users whose first and second user contact information matches is 1; and the number of second users whose first and second user contact information match and whose first and second location match is 1. Optionally, the first value can also be set to a value other than 1.

[0050] Optionally, the user code may be, for example, the Unified Social Credit Code. User contact information may include contact person's name, phone number, broadband information, etc. Refer to Table 2 below to determine if the user's location is consistent.

[0051] Table 2

[0052]

[0053] Step S302: The user data sample whose number of the second user is the second value is determined as the negative user data sample.

[0054] The second value is different from the first value; for example, the second value is 0, or the second value is a positive integer greater than 1.

[0055] From the data sample that the first user can directly associate with the archived user data, remove the above positive user data samples to obtain negative user data samples.

[0056] In one possible implementation, the user data of the negative user data sample also needs to satisfy at least one of the following: the first user code is the same as the second user code of the second user, and the first location is different from the second location; the similarity between the first user name and the second user name is 100%, and the first location is different from the second location; the first user contact information is the same as the second user contact information of the second user, and the similarity between the first user name and the second user name is lower than the first threshold; the first location is the same as the second location, and the similarity between the first user name and the second user name is lower than the first threshold; the first location is the same as the second location, and the similarity between the first user name and the second user name is a third value, wherein the third value satisfies a second predetermined condition.

[0057] For example, negative user data samples may meet at least one of the following conditions: The Unified Social Credit Code of the first user is the same as that of the second user; the district to which the first user belongs is different from that of the second user; and the number of identified second users is greater than one. Alternatively, the district to which the first user belongs is the same as that of the second user; the similarity between the first user's name and the second user's name is greater than 50%; and the number of second users is greater than one. In this case, the second user with the highest name similarity is removed, and the remaining second user data is used as the negative user sample data. Another possibility is that the similarity between the first user's name and the second user's name is 100%, and the district to which the first user belongs is different from that of the second user. Yet another possibility is that the district to which the first user belongs is different from that of the second user; and the similarity between the first user's name and the second user's name is 0. In this case, user data meeting the above conditions is randomly selected, accounting for one-fifth of the total negative sample. Finally, the contact phone number of the first user is the same as that of the second user, and the similarity between the first user's name and the second user's name is 0.

[0058] Step S303: Train the matching model using the positive user data samples and the negative user data samples.

[0059] The positive and negative user data samples identified above provide the data foundation for model training. The matching model trained based on these samples can be used to determine the sample to be judged. Optionally, the matching model is an XGBoost model. Inputting the sample to be judged into the XGBoost model trained with the positive and negative user data samples determines the probability value of a match between the first user and the second user, and selects the second user with the highest probability value as the archived user corresponding to the first user. It is understandable that this matching model may not find a second user matching the first user; in this case, a prompt such as "no matching user" can be output.

[0060] Step S310: Input the first user's first user data into the matching model.

[0061] Step S320: Output the second user data.

[0062] Steps S310 and S320 can be described using the corresponding steps in the previous embodiment. For repeatable parts, they will not be described again here.

[0063] In one possible implementation, the XGBoost model, due to its powerful parallel computing efficiency, missing value handling, and prediction performance, can be used for algorithm fitting and parameter tuning. As the data is updated, the implementation fits a residual tree in each iteration, improving data matching accuracy through machine learning. Furthermore, the output results can be evaluated. This includes using the Receiver Operating Characteristic (ROC) curve to represent different thresholds, with the area under the curve representing the model's "accuracy" and "error rate." A larger area under the curve indicates higher accuracy and better model performance. Alternatively, the Discriminant metric (KS) can be used to measure the maximum difference between the distributions of positive and negative user data samples under the model, further evaluating the model's output results.

[0064] In one possible implementation, the second user data includes the first user data and the matching relationship between the first user and the second user.

[0065] Matching relationships can be used to indicate whether a second user matches the first user. If a matching second user exists, the first and second users are associated to complete their base data.

[0066] The output of the second user data may also include a series of data determined based on the user's name, location, code, and contact information. Specifically, this includes: the user's industry attributes, such as shopping services, catering services, lifestyle services, automotive services, healthcare services, and accommodation services; the user's market attributes, such as buildings, street-front shops, hotels, industrial parks, and industrial markets; the user's location attributes, such as latitude and longitude, and the specific location of the store down to the city, district, grid, and administrative village levels; and basic user attributes, such as key personnel, company name, company ID, business district, status, broadband usage, and IT needs.

[0067] In this embodiment, by integrating data from different data sources, user matching is performed based on the XGBoost model, outputting user information including matching relationships between users, making the dimensions of user matching more complete. In addition to user names, user matching is also performed based on user coordinates, user contact information, user location, etc., making multi-dimensional user matching based on user data more accurate and reasonable, and better meeting actual business needs.

[0068] In one possible implementation, the user matching method also includes constructing a unified user identifier unit, constructing a similarity matching model unit, and constructing a model output unit.

[0069] The Unified User Identification Unit (UUID) is used to determine whether users obtained from multiple channels belong to the same user by performing word segmentation and semantic recognition on the Unified Social Credit Code, key person information, store name, and location. For data belonging to the same user, a unique user ID is assigned.

[0070] The similarity matching model unit compares users with unified IDs to archived users, first matching users who can be directly associated. Direct association includes the following three scenarios: a user's unified social credit code is associated with an archived user; a user's contact information is associated with an archived user's members or key persons; and a user's broadband application information is associated with an archived user's service application information. For users who cannot be directly associated, their associated users are determined based on the model's output. Before use, the model is trained using positive and negative samples. Samples that can be directly associated with and uniquely match an archived user are considered positive samples, while samples that can be associated with some non-unique archived users or cannot match any archived user are considered negative samples.

[0071] The model output unit is used for the initial screening of the range of users to be matched and the output of model results. The range of users to be matched after the initial screening is used as the target sample input into the similarity matching model, and the output is information on the group affiliation, industry attributes, and market affiliation of SMEs.

[0072] By combining, judging, and outputting data, this method can determine the similarity of user data from multiple dimensions. Compared to traditional methods of directly matching usernames or using mining algorithms (such as word2vec) to associate usernames, it has a higher matching rate and can avoid situations where some users are actually the same but cannot be matched due to name differences. This makes the user similarity judgment more reasonable and accurate.

[0073] Figure 4 The diagram shows a user matching device according to an embodiment of this application. The device 400 includes an input module 401 and an output module 402.

[0074] The input module 401 is used to input the first user data of the first user into the matching model. The first user data includes at least two of the first user name, first user code, first user location, and first user contact information. The output module 402 is used to output the second user data through the matching model. The second user data is used to indicate the second user in the archived users who matches the first user.

[0075] In one possible implementation, the output module 402 is used to output the second user data based on the probability value of each of the archived users matching the first user, using the matching model.

[0076] In one possible implementation, the user matching device further includes a determining module, configured to determine, before outputting the second user data, that the probability of the second user matching the first user is 0 if at least one of the following conditions is met: if the second location of the second user is known, the second location of the second user is different from the first location of the first user; if the second location is known, the similarity between the second user's second username and the first user's username is lower than a first threshold; if the second location is unknown, the distance between the second location coordinates of the second user and the first location coordinates of the first user is higher than a second threshold; if the second location is unknown, the second location coordinates do not meet a predetermined location condition, and the similarity between the second user's name and the first user's name is lower than a third threshold.

[0077] In one possible implementation, the second position coordinates not satisfying a predetermined position condition includes: the second position coordinates being located outside the intersection area of ​​multiple target circle center areas, wherein the target circle center area is an area determined based on the third position coordinates of each core member of the second user and a predetermined radius, and the third position coordinates satisfying a first predetermined condition.

[0078] In one possible implementation, the user matching apparatus further includes a model training module, used to determine user data samples in which the number of the second users is a first value as positive user data samples before outputting the second user data through the matching model; to determine user data samples in which the number of the second users is a second value as negative user data samples, the second value being different from the first value; and to train the matching model using the positive user data samples and the negative user data samples.

[0079] In one possible implementation, the user data of the negative user data sample satisfies at least one of the following: the first user code is the same as the second user code of the second user, and the first location is different from the second location; the similarity between the first user name and the second user name is 100%, and the first location is different from the second location; the first user contact information is the same as the second user contact information of the second user, and the similarity between the first user name and the second user name is lower than the first threshold; the first location is the same as the second location, and the similarity between the first user name and the second user name is lower than the first threshold; the first location is the same as the second location, and the similarity between the first user name and the second user name is a third value, wherein the third value satisfies a second predetermined condition.

[0080] In one possible implementation, the second user data includes the first user data and the matching relationship between the first user and the second user.

[0081] The device 400 provided in this application embodiment can execute the methods described in the preceding method embodiments and achieve the functions and beneficial effects of the methods described in the preceding method embodiments, which will not be repeated here.

[0082] Figure 5 This diagram illustrates the hardware structure of an electronic device executing the user matching method provided in the embodiments of this application. Referring to the diagram, at the hardware level, the electronic device includes a processor 510, and optionally includes an internal bus 520, a network interface 530, and a memory. The memory may include main memory 540, such as high-speed random-access memory (RAM), and may also include non-volatile memory 550, such as at least one disk storage device. Of course, the electronic device may also include other hardware required for other services.

[0083] The processor 510, network interface 530, and memory can be interconnected via an internal bus 520. This internal bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be categorized as an address bus, data bus, control bus, etc. For ease of illustration, only a single bidirectional arrow is used in this diagram, but this does not imply that there is only one bus or one type of bus.

[0084] The memory is used to store programs. Specifically, the program may include program code, which includes computer operation instructions. The memory may include main memory 540 and non-volatile memory 550, and provides instructions and data to the processor 510.

[0085] Processor 510 reads the corresponding computer program from non-volatile memory 550 into memory 540 and then runs it, forming a device for locating the target user at the logical level. Processor 510 executes the program stored in memory and specifically performs... Figures 1 to 3 The method described in the embodiments achieves the same or corresponding technical effects.

[0086] The above is as stated in this application. Figures 1 to 3The methods disclosed in the illustrated embodiments can be applied to a processor or implemented by processor 510. Processor 510 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by integrated logic circuits in the hardware of processor 510 or by instructions in software form. The processor 510 can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software module can reside in a mature storage medium in the field, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in the memory, and the processor 510 reads the information from the memory and, in conjunction with its hardware, completes the steps of the above method.

[0087] The electronic device can also execute the methods described in the preceding method embodiments and achieve the functions and beneficial effects of the methods described in the preceding method embodiments, which will not be repeated here.

[0088] Of course, in addition to software implementation, the electronic device of this application does not exclude other implementation methods, such as logic devices or a combination of hardware and software, etc. In other words, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.

[0089] This application also proposes a computer-readable storage medium that stores one or more programs, which, when executed by an electronic device including multiple applications, cause the electronic device to perform... Figures 1 to 3 The method described in the embodiments achieves the same or corresponding technical effects.

[0090] The computer-readable storage medium includes read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks, etc.

[0091] Furthermore, embodiments of this application also provide a computer program product, the computer program product including a computer program stored on a non-transitory computer-readable storage medium, the computer program including program instructions, which, when executed by a computer, implement... Figures 1 to 3 The method described in the embodiments achieves the same or corresponding technical effects.

[0092] In summary, the above description is merely a preferred embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.

[0093] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.

[0094] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0095] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0096] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.

Claims

1. A method for user matching, characterized in that, include: The user data sample whose number of second users is the first value is determined as the positive user data sample; The user data sample whose number of the second user is the second value is determined as the negative user data sample; The matching model is trained using the positive user data samples and the negative user data samples; The first user's first user data is input into the matching model, and the first user data includes at least two of the first user name, first user code, first user location, and first user contact information. Based on the probability value of each archived user matching the first user, the matching model outputs second user data, which is used to indicate the second user among the archived users who matches the first user. The user location includes the place of origin and location coordinates. Before outputting the second user data based on the probability values ​​of each of the archived users matching the first user using the matching model, the method further includes: The probability of the second user matching the first user is determined to be 0, and the data of mismatched users is discarded, provided that the following conditions are met: If the second user's second home location is unknown, the second user's second location coordinates do not meet the predetermined location conditions, and the similarity between the second user's name and the first user's name is lower than the third threshold. The second position coordinates do not meet the predetermined position conditions, including: The second position coordinates are located outside the intersection area of ​​multiple target circle center areas, wherein the target circle center area is an area determined according to the third position coordinates of each core member of the second user and a predetermined radius, and the third position coordinates satisfy a first predetermined condition; The second user data includes the first user data and the matching relationship between the first user and the second user.

2. The method according to claim 1, characterized in that, The user location includes the place of origin and location coordinates. Before outputting the second user data based on the probability values ​​of each of the archived users matching the first user using the matching model, the method further includes: The probability that the second user matches the first user is 0 if at least one of the following conditions is met: If the second user's second home location is known, the second user's second home location is different from the first user's first home location; If the second place of origin is known, the similarity between the second user's second username and the first username is less than the first threshold. When the second location is unknown, the distance between the second user's second location coordinates and the first user's first location coordinates is higher than a second threshold.

3. The method according to claim 1, characterized in that, The second value is different from the first value.

4. The method according to claim 2, characterized in that, The user data in the negative user data sample satisfies at least one of the following: The first user code is the same as the second user code of the second user, and the first location is different from the second location; The similarity between the first username and the second username is 100%, and the first location and the second location are different; The first user's contact information is the same as the second user's second user contact information, and the similarity between the first user name and the second user name is lower than the first threshold. The first location is the same as the second location, and the similarity between the first username and the second username is lower than the first threshold. The first location is the same as the second location, and the similarity between the first username and the second username is a third value, which satisfies a second predetermined condition.

5. A user matching device, characterized in that, include: The model training module is used to determine the user data samples whose number of second users is the first value as positive user data samples; The user data sample whose number of the second user is the second value is determined as the negative user data sample; The matching model is trained using the positive user data samples and the negative user data samples; The input module is used to input the first user's first user data into the matching model. The first user data includes at least two of the first user's name, first user code, first user location, and first user's contact information. The output module is used to output second user data based on the probability value of each archived user matching the first user through the matching model. The second user data is used to indicate the second user among the archived users who matches the first user. The user location includes the place of origin and location coordinates. Before outputting the second user data based on the probability values ​​of each of the archived users matching the first user using the matching model, the method further includes: The probability of the second user matching the first user is determined to be 0, and the data of mismatched users is discarded, provided that the following conditions are met: If the second user's second home location is unknown, the second user's second location coordinates do not meet the predetermined location conditions, and the similarity between the second user's name and the first user's name is lower than the third threshold. The second position coordinates do not meet the predetermined position conditions, including: The second position coordinates are located outside the intersection area of ​​multiple target circle center areas, wherein the target circle center area is an area determined according to the third position coordinates of each core member of the second user and a predetermined radius, and the third position coordinates satisfy a first predetermined condition; The second user data includes the first user data and the matching relationship between the first user and the second user.

6. An electronic device, comprising: processor; as well as A memory configured to store computer-executable instructions, which, when executed, use the processor to perform the user matching method of any one of claims 1-4.

7. A computer-readable medium storing one or more programs, which, when executed by an electronic device including a plurality of applications, cause the electronic device to perform the user matching method according to any one of claims 1-4.

Citation Information

Patent Citations

  • Hotel matching method and device, electronic equipment and storage medium

    CN114358979A

  • Matching method and device for offline merchant information and storage medium

    CN115392961A