A method, apparatus, device and medium for identifying a target user

By generating and traversing the similarity set of target statement vectors, the problem of not being able to accurately find the matching user in the search scenario is solved, and the accurate identification of the matching user is achieved.

CN118981644BActive Publication Date: 2025-05-27ZHEJIANG MEIRI HUDONG NETWORK TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411458344.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-18
Publication Date
2025-05-27
Estimated Expiration
2044-10-18

AI Technical Summary

Technical Problem

In the search scenario, users that are suitable for the text or statement to be matched cannot be accurately found, resulting in the inability to obtain matching user information.

Method used

By getting the target statement vector, the target similarity set is generated and iterated over the set to get the target user. The specific steps include obtaining the initial text set of the initial user group, converting it into the initial text vector set, generating the target similarity set of the target statement vector, and traversing the set to determine the target user.

Benefits of technology

It realizes accurate acquisition of users that are suitable for matching text or statements, and improves the accuracy of user matching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118981644B_ABST
    Figure CN118981644B_ABST
Patent Text Reader

Abstract

This application relates to the technical field of user identification, and particularly to a method, apparatus, device, and medium for identifying a target user. The method includes: obtaining a target sentence vector; obtaining a target similarity set of the target sentence vector according to the target sentence vector; traversing the target similarity set to obtain a target user; and being able to accurately obtain a user suitable for the matching text or sentence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of user identification, and in particular to a method, apparatus, device and medium for identifying a target user. Background Art

[0002] In a search scenario, a user inputs a text or statement to be matched, and the system needs to find the content in a corpus that is as similar as possible to the text to be matched and return the matching result to the user; however, it is impossible to accurately find the user suitable for the matching text or statement. Therefore, how to accurately obtain the user suitable for the matching text or statement is an urgent problem to be solved. Summary of the Invention

[0003] The purpose of the present invention is to provide a method for identifying a target user, and the method includes the following steps:

[0004] Obtain a target statement vector.

[0005] According to the target statement vector, obtain a target similarity set of the target statement vector.

[0006] Traverse the target similarity set to obtain a target user.

[0007] Specifically, the target statement vector refers to the vector of the statement input by the user.

[0008] Specifically, the step of obtaining a target similarity set of the target statement vector according to the target statement vector further includes the following steps:

[0009] Obtain an initial text set of an initial user group.

[0010] According to the initial text set of the initial user group, obtain an initial text vector set of the initial user group.

[0011] Generate a target similarity set of the target statement vector according to the target statement vector and the initial text vector set of the initial user group.

[0012] Specifically, the step of traversing the target similarity set to obtain a target user further includes the following steps:

[0013] Traverse the target similarity set to obtain several key similarities.

[0014] Take the initial user corresponding to any one of the key similarities as the target user.

[0015] The present invention also provides an apparatus for identifying a target user, and the apparatus includes:

[0016] A target statement vector acquisition module, configured to acquire a target statement vector;

[0017] A target similarity set acquisition module, configured to obtain a target similarity set of the target statement vector according to the target statement vector;

[0018] A target user acquisition module, configured to traverse the target similarity set to obtain a target user.

[0019] Specifically, the target statement vector refers to the vector of the statement input by the user.

[0020] Specifically, the target similarity set acquisition module further includes:

[0021] An initial text set acquisition module, configured to obtain an initial text set of an initial user group.

[0022] An initial text vector set acquisition module, configured to obtain an initial text vector set of the initial user group according to the initial text set of the initial user group.

[0023] A target similarity set generation module, configured to generate a target similarity set of the target statement vector according to the target statement vector and the initial text vector set of the initial user group.

[0024] Specifically, the target user acquisition module includes:

[0025] A key similarity acquisition module, configured to traverse the target similarity set to obtain a plurality of key similarities.

[0026] A target user determination module, configured to use the initial user corresponding to any one of the key similarities as the target user.

[0027] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, where when the processor executes the computer program, the above-mentioned identification of the target user is implemented.

[0028] The present invention also provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned identification of the target user is implemented.

[0029] The present invention has at least the following beneficial effects compared with the prior art:

[0030] The present invention obtains a target statement vector; according to the target statement vector, obtains a target similarity set of the target statement vector; traverses the target similarity set to obtain a target user; and can accurately obtain a user suitable for the matching text or statement. Description of the Drawings

[0031] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0032] Figure 1 It is a flowchart of a method for identifying a target user provided in Embodiment 1 of the present invention;

[0033] Figure 2 It is a schematic structural diagram of a device for identifying a target user provided in Embodiment 2 of the present invention. Detailed implementation manners

[0034] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.

[0035] Embodiment 1

[0036] As Figure 1 shown, Embodiment 1 of the present invention provides a method for identifying a target user, and the method includes the following steps:

[0037] S1. Obtain a target sentence vector.

[0038] Specifically, the target sentence vector refers to the vector of the input sentence.

[0039] S2. According to the target sentence vector, obtain a target similarity set of the target sentence vector.

[0040] Specifically, the step of obtaining a target similarity set of the target sentence vector according to the target sentence vector further includes the following steps:

[0041] S21. Obtain an initial text set of the initial user group.

[0042] S22. According to the initial text set of the initial user group, obtain an initial text vector set of the initial user group.

[0043] S23. According to the target sentence vector and the initial text vector set of the initial user group, generate a target similarity set of the target sentence vector.

[0044] In a specific embodiment, the initial text set of the initial text includes a number of the initial texts, where the initial text refers to the text used to characterize the user characteristic features of the initial user group. It can be understood that the initial text set A of the initial user group = {A 1 , A 2 , ……, A i , ……, A m}, A i = {A i1 , A i2 , ……, A ij , ……, A in(i)}, A ij = {A 1 ij , A 2 ij , ……, A r ij , ……, A s(ij) ij}, A r ij refers to the text of the r-th user characteristic feature of the j-th sub-initial user group of the i-th type of initial user group. The value range of i is from 1 to m, where m is the number of initial user group types. The value range of j is from 1 to n(i), and n(i) refers to the number of sub-initial user groups in the i-th type of initial user group. The value range of r is from 1 to s(ij), and s(ij) refers to the number of text of user characteristic features in the j-th sub-initial user group of the i-th type of initial user group. For example, the characteristic feature of the sub-initial user group is the user characteristic feature of users who have purchased a certain product more than 5 times.

[0045] Further, different sub-initial user groups in a certain type of initial user group refer to a single sub-user group in a certain initial user group.

[0046] Further, different types of initial user groups mean that the data sources of the initial user groups are different.

[0047] In a specific embodiment, the initial text vector set of the initial user group includes a number of initial text vectors. It can be understood that the initial text vector set B of the initial user group = {B 1 , B 2 , ……, B i , ……, B m}, B i = {B i1 , B i2 , ……, B ij , ……, B in(i)}, B ij = {B 1ij , B 2 ij , ……, B r ij , ……, B s(ij) ij}, B r ij refers to A r ij The corresponding text vector, where those skilled in the art know the method of converting text into a text vector, which will not be elaborated here. For example, the text is input into the Bert model to obtain the text vector.

[0048] In a specific embodiment, the target similarity set of the target statement vectors includes the target similarities of a number of target statement vectors, which can be understood as: the target similarity set F = {F 1 , F 2 , ……, F x , ……, F p}, F x = {F xi1 , F xi2 , ……, F xy , ……, F xq(x)}, F xy = {F 1 xy , F 2 xy , ……, F g xy , ……, F z(xy) xy}, F g xy refers to the similarity between the target statement vector and the text vector corresponding to the g-th user characteristic feature of the y-th sub-middle user group of the middle user group of the x-th category. The value range of g is from 1 to p, where p is the number of middle user group types, the value range of y is from 1 to q(x), q(x) refers to the number of sub-middle user groups in the i-th category of middle user groups, and the value range of g is from 1 to z(xy), where z(xy) refers to the number of similarities between the text vectors corresponding to the user characteristic features within the y-th sub-middle user group of the x-th category of middle user groups.

[0049] Specifically, it further includes the following steps:

[0050] S100, obtain the first specified text corresponding to A ij which can be understood as: After separating A 1 ij to A s(ij) ij with a delimiter character, they are merged into a unified text.

[0051] S200, from A ij Extract the first keyword occurrence list C from the corresponding first specified text ij ={C 1 ij , C 2 ij , ……, C e ij , ……, C h ij}, C e ij refers to the occurrence times of the e-th keyword extracted from the corresponding first specified text of A, where the value range of e is from 1 to h, and h refers to the number of keywords extracted from the corresponding first specified text of A ij ; it can be understood that: take the number of initial texts of the keywords extracted from the corresponding first specified text of A as the occurrence times of the keywords ij ; it can be understood that: take the number of initial texts of the keywords extracted from the corresponding first specified text of A as the occurrence times of the keywords ij ; it can be understood that: take the number of initial texts of the keywords extracted from the corresponding first specified text of A as the occurrence times of the keywords

[0052] S300, when C e ij ≥C 0 , determine the corresponding first keyword as the first intermediate keyword, where C e ij is the first threshold of the occurrence times of the preset keyword. Those skilled in the art can set the threshold of the occurrence times of the preset keyword according to actual needs, which will not be elaborated here 0 ; it can be understood that: take the number of initial texts of the keywords extracted from the corresponding first specified text of A as the occurrence times of the keywords

[0053] S400, obtain the corresponding second specified text of A, which can be understood as: after separating the corresponding first specified text of A to the corresponding first specified text of A with a delimiter character, merge them into a unified text, where the acquisition method of the corresponding first specified text of A to the corresponding first specified text of A i refers to the acquisition method of the corresponding first specified text of A, which is referred to in step S100 and will not be elaborated here i1 to the corresponding first specified text of A in(i) ; it can be understood that: take the number of initial texts of the keywords extracted from the corresponding first specified text of A as the occurrence times of the keywords i1 to the corresponding first specified text of A in(i) in the corresponding first specified text of A ij to the corresponding first specified text of A

[0054] S500, extract the second keyword occurrence list D from the corresponding second specified text of A i ={D i , D i1 , ……, D i2 , ……, D it , ……, D ik}, D it refers to the occurrence times of the t-th keyword extracted from the corresponding second specified text of A, where the value range of t is from 1 to k, and k refers to the number of keywords extracted from the corresponding second specified text of A i ; it can be understood that: take the number of initial texts of the keywords extracted from the corresponding second specified text of A as the occurrence times of the keywordsi The number of keywords extracted from the corresponding second specified text; it can be understood that when A appears i The initial number of texts of the keywords extracted from the corresponding second specified text is used as the number of occurrences of the keywords.

[0055] S600, when D it ≥ D 0 At this time, determine D it The corresponding first keyword is used as the second intermediate keyword, where D 0 Is the second threshold of the number of occurrences of the preset keyword. Those skilled in the art can set the threshold of the number of occurrences of the preset keyword according to actual needs, which will not be elaborated here.

[0056] Furthermore, C 0 Is inconsistent with D 0

[0057] S700, obtain the similarity between the target sentence vector and the word vector of any second intermediate keyword. Those skilled in the art know the method of obtaining the similarity between two vectors in the prior art, which will not be elaborated here.

[0058] S800, when the similarity between the target sentence vector and the word vector of any second intermediate keyword is not less than the preset first similarity threshold, obtain the initial user group corresponding to the determined second intermediate keyword as the intermediate user group. Those skilled in the art set the first similarity threshold according to actual needs, which will not be elaborated here.

[0059] S900, obtain the similarity between the target sentence vector and the word vector of any first intermediate keyword corresponding to a certain intermediate user group. Those skilled in the art know the method of obtaining the similarity between two vectors in the prior art, which will not be elaborated here.

[0060] S1000, when the similarity between the target sentence vector and the word vector of any first intermediate keyword corresponding to a certain intermediate user group is not less than the preset second similarity threshold, obtain the sub-initial user group corresponding to any first intermediate keyword corresponding to the determined intermediate user group as the sub-intermediate user group. Those skilled in the art set the second similarity threshold according to actual needs, which will not be elaborated here; it can be understood that the method for obtaining any first intermediate keyword corresponding to a certain intermediate user group can refer to the steps of S100 - S300.

[0061] S1100, use the similarity between the target sentence vector and the word vector of any first intermediate keyword corresponding to a certain intermediate user group that is not less than the preset second similarity threshold as F g xy

[0062] ​​As described above, for different types of user groups, the applicable user group is selected by the keywords extracted from the text of the user group and based on the similarity between the keywords extracted from the text of the user group and the target sentence vector, so that the target similarity can be accurately determined, and according to the target similarity, the target user matching the target sentence vector is selected.

[0063] S3. Traverse the set of target similarities to obtain the target user.

[0064] Specifically, the step of traversing the set of target similarities to obtain the target user further includes the following steps:

[0065] S31. Traverse the set of target similarities to obtain a number of key similarities; it can be understood that: traverse F and when F g xy ≥F 0 At this time, take F g xy As the key similarity, where F 0 Is a preset third similarity threshold, and those skilled in the art set the third similarity threshold according to actual needs, which will not be elaborated here.

[0066] S32. Take the initial user corresponding to any of the key similarities as the target user.

[0067] As described above, by selecting through the similarity of the text of different types of user groups, the user who meets the input sentence vector can be accurately selected, so as to recommend an appropriate user.

[0068] Embodiment 2

[0069] As Figure 2 Shown, Embodiment 2 of the present invention provides a target user recognition device, and the device includes the following steps:

[0070] Target sentence vector acquisition module 1, configured to acquire a target sentence vector.

[0071] Specifically, the target sentence vector refers to the vector of the input sentence.

[0072] Target user acquisition module 2, configured to acquire a set of target similarities of the target sentence vector according to the target sentence vector.

[0073] Specifically, the target user acquisition module 2 further includes:

[0074] Initial text set acquisition module 21, configured to acquire an initial text set of an initial user group.

[0075] Initial text vector set acquisition module 22, configured to acquire an initial text vector set of the initial user group according to the initial text set of the initial user group.

[0076] The target similarity set generation module 23 is configured to generate a target similarity set of the target statement vector according to the target statement vector and the initial text vector set of the initial user group.

[0077] In a specific embodiment, the initial text set of the initial text includes a plurality of the initial texts, where the initial text refers to the text used to characterize the user characteristic features of the initial user group, and it can be understood that: the initial text set A of the initial user group = {A 1 , A 2 , ……, A i , ……, A m}, A i = {A i1 , A i2 , ……, A ij , ……, A in(i)}, A ij = {A 1 ij , A 2 ij , ……, A r ij , ……, A s(ij) ij}, A r ij refers to the text of the r-th user characteristic feature of the j-th sub-initial user group of the i-th class of the initial user group. The value range of i is from 1 to m, where m is the number of initial user group types. The value range of j is from 1 to n(i), and n(i) refers to the number of sub-initial user groups in the i-th class of the initial user group. The value range of r is from 1 to s(ij), and s(ij) refers to the number of texts of user characteristic features in the j-th sub-initial user group of the i-th class of the initial user group; for example, the characteristic feature of the sub-initial user group is the user characteristic feature of purchasing a certain product more than 5 times.

[0078] Further, different sub-initial user groups in a certain type of initial user group refer to a single sub-user group in a certain type of initial user group.

[0079] Further, different types of initial user groups mean that the data sources of the initial user groups are different.

[0080] In a specific embodiment, the initial text vector set of the initial user group includes a plurality of initial text vectors, and it can be understood that: the initial text vector set B of the initial user group = {B 1 , B 2 , ……, B i , ……, B m}, B i={B i1 ,B i2 ,……,B ij ,……,B in(i)},B ij ={B 1 ij ,B 2 ij ,……,B r ij ,……,B s(ij) ij},B r ij refers to the text vector corresponding to A r ij Those skilled in the art know the method of converting text into a text vector, which will not be elaborated here. For example, the text is input into the Bert model to obtain the text vector.

[0081] In a specific embodiment, the target similarity set of the target statement vector includes the target similarities of a number of target statement vectors. It can be understood that: the target similarity set F = {F 1 ,F 2 ,……,F x ,……,F p}, F x ={F xi1 ,F xi2 ,……,F xy ,……,F xq(x)}, F xy ={F 1 xy ,F 2 xy ,……,F g xy ,……,F z(xy) xy}, F g xy refers to the similarity between the target statement vector and the text vector corresponding to the g-th user characteristic feature of the y-th sub-middle user group of the middle user group of the x-th category. The value range of g is from 1 to p, where p is the number of middle user group types. The value range of y is from 1 to q(x), where q(x) refers to the number of sub-middle user groups in the i-th category of middle user groups. The value range of g is from 1 to z(xy), where z(xy) refers to the number of similarities between the text vectors corresponding to the user characteristic features within the y-th sub-middle user group of the x-th category of middle user groups.

[0082] Specifically, it further includes:

[0083] The first specified text acquisition module 100 is used to obtain A ijThe corresponding first specified text can be understood as: 1 ij To A s(ij) ij After being separated by a separator character, they are merged into a unified text.

[0084] The first keyword occurrence count list acquisition module 200 is used to obtain the number of occurrences of the first keyword from A ij The first keyword occurrence count list C is extracted from the corresponding first specified text ij ={C 1 ij , C 2 ij , ..., C e ij , ..., C h ij}, C e ij A ij The number of occurrences of the e-th keyword in the corresponding first specified text is extracted, and the value range of e is 1 to h, where h refers to A ij The number of keywords extracted from the corresponding first specified text can be understood as: ij The number of initial texts of the keyword extracted from the corresponding first designated text is used as the number of occurrences of the keyword.

[0085] The first intermediate keyword determination module 300 is used when C e ij ≥C 0 When C e ij The corresponding first keyword is used as the first intermediate keyword, where C 0 It is a first threshold value of the number of occurrences of a preset keyword. Those skilled in the art can set the threshold value of the number of occurrences of the preset keyword according to actual needs, which will not be described in detail here.

[0086] The second designated text acquisition module 400 is used to acquire A i The corresponding second specified text can be understood as: i1 The first specified text corresponding to A in(i) The corresponding first specified text is separated by a separator character and then merged into a unified text, where A i1 The first specified text corresponding to A in(i) The corresponding first specified text A ij The corresponding method for obtaining the first designated text refers to step S100 and will not be repeated here.

[0087] The second keyword occurrence count list acquisition module 500 is used to obtain the number of occurrences of the second keyword from A iExtract the list D of the occurrence times of the second keyword from the corresponding second specified text i ={D i1 , D i2 , ……, D it , ……, D ik}, D it refers to A i Extract the occurrence times of the t-th keyword from the corresponding second specified text, where the value range of t is from 1 to k, and k refers to A i The number of keywords extracted from the corresponding second specified text; it can be understood that: occurrence A i Take the initial text quantity of the keywords extracted from the corresponding second specified text as the occurrence times of the keywords.

[0088] The second intermediate keyword determination module 600 is used to determine the first keyword corresponding to D it ≥D 0 as the second intermediate keyword, where D it is the second threshold of the occurrence times of the preset keyword. Those skilled in the art can set the threshold of the occurrence times of the preset keyword according to actual needs, which will not be elaborated here. 0

[0089] Furthermore, C 0 is inconsistent with D 0 .

[0090] The first execution module 700 is used to obtain the similarity between the target statement vector and the word vector of any second intermediate keyword. Those skilled in the art know the method of obtaining the similarity between two vectors in the prior art and will not elaborate here.

[0091] The second execution module 800 is used to obtain the initial user group corresponding to the determined second intermediate keyword as the intermediate user group when the similarity between the target statement vector and the word vector of any second intermediate keyword is not less than the preset first similarity threshold. Those skilled in the art set the first similarity threshold according to actual needs and will not elaborate here.

[0092] The third execution module 900 is used to obtain the similarity between the target statement vector and the word vector of any first intermediate keyword corresponding to a certain intermediate user group. Those skilled in the art know the method of obtaining the similarity between two vectors in the prior art and will not elaborate here.

[0093] ​The fourth execution module 1000 is configured to, when the similarity between the target statement vector and the word vector of any first intermediate keyword corresponding to a certain intermediate user group is not less than a preset second similarity threshold, obtain the sub-initial user group corresponding to any first intermediate keyword corresponding to the determined intermediate user group as the sub-intermediate user group. Those skilled in the art can set the second similarity threshold according to actual needs and will not elaborate further here. It can be understood that the method for obtaining any first intermediate keyword corresponding to a certain intermediate user group can refer to the steps of S100 - S300.

[0094] The fifth execution module 1100 is configured to use the similarity between the target statement vector whose similarity is not less than the preset second similarity threshold and the word vector of any first intermediate keyword corresponding to a certain intermediate user group as F. g xy .

[0095] As described above, for different types of user groups, by extracting keywords from the text of the user group and selecting the applicable user group according to the similarity between the keywords extracted from the text of the user group and the target statement vector, the target similarity can be accurately determined, and according to the target similarity, the target user that matches the target statement vector can be selected.

[0096] The target user acquisition module 3 is configured to traverse the target similarity set to obtain the target user.

[0097] Specifically, the target user acquisition module 3 further includes:

[0098] The key similarity acquisition module 31 is configured to traverse the target similarity set to obtain several key similarities. It can be understood that by traversing F and when F g xy ≥F 0 , F g xy is used as the key similarity, where F 0 is the preset third similarity threshold. Those skilled in the art can set the third similarity threshold according to actual needs and will not elaborate further here.

[0099] The target user determination module 32 is configured to use the initial user corresponding to any of the key similarities as the target user.

[0100] As described above, by selecting through the similarity of the text of different types of user groups, the user that meets the input statement vector can be accurately selected, so as to recommend an appropriate user.

[0101] Embodiment III

[0102] This embodiment provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented:

[0103] Obtain a target statement vector;

[0104] According to the target statement vector, obtain a target similarity set of the target statement vector;

[0105] Traverse the target similarity set to obtain a target user.

[0106] Embodiment 4

[0107] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0108] Obtain a target statement vector;

[0109] According to the target statement vector, obtain a target similarity set of the target statement vector;

[0110] Traverse the target similarity set to obtain a target user.

[0111] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0112] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules as needed, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above.

[0113] Although some specific embodiments of the present invention have been described in detail by way of examples, those skilled in the art should understand that the above examples are only for illustration and not for limiting the scope of the present invention. Those skilled in the art should also understand that various modifications can be made to the embodiments without departing from the scope and spirit of the present invention. The scope of the present invention is defined by the appended claims.

Claims

1. A method for identifying a target user, characterized in that: The method comprises the following steps: Get the target sentence vector; According to the target sentence vector, a target similarity set of the target sentence vector is obtained; wherein the target similarity set F={F1, F2, ..., F x , ..., F p }, F x ={F xi1 , F xi2 , ..., F xy , ..., F xq(x) }, F xy ={F 1 xy , F 2 xy , ..., F g xy , ..., F z(xy) xy }, F g xy It refers to the similarity between the target sentence vector and the text vector corresponding to the text of the g-th user characteristic feature of the y-th sub-intermediate user group of the x-th class intermediate user group. The value range of g is 1 to p, p is the number of intermediate user group types, the value range of y is 1 to q(x), q(x) refers to the number of sub-intermediate user groups in the ith class intermediate user group, and the value range of g is 1 to z(xy), z(xy) refers to the number of similarities between the text vectors corresponding to the text of the user characteristic feature in the y-th sub-intermediate user group in the x-th class intermediate user group. Traversing the target similarity set to obtain the target user; The following steps are also included: S100, get A ij The corresponding first specified text; S200, from A ij The first keyword occurrence count list C is extracted from the corresponding first specified text ij ={C 1 ij , C 2 ij , ..., C e ij , ..., C h ij }, C e ij A ij The number of occurrences of the e-th keyword in the corresponding first specified text is extracted, and the value range of e is 1 to h, where h refers to A ij The number of keywords extracted from the corresponding first specified text; S300, when C e ij ≥C0, determine C e ij The corresponding first keyword is used as the first intermediate keyword, wherein C0 is the first threshold value of the number of occurrences of the preset keyword; S400, get A i The corresponding second specified text is A i1 The first specified text corresponding to A in(i) The corresponding first specified text is separated by a separator character and then merged into a unified text, where A i1 The first specified text corresponding to A in(i) The corresponding first specified text A ij The corresponding method for obtaining the first specified text is as in step S100; S500, from A i The corresponding second specified text extracts the second keyword occurrence count list D i ={D i1 , D i2 , ..., D it , ..., D ik }, D it A i The number of occurrences of the t-th keyword in the corresponding second specified text is extracted, and the value range of t is 1 to k, where k refers to A i the number of keywords extracted from the corresponding second specified text; S600, when D it ≥D0, determine D it The corresponding first keyword is used as the second intermediate keyword, wherein D0 is the second threshold value of the number of occurrences of the preset keyword; S700, obtaining the similarity between the target sentence vector and the word vector of any second intermediate keyword; S800, when the similarity between the target sentence vector and the word vector of any second intermediate keyword is not less than a preset first similarity threshold, obtaining and determining an initial user group corresponding to the second intermediate keyword as the intermediate user group; S900, obtaining the similarity between the target sentence vector and the word vector of any first intermediate keyword corresponding to a certain intermediate user group; S1000, when the similarity between the target sentence vector and the word vector of any first intermediate keyword corresponding to a certain intermediate user group is not less than a preset second similarity threshold, obtaining a sub-initial user group corresponding to any first intermediate keyword corresponding to the intermediate user group as a sub-intermediate user group; S1100: The similarity between the target sentence vector that is not less than the preset second similarity threshold and the word vector of any first intermediate keyword corresponding to a certain intermediate user group is taken as F g xy .

2. The target user identification method according to claim 1, characterized in that: The target sentence vector refers to the vector of the sentence input by the user.

3. The target user identification method according to claim 1, characterized in that: According to the target sentence vector, the step of obtaining a target similarity set of the target sentence vector also includes the following steps: Obtaining an initial text set of an initial user group; Acquire an initial text vector set of the initial user group according to the initial text set of the initial user group; A target similarity set of the target sentence vector is generated according to the target sentence vector and the initial text vector set of the initial user group.

4. The target user identification method according to claim 1, characterized in that: The step of traversing the target similarity set to obtain the target user also includes the following steps: Traversing the target similarity set to obtain a number of key similarities; An initial user corresponding to any of the key similarities is taken as a target user.

5. A target user identification device, characterized in that: The device comprises: A target sentence vector acquisition module is used to acquire a target sentence vector; The target similarity set acquisition module is used to acquire the target similarity set of the target sentence vector according to the target sentence vector; wherein the target similarity set F={F1, F2, ..., F x , ..., F p }, F x ={F xi1 , F xi2 , ..., F xy , ..., F xq(x) }, F xy ={F 1 xy , F 2 xy , ..., F g xy , ..., F z(xy) xy }, F g xy It refers to the similarity between the target sentence vector and the text vector corresponding to the text of the g-th user characteristic feature of the y-th sub-intermediate user group of the x-th class intermediate user group. The value range of g is 1 to p, p is the number of intermediate user group types, the value range of y is 1 to q(x), q(x) refers to the number of sub-intermediate user groups in the ith class intermediate user group, and the value range of g is 1 to z(xy), z(xy) refers to the number of similarities between the text vectors corresponding to the text of the user characteristic feature in the y-th sub-intermediate user group in the x-th class intermediate user group. A target user acquisition module, used to traverse the target similarity set and acquire the target user; Among them, it also includes: The first designated text acquisition module 100 is used to acquire A ij The corresponding first specified text; The first keyword occurrence count list acquisition module 200 is used to obtain the number of occurrences of the first keyword from A ij The first keyword occurrence count list C is extracted from the corresponding first specified text ij ={C 1 ij , C 2 ij , ..., C e ij , ..., C h ij }, C e ij A ij The number of occurrences of the e-th keyword in the corresponding first specified text is extracted, and the value range of e is 1 to h, where h refers to A ij The number of keywords extracted from the corresponding first specified text; The first intermediate keyword determination module 300 is used when C e ij ≥C0, determine C e ij The corresponding first keyword is used as the first intermediate keyword, wherein C0 is the first threshold value of the number of occurrences of the preset keyword; The second designated text acquisition module 400 is used to acquire A i The corresponding second specified text is A i1 The first specified text corresponding to A in(i) The corresponding first specified text is separated by a separator character and then merged into a unified text, where A i1 The first specified text corresponding to A in(i) The corresponding first specified text A ij The corresponding method for obtaining the first specified text is as in step S100; The second keyword occurrence count list acquisition module 500 is used to obtain the number of occurrences of the second keyword from A i The corresponding second specified text extracts the second keyword occurrence count list D i ={D i1 , D i2 , ..., D it , ..., D ik }, D it A i The number of occurrences of the t-th keyword in the corresponding second specified text is extracted, and the value range of t is 1 to k, where k refers to A i the number of keywords extracted from the corresponding second specified text; The second intermediate keyword determination module 600 is used when D it ≥D0, determine D it The corresponding first keyword is used as the second intermediate keyword, wherein D0 is the second threshold value of the number of occurrences of the preset keyword; A first execution module 700 is used to obtain the similarity between the target sentence vector and the word vector of any second intermediate keyword; The second execution module 800 is used to obtain and determine the initial user group corresponding to the second intermediate keyword as the intermediate user group when the similarity between the target sentence vector and the word vector of any second intermediate keyword is not less than a preset first similarity threshold; The third execution module 900 is used to obtain the similarity between the target sentence vector and the word vector of any first intermediate keyword corresponding to a certain intermediate user group; The fourth execution module 1000 is used to obtain a sub-initial user group corresponding to any first intermediate keyword corresponding to a certain intermediate user group as a sub-intermediate user group when the similarity between the target sentence vector and the word vector of any first intermediate keyword corresponding to a certain intermediate user group is not less than a preset second similarity threshold; The fifth execution module 1100 is used to use the similarity between the target sentence vector that is not less than the preset second similarity threshold and the word vector of any first intermediate keyword corresponding to a certain intermediate user group as F g xy .

6. The target user identification device according to claim 5, characterized in that: The target sentence vector refers to the vector of the sentence input by the user.

7. The target user identification device according to claim 5, characterized in that: The target similarity set acquisition module also includes: An initial text set acquisition module, used to acquire an initial text set of an initial user group; An initial text vector set acquisition module, used to acquire an initial text vector set of the initial user group according to the initial text set of the initial user group; The target similarity set generation module is used to generate a target similarity set of the target sentence vector according to the target sentence vector and the initial text vector set of the initial user group.

8. The target user identification device according to claim 5, characterized in that: The target user acquisition module includes: A key similarity acquisition module, used to traverse the target similarity set and acquire a number of key similarities; The target user determination module is used to take the initial user corresponding to any of the key similarities as the target user.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the target user identification method according to any one of claims 1 to 4 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the target user identification method according to any one of claims 1 to 4 is implemented.

Citation Information

Patent Citations

  • Article author identity recognition and evaluation model training method and device and storage medium

    CN110059180A