Data tag determination method, electronic equipment and storage medium

By using a large model to refine and determine the target field name, the problem of inapplicable potential direction acquisition of data applications in the prior art is solved, and accurate and comprehensive data label acquisition is achieved.

CN120144588AInactive Publication Date: 2025-06-13ZHEJIANG MEIRI HUDONG NETWORK TECH CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510241488.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-06-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art is not applicable enough when acquiring the potential direction of data application, and the principal component analysis method has high requirements for data and it is difficult to directly define components.

Method used

By obtaining the data list of the target user, using the big model to label the target field name, and combining the quantity weight and the field value list, the target field and target label of the target field name are determined.

Benefits of technology

Accurate and comprehensive data labels are achieved, all possible labels are determined based on the big model and field names, and target labels are determined based on the domain.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144588A_ABST
    Figure CN120144588A_ABST
Patent Text Reader

Abstract

The invention provides a data label determination method, electronic equipment and a storage medium, and relates to the technical field of data processing.The method comprises the steps that a data list of a target user is obtained, a single field name corresponding to a field is obtained, and the target field name is input into a target large model; obtaining a first purpose label list corresponding to a target field name and an initial probability value corresponding to each first purpose label of the target field name, obtaining a combined field name list, inputting the combined field name into a target large model, obtaining a second purpose label list corresponding to the combined field name, obtaining a quantity weight list, and inputting the quantity weight list into the target large model; obtaining a weight corresponding to the second purpose label and marking the weight as a second purpose label weight; obtaining a domain value list corresponding to the second purpose label; determining a target domain corresponding to the target field name; obtaining a domain probability value of the first purpose label in the target domain; therefore, the data label can be obtained more accurately and comprehensively.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly to a method for determining data tags, an electronic device, and a storage medium. Background Art

[0002] With the development of big data, enterprises attach increasing importance to data assets. At the same time, enterprises pay more and more attention to promoting the construction of data basic systems, overall planning of data resource integration, sharing, development and utilization, and promoting the planning and construction of Digital China, digital economy, and digital society. As the assets of enterprises, enterprises need to take inventory of their data and know the potential application directions of the data.

[0003] In the prior art, in the data analysis method of dividing data points into different groups through cluster analysis, the data points within the group have high similarity, but cluster analysis is often applicable to the user segmentation field and is not very suitable for obtaining the potential directions of data applications; the principal component analysis method (PCA) is used for label extraction, but the principal component analysis method has high requirements for data and it is difficult to directly define the components. Summary of the Invention

[0004] In view of the above technical problems, the technical solution adopted by the present invention is: a method for determining data tags, the method comprising the following steps:

[0005] S100, obtaining a data list of a target user, the data list of the target user including a plurality of fields;

[0006] S200, obtaining a single field name corresponding to the field;

[0007] S300, marking any single field name as a target field name, and obtaining a target tag corresponding to the target field name;

[0008] Wherein, S300 includes the following steps:

[0009] S310, inputting the target field name into a target large model, obtaining a list of first usage tags corresponding to the target field name and an initial probability value corresponding to each first usage tag of the target field name, the list of first usage tags including a plurality of first usage tags;

[0010] S320, combining the target field name with a plurality of other single field names except the target field name to obtain a list of combined field names, and inputting the list of combined field names into the target large model to obtain a list of second usage tags corresponding to the combined field names; the list of combined field names includes a plurality of combined field names, and the list of second usage tags includes a plurality of second usage tags;

[0011] S330. Obtain a quantity weight list, and based on the quantity weight list and the combined field name, obtain the weight corresponding to the second usage label and mark it as the second usage label weight; wherein, the quantity weight list includes a number of quantity weights, and the quantity weight is the weight between the number of field names in the combined field name and the field;

[0012] S340. Input the second usage label into the target large model to obtain a list of field values corresponding to the second usage label, and the list of field values includes probability values of the second usage label in a number of preset fields;

[0013] S350. Based on the second usage label weight and the list of field values corresponding to the second usage label, determine the target field corresponding to the target field name;

[0014] S360. Obtain the field probability value of the first usage label in the target field, and based on the field probability value of the first usage label in the target field and the initial probability value corresponding to the first usage label, determine the target label corresponding to the target field name.

[0015] According to another aspect of the present invention, there is provided a non-transitory computer-readable storage medium, in which at least one instruction or at least one program segment is stored, and the at least one instruction or the at least one program segment is loaded and executed by a processor to implement the foregoing method.

[0016] According to still another aspect of the present invention, there is provided an electronic device, including a processor and the foregoing non-transitory computer-readable storage medium.

[0017] The present invention has at least the following beneficial effects: In summary, obtain the data list of the target user, obtain the single field name corresponding to the field, input the target field name into the target large model, obtain the list of first usage labels corresponding to the target field name and the initial probability value corresponding to each first usage label of the target field name, combine the target field name with several other single field names except the target field name to obtain a list of combined field names, input the combined field name into the target large model, obtain the list of second usage labels corresponding to the combined field name, obtain the quantity weight list, and based on the quantity weight list and the combined field name, obtain the weight corresponding to the second usage label and mark it as the second usage label weight, input the second usage label into the target large model, obtain the list of field values corresponding to the second usage label, based on the second usage label weight and the list of field values corresponding to the second usage label, determine the target field corresponding to the target field name, obtain the field probability value of the first usage label in the target field, and based on the field probability value of the first usage label in the target field and the initial probability value corresponding to the first usage label, determine the target label corresponding to the target field name. The present invention determines all possible labels of the field name based on the large model and the field name, and determines the target label based on the field, accurately and comprehensively obtaining the data label. Brief Description of the Drawings

[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required in the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0019] Figure 1 It is a flowchart of a method for determining data tags provided by an embodiment of the present invention. Detailed Embodiments

[0020] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0021] Embodiment 1

[0022] Embodiment 1 of the present invention provides a method for determining data tags, as Figure 1 shown, the method includes the following steps:

[0023] S100. Obtain a data list of a target user, where the data list of the target user includes several fields.

[0024] Specifically, the data list A of the target user = {A 1 , A 2 , …, A i , …, A m}, A i is the i-th field, and the value range of i is from 1 to m, where m is the number of fields of the target user.

[0025] S200. Obtain a single field name corresponding to the field.

[0026] S300. Mark any single field name as a target field name, and obtain a target tag corresponding to the target field name.

[0027] Among them, S300 includes the following steps:

[0028] S310. Input the target field name into the target large model, and obtain a list of first use tags corresponding to the target field name and an initial probability value for each first use tag of the target field name. The list of first use tags includes several first use tags.

[0029] Specifically, the target large model is a target large language model. Specifically, historical field names and historical usage labels corresponding to the historical field names are obtained, an initial large language model is constructed, and the initial large language model is trained using the historical field names and the historical usage labels corresponding to the historical field names to obtain a target large language model.

[0030] S320, combine the target field name with several other single field names except the target field name, obtain a list of combined field names, and input the combined field names into the target large model to obtain a second purpose label list corresponding to the combined field names; the combined field name list includes several combined field names, and the second purpose label list includes several second purpose labels.

[0031] Specifically, the target field name N i The corresponding combined field name list C i ={C i1 , C i2 , …, C ij , …, C in}, C ij A i The jth combination field name of C, where j ranges from 1 to n and n is the number of combination field names. ij Input the target large model and obtain C ij Corresponding second purpose label list D ij ={D ij1 , D ij2 , …, D ijr , …, D ijs}, D ijr It is C ij The corresponding r-th second purpose label, where r ranges from 1 to s, and s is the number of second purpose labels.

[0032] S330, obtaining a quantity weight list, and based on the quantity weight list and the combined field name, obtaining a weight corresponding to the second use label and marking it as a second use label weight; wherein the quantity weight list includes a plurality of quantity weights, and the quantity weight is the weight between the quantity of field names in the combined field name and the field. Specifically, obtaining the second use weight list F ij ={F ij1 , F ij2 , …, F ijr , …, F ijs}, F ijr Yes D ijr The corresponding weight.

[0033] S340. Input the second usage label into the target large model to obtain a list of domain values corresponding to the second usage label. The list of domain values includes probability values of the second usage label in a number of preset domains. Input the second usage label D ijr into the target large model to obtain the second usage label D ijr corresponding list of domain values E ijr ={E ijr1 , E ijr2 , …, E ijrg , …, E ijrz}, H={H 1 , …, H g , …, H z}, E ijrg is the probability value of D ijr in H g , and H g is the g-th preset domain.

[0034] S350. Based on the second usage label weight and the list of domain values corresponding to the second usage label, determine the target domain corresponding to the target field name.

[0035] S360. Obtain the domain probability value of the first usage label in the target domain, and based on the domain probability value of the first usage label in the target domain and the initial probability value corresponding to the first usage label, determine the target label corresponding to the target field name.

[0036] In summary, obtain the data list of the target user, obtain a single field name corresponding to the field, input the target field name into the target large model, obtain the list of first usage labels corresponding to the target field name and the initial probability value corresponding to each first usage label for the target field name, combine the target field name with several other single field names except the target field name to obtain a list of combined field names, input the combined field names into the target large model, obtain the list of second usage labels corresponding to the combined field names, obtain the quantity weight list, and based on the quantity weight list and the combined field names, obtain the weight corresponding to the second usage label and mark it as the second usage label weight, input the second usage label into the target large model, obtain the list of domain values corresponding to the second usage label, based on the second usage label weight and the list of domain values corresponding to the second usage label, determine the target domain corresponding to the target field name, obtain the domain probability value of the first usage label in the target domain, and based on the domain probability value of the first usage label in the target domain and the initial probability value corresponding to the first usage label, determine the target label corresponding to the target field name. The present invention determines all possible labels of the field name based on the large model and the field name, and determines the target label based on the domain, accurately and comprehensively obtaining the data label.

[0037] Specifically, S200 further includes:

[0038] S210. Obtain a list of preset field features, where the list of preset field features includes several preset field features. The preset field features include the total length of the field, the number of numeric characters in the field, the number of English characters in the field, etc.

[0039] S220. Obtain a corresponding list of preset field feature values, where the list of preset field feature values includes the preset field feature values corresponding to several preset field features.

[0040] S230. Determine the single field name of the field based on the list of preset field feature values.

[0041] In summary, obtain a list of preset field features, obtain a corresponding list of preset field feature values, and determine the single field name of the field based on the list of preset field feature values. The present invention determines the single field name based on the list of preset field feature values, rather than directly obtaining the field name of the filled field, to ensure the accuracy of the field name.

[0042] Specifically, S320 includes:

[0043] S321. Obtain a list of single field names N = {N 1 , N 2 , …, N i , …, N m}, and obtain a list of other field names P = {P i , P 1 , …, P 2 , …, P p , …, P n-1}, where P p is the p-th other field name, and the value range of p is from 1 to n - 1.

[0044] S322. Take any number of other field names from the list of other field names P and combine them with the target field name N i to obtain a combined field name. Extract through the extraction method of combination numbers to obtain the combined field name, so as to ensure that all combination methods of the target field name N i are combined.

[0045] Specifically, the quantity weight list is obtained through the following steps in S330:

[0046] S331. Obtain a list of historical field quantities R = {R 1 , R 2 , …, R a , …, R b}, where R a = {R a1 , R a2 , …, R ac , …, R ad}, R ac is the c-th sample of the a-th historical field quantity, where c ranges from 1 to d, d is the number of samples, and a ranges from 1 to b, b is the number of historical field quantities; among them, the a-th historical field quantity includes a field names.

[0047] S332. Obtain the historical domain list set T = {T 1 , T 2 , …, T a , …, T b}, where the a-th historical domain list T a= {T a1 , T a2 , …, T ac , …, T ad}, and T ac is the domain corresponding to R ac .

[0048] S333. Based on the historical field quantity list R and the historical domain list T, obtain the correlation coefficient list U = {U 1 , U 2 , …, U a , …, U b}, where the a-th correlation coefficient U a = [∑ d c=1 [(R ac - ER a ) × (T ac - ET a )]] / [(∑ d c=1 ((R ac - ER a )) 2 ) 1 / 2 × ∑ d c=1 ((T ac - ET a )) 2 )]] 1 / 2 , ER a = (∑ d c=1 R ac ) / d, ET a = (∑ d c=1 T ac ) / d.

[0049] S334. Normalize U 1 to U b to obtain the quantity weight list.

[0050] Specifically, those skilled in the art know that any method of normalization in the prior art belongs to the protection scope of the present invention and will not be elaborated here.

[0051] In summary, obtain the historical field quantity list, obtain the historical domain list set, based on the historical field quantity list R and the historical domain list T, obtain the correlation coefficient list, and perform normalization processing on U 1 to U b to obtain the quantity weight list, calculate the weight through historical data, and more accurately obtain the weight relationship between the field name quantity and the domain.

[0052] Further, in S360, determining the target domain corresponding to the target field name based on the second use label weight and the domain value list corresponding to the second use label further includes:

[0053] S361, obtain the domain integral value list J i ={J i1 , J i2 , …, J ig , …, J iz}, the domain integral value J ig corresponding to the g-th preset domain = ∑ n j=1 ∑ s r=1 (F ijr ×E ijrg ), where F ijr is the second use label weight corresponding to D ijr , D ijr is the r-th second use label corresponding to C ij , C ij is the j-th combined field name corresponding to N i , the value range of j is from 1 to n, n is the number of combined field names corresponding to N i , the value range of r is from 1 to s, s is the number of second use labels, and the value range of g is from 1 to z, z is the number of preset domains.

[0054] S362, sort J i1 to J iz from largest to smallest, and obtain the top q preset domains as the target domain corresponding to the target field name N i . Specifically, sort J i1 to J iz from largest to smallest, and obtain the sorted list K i ={K i1 , K i2 , …, K ig , …, K iz}, and obtain the top q as the target domain TKi = {TK i1 , TK i2 , …, TK ix , …, TK iq}, TK ix is the x-th target field corresponding to N i , where the value range of x is from 1 to q, q is the number of target fields. Optionally, q = 1.

[0055] Furthermore, S360 further includes:

[0056] S3601, obtaining the target value corresponding to the first usage label, where the target value is the weighted sum value of the field probability value and the initial probability value.

[0057] S3602, obtaining the maximum target value and using the first usage label corresponding to the maximum target value as the target label.

[0058] In an embodiment of the present invention, in S3602, when calculating the target value, the weight of the field probability value is equal to the weight of the initial probability value.

[0059] In summary, obtain the target value corresponding to the first usage label, obtain the maximum target value, and use the first usage label corresponding to the maximum target value as the target label.

[0060] Embodiment 2

[0061] After obtaining the target field and the target label corresponding to the target field, using the target field and the target label corresponding to the target field as the target data list, Embodiment 2 of the present invention provides a path determination method based on a large model, and the method includes the following steps:

[0062] S001, obtaining the user question, inputting the user question into the target large model, and obtaining the implementation path list AA = {AA 1 , AA 2 , …, AA e , …, AA f} and the implementation probability list AB = {AB 1 , AB 2 , …, AB e , …, AB f}, where AB e is the implementation probability value of AA e , and the e-th implementation path AA e = {AA e1 , AA e2 , …, AA eh , …, AA eu}, AA ehis the h-th implementation field of the e-th actual path and the implementation label corresponding to the h-th implementation field, where the value range of h is from 1 to u, u is the number of implementation fields, and the value range of e is from 1 to f, f is the number of actual paths; among them, AB e is greater than AB e+1 .

[0063] S002, traverse AA, for AAe, use AA e1 to AA eu to match with the target data list, where the target data list includes several target fields and the target label corresponding to each target field. Specifically, traverse AA, in the order from AA 1 to AA f to match with the target data list.

[0064] Specifically, use the implementation field and the implementation label corresponding to the implementation field and the target field and the target label corresponding to the target field to match, and obtain that the similarity between the implementation field and the target field is greater than the preset similarity threshold, and the similarity between the implementation label corresponding to the implementation field and the target label corresponding to the target field is greater than the preset similarity threshold, then it is considered that the implementation field and the implementation label corresponding to the implementation field match successfully.

[0065] S003, if AA e1 to AA eu all match successfully with the target data list, take AA e as the final path. And use the final path to solve the user's problem.

[0066] S004, if the matching fails, obtain the list to be matched AC = {AC 1 , AC 2 , …, AC e , …, AC f}, the data to be matched AC e corresponding to AA e = {AC e1 , AC e2 , …, AC ev , …, AC ew}, AC ev is the v-th implementation field that fails to match in AA e and the implementation label corresponding to the v-th implementation field that fails to match, where the value range of v is from 1 to w, and the number of implementation fields that fail to match w ≤ u.

[0067] S005, for AC ev , obtain the relevant field AF ev of AC ev , and obtain AC ev to AF evThe target weighted value AG ev , where AF ev and AF ev The matching results of the corresponding relevant tags and the target data list are successful; the target weighted value AG ev is the sum of the preset weights of the nodes passed by the implementation path from AC ev to AF ev . Among them, C ev to AF ev The preset weights of the nodes passed by the implementation path can be determined according to the actual situation. In an embodiment of the present invention, C ev to AF ev The preset weights of the nodes passed by the implementation path increase step by step, that is, C ev to AF ev The preset weight of the α-th node passed by the implementation path is greater than that of C ev to AF ev The preset weight of the (α + 1)-th node passed by the implementation path.

[0068] Furthermore, S005 further includes: if the relevant field AF ev of AC ev is obtained fails, AA e is deleted from the implementation path list AA.

[0069] S006, determining the final path based on AG ev and AB e .

[0070] In summary, obtain the user's question, input the user's question into the target large model, obtain the implementation path list and the implementation probability list corresponding to the implementation path list, traverse AA, for AA e , use AA e1 to AA eu to match with the target data list. If AA e1 to AA eu all match successfully with the target data list, AA e is used as the final path. If the match fails, obtain the list to be matched AC, for AC ev , obtain the relevant field AF ev of AC ev , and obtain the target weighted value AG ev from AC ev to AF ev , determine the final path based on AG ev and AB e . The present invention obtains the implementation path of the user's question through the target large model, and determines the final path based on the implementation probability and the target weighted value of the relevant field, improving the efficiency of obtaining the final path of the user's question.

[0071] Furthermore, S006 also includes:

[0072] S061, obtaining the target sum value AG e0 = ∑ w v=1 AG ev , and based on the target sum value AG e0 obtaining the final value AH e = β 1 × AG e0 + β 2 × AB e , where β 1 is the first probability factor, and β 2 is the second probability factor. In an embodiment of the present invention, β 1 = β 2 .

[0073] S062, obtaining AH 0 = max{AH 1 , AH 2 , …, AH e , …, AH f}, obtaining the implementation path AA 0 corresponding to AH 0 .

[0074] S063, taking the relevant paths from AA 0 to AF ev as the final path. ev In an embodiment of the present invention, the A* algorithm is used to obtain the relevant path from AC

[0075] to AF ev . In another embodiment of the present invention, the ant colony algorithm is used to obtain the relevant path from AC ev to AF ev . ev

[0076] In summary, obtain the target sum value, obtain AH 0 and obtain the implementation path AA 0 corresponding to AH 0 , and take the relevant paths from AA 0 to AF ev as the final path. ev

[0077] ​​An embodiment of the present invention also provides a non-transitory computer-readable storage medium, which can be disposed in an electronic device to store at least one instruction or at least one program related to a method in the method embodiment. The at least one instruction or the at least one program is loaded and executed by the processor to implement the method provided in the above embodiment.

[0078] An embodiment of the present invention also provides an electronic device, including a processor and the aforementioned non-transitory computer-readable storage medium.

[0079] Although some specific embodiments of the present invention have been described in detail by way of examples, those skilled in the art should understand that the above examples are for illustrative purposes only and not for limiting the scope of the present invention. Those skilled in the art should also understand that various modifications can be made to the embodiments without departing from the scope and spirit of the present invention.

Claims

1. A method for determining a data label, characterized in that: The method comprises the following steps: S100, obtaining a data list of a target user, wherein the data list of the target user includes several fields; S200, obtaining a single field name corresponding to the field; S300, marking any single field name as a target field name, and obtaining a target label corresponding to the target field name; Wherein, S300 includes the following steps: S310, inputting the target field name into the target macro model, obtaining a first usage label list corresponding to the target field name and an initial probability value corresponding to each first usage label of the target field name, wherein the first usage label list includes a plurality of first usage labels; S320, combining the target field name with several other single field names except the target field name, obtaining a list of combined field names, and inputting the combined field names into the target large model to obtain a list of second purpose labels corresponding to the combined field names; the list of combined field names includes several combined field names, and the second purpose label list includes several second purpose labels; S330, obtaining a quantity weight list, and based on the quantity weight list and the combined field name, obtaining a weight corresponding to the second purpose label and marking it as a second purpose label weight; wherein the quantity weight list includes a plurality of quantity weights, and the quantity weight is a weight between the quantity of field names in the combined field name and the field; S340, inputting the second usage tag into the target large model, obtaining a domain value list corresponding to the second usage tag, wherein the domain value list includes probability values ​​of the second usage tag in a plurality of preset domains; S350, determining a target field corresponding to the target field name based on the second usage tag weight and the field value list corresponding to the second usage tag; S360, obtaining a domain probability value of the first usage tag in the target domain, and determining a target tag corresponding to the target field name based on the domain probability value of the first usage tag in the target domain and an initial probability value corresponding to the first usage tag.

2. The data tag determination method according to claim 1, characterized in that: The S200 also includes: S210, obtaining a preset field feature list, wherein the preset field feature list includes a plurality of preset field features; S220, obtaining a corresponding preset field characteristic value list, wherein the preset field characteristic value list includes preset field characteristic values ​​corresponding to a plurality of preset field characteristics; S230: Determine a single field name of the field based on a preset field characteristic value list.

3. The data tag determination method according to claim 1, characterized in that: The target large model is a target large language model.

4. The data tag determination method according to claim 1, characterized in that: S320 includes: S321, obtain a single field name list N = {N1, N2, ..., N i ,…,N m }, and get all the fields except the target field name N i List of other field names P = {P1, P2, ..., P p ,…,P n-1 }, P p is the pth other field name, and the value range of p is 1 to n-1; S322, take any other field names from the other field name list P, and compare them with the target field name N i Combine and obtain the combined field name.

5. The data tag determination method according to claim 1, characterized in that: In S330, the quantity weight list is obtained through the following steps: S331, obtain the history field quantity list R = {R1, R2, ..., R a , …, R b }, R a = {R a1 , R a2 , …, R ac , …, R ad }, R ac is the cth sample of the ath historical field quantity, where c ranges from 1 to d, d is the number of samples, and a ranges from 1 to b, b is the number of historical fields; where the ath historical field quantity includes a field names; S332, obtain the historical domain list set T = {T1, T2, ..., T a ,…,T b }, the ath history field list T a= {T a1 , T a2 ,…,T ac ,…,T ad }, T ac YesR ac Corresponding fields; S333, based on the historical field quantity list R and the historical domain list T, obtain the correlation coefficient list U = {U1, U2, ..., U a , …, U b }, where the ath correlation coefficient U a =[∑ d c=1 [(R ac -ER a )×(T ac -ET a )]] / [(∑ d c=1 ((R ac -ER a ) 2 )) 1 / 2 ×∑ d c=1 ((T ac -ET a ) 2 )) 1 / 2 ], ER a =(∑ d c=1 R ac ) / d,ET a =(∑ d c=1 T ac ) / d; S334, for U1 to U b Perform normalization processing to obtain a list of quantity weights.

6. The data tag determination method according to claim 4, characterized in that: In S360, determining the target domain corresponding to the target field name based on the second usage tag weight and the domain value list corresponding to the second usage tag further includes: S361, obtain domain score value list J i ={J i1 , J i2 , …, J ig , …, J iz }, the domain integral value J corresponding to the g-th preset domain ig =∑ n j=1 ∑ s r=1 (F ijr ×E ijrg ), where F ijr Yes D ijr The corresponding second purpose label weight, D ijr It is C ij The corresponding r-th second purpose label, C ij YesN i The corresponding j-th combined field name, j ranges from 1 to n, and n is N i The number of corresponding combined field names, r ranges from 1 to s, s is the number of second purpose tags, g ranges from 1 to z, z is the number of preset fields; S362, for J i1 To J iz Sort by large to small, and get the first q preset fields as the target field name N i The corresponding target area.

7. The data tag determination method according to claim 1, characterized in that: S360 also includes: S3601, obtaining a target value corresponding to the first usage tag, where the target value is a weighted sum of a domain probability value and an initial probability value; S3602, obtaining a maximum target value, and using a first usage tag corresponding to the maximum target value as a target tag.

8. The data tag determination method according to claim 7, characterized in that: In S3602, the weight of the domain probability value is equal to the weight of the initial probability value.

9. A non-transitory computer-readable storage medium, wherein at least one instruction or at least one program is stored in the storage medium, characterized in that: The at least one instruction or the at least one program is loaded and executed by the processor to implement the data tag determination method as described in any one of claims 1-8.

10. An electronic device, characterized in that: The invention comprises a processor and the non-transitory computer-readable storage medium as claimed in claim 9.

Citation Information

Patent Citations

  • Label establishing method and device, electronic equipment and medium

    CN111553148A

  • Label generation method and device, electronic equipment and computer readable storage medium

    CN112214556A

  • Cross-business field field matching method and device and storage medium

    CN115827645A

  • Field classification method, electronic equipment, storage medium and computer program product

    CN117763150A

  • Image category prediction processing method for field generalization based on feature decoupling and category center matching

    CN118506047A