Method, apparatus, device, medium and product for determining classification information

CN117573863BActive Publication Date: 2026-09-25TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210935490.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-04
Publication Date
2026-09-25
Estimated Expiration
2042-08-04

AI Technical Summary

Technical Problem

基于有限的检索结果进行商户分类预测,严重影响了预测准确性

Benefits of technology

[0017]本申请实施例提供的分类信息的确定方法、装置、设备、介质和产品,可以获取互联网平台用户对商户(例如,本申请实施例所述的待预测对象)的评价信息,利用分类模型对评价信息进行处理,获得商户的分类信息。具体地,对待预测对象的评价信息进行特征提取,获得评价信息的语义表达特征,还可以对语义表达特征进行权重分配,获得中间特征。最后将待预测对象的所有评价信息对应的中间特征进行融合,并基于融合结果确定待预测对象的分类信息。例如,待预测对象的预测类型,和/或,待预测对象在该预测类型下的预测属性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117573863B_ABST
    Figure CN117573863B_ABST
Patent Text Reader

Abstract

The application discloses a method and device for determining classification information, equipment, medium and product, and relates to the technical field of big data. The method comprises the following steps: obtaining at least one evaluation information of a to-be-predicted object, performing semantic extraction on each evaluation information based on a classification model to obtain semantic expression features of each evaluation information; performing weight distribution on the semantic expression features based on the classification model to obtain intermediate features of each evaluation information; performing fusion processing on all intermediate features corresponding to the at least one evaluation information based on the classification model, and determining classification information of the to-be-predicted object according to a fusion result; the classification information comprises at least one of a prediction type and a prediction attribute under the prediction type, direct prediction of a merchant type is realized, and the accuracy of merchant classification prediction is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure generally relates to the field of big data technology, specifically to the field of cybersecurity technology, and in particular to a method, apparatus, device, medium, and product for determining classified information. Background Technology

[0002] With the rapid development of internet technology, merchants can join internet platforms to receive orders online. These platforms also provide users with various feedback channels, allowing them to obtain user feedback on merchants online (e.g., reviews, complaints).

[0003] Currently, massive amounts of review information can be retrieved using keywords (e.g., refund, negative reviews). Furthermore, analyzing review information containing keywords can determine the type of individual reviews, and even predict the type of merchant based on the type of review information.

[0004] The sheer volume and diversity of merchant reviews, coupled with the limited number of preset keywords, means that the amount of review information retrieved through keyword searches is also quite limited. Using these limited search results for merchant classification predictions severely impacts prediction accuracy. Summary of the Invention

[0005] In view of the above-mentioned defects or deficiencies in the prior art, it is desirable to provide a method, apparatus, device, medium and product for determining classification information, which can determine the classification information of merchants based on the merchants' evaluation information, realize the direct prediction of merchant types and improve the accuracy of merchant classification.

[0006] Firstly, a method for determining classification information is provided, including:

[0007] Obtain at least one evaluation information of the object to be predicted, and extract the semantics of each evaluation information based on the classification model to obtain the semantic expression features of each evaluation information.

[0008] Based on the classification model, the semantic expression features are weighted to obtain the intermediate features of each evaluation information.

[0009] Based on the classification model, all intermediate features corresponding to at least one evaluation information are fused, and the classification information of the object to be predicted is determined according to the fusion result; the classification information includes at least one of the prediction type and the prediction attribute under the prediction type.

[0010] Secondly, a classification information determination device is provided, comprising:

[0011] The semantic extraction unit is used to obtain at least one evaluation information of the object to be predicted, and performs semantic extraction on each evaluation information based on the classification model to obtain the semantic expression features of each evaluation information.

[0012] The weight allocation unit is used to allocate weights to semantic expression features based on the classification model to obtain intermediate features for each evaluation information.

[0013] The determination unit is used to perform fusion processing on all intermediate features corresponding to at least one evaluation information based on the classification model, and determine the classification information of the object to be predicted based on the fusion result; the classification information includes at least one of the prediction type and prediction attributes under the prediction type.

[0014] Thirdly, embodiments of this application provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method described in embodiments of this application.

[0015] Fourthly, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in embodiments of this application.

[0016] Fifthly, embodiments of this application provide a computer program product including instructions that, when executed, cause the method described in embodiments of this application to be performed.

[0017] The classification information determination method, apparatus, device, medium, and product provided in this application can obtain evaluation information of merchants (e.g., the object to be predicted described in this application) from internet platform users, process the evaluation information using a classification model, and obtain the merchant's classification information. Specifically, feature extraction is performed on the evaluation information of the object to be predicted to obtain semantic expression features of the evaluation information. Weights can also be assigned to the semantic expression features to obtain intermediate features. Finally, the intermediate features corresponding to all the evaluation information of the object to be predicted are fused, and the classification information of the object to be predicted is determined based on the fusion result. For example, the prediction type of the object to be predicted, and / or the prediction attribute of the object to be predicted under that prediction type.

[0018] It is evident that this application can directly predict merchant types based on merchant review information. Compared with existing technologies that rely on limited keywords to determine the type of review information and then determine the merchant type based on the type of review information, this application is not limited by the keyword hit rate of review information. It directly determines the merchant classification information based on the semantics of review information, thereby improving the accuracy of merchant classification. Attached Figure Description

[0019] Other features, objects, and advantages of this application will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:

[0020] Figure 1This is a schematic diagram of the implementation environment of an embodiment of this application;

[0021] Figure 2 A flowchart illustrating the method for determining classification information provided in an embodiment of this application;

[0022] Figure 3 A flowchart illustrating the training method for the classification model provided in this application embodiment;

[0023] Figure 4 This is a schematic diagram of the structure of the classification model provided in the embodiments of this application;

[0024] Figure 5 Another structural schematic diagram of the classification model provided in the embodiments of this application;

[0025] Figure 6 A classification diagram of the training phase provided for an embodiment of this application;

[0026] Figure 7 A schematic diagram illustrating the classification of application stages provided in this application embodiment;

[0027] Figure 8 This is a schematic diagram of label completion provided for an embodiment of this application;

[0028] Figure 9 This is another illustration of label completion provided for an embodiment of this application;

[0029] Figure 10 A schematic diagram illustrating threshold determination provided in an embodiment of this application;

[0030] Figure 11 This is a schematic diagram of data filtering provided for an embodiment of this application;

[0031] Figure 12 This is a schematic diagram of the structure of the classification information determination device provided in the embodiments of this application;

[0032] Figure 13 Another structural schematic diagram of the classification information determination device provided in the embodiments of this application;

[0033] Figure 14 A schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0034] The present application will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings.

[0035] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.

[0036] Figure 1 The internet system to which the embodiments of this application are adapted. Reference Figure 1 The system includes an internet platform 10, a merchant terminal 20, and a user terminal 30. The internet platform includes a front-end entry point 101 and a back-end server 102. The front-end entry point can be a webpage, an application (APP), etc.

[0037] Internet platform 10 can provide merchants with various online transaction models based on the internet. Merchants can register through merchant terminal 20 and accept orders and settle payments through the front-end entry 102 provided by the internet platform. Users can access the front-end entry 102 through user terminal 30 to place orders online and provide merchant reviews. In this embodiment, the evaluation information of merchants by users through the front-end entry can be user feedback information on merchants, such as user complaints about merchants.

[0038] The backend server 102 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.

[0039] Merchant terminal 20 and user terminal 30 may be devices including but not limited to personal computing, platform computers, smartphones, vehicle terminals, etc., and this application embodiment does not limit them.

[0040] Currently, internet platforms can identify registered merchants to take appropriate action against those exhibiting irregularities. Existing technology primarily relies on keywords to retrieve review information, further categorizes the retrieved reviews based on those keywords, and finally predicts the merchant type based on the categorization results. Clearly, existing technology cannot directly predict merchant type; it relies on the categorization results of review information. Furthermore, since the review information retrieved via keywords is relatively limited, categorizing merchants based on these limited results severely impacts the accuracy of the classification.

[0041] Based on this, this application proposes a method for determining classification information, which can determine the classification information of merchants based on their rating information, enabling direct prediction of merchant types and improving the accuracy of merchant classification. This method can be applied to... Figure 1 The backend server 10 is shown. (For example...) Figure 2 As shown, the method includes the following steps:

[0042] 201. Obtain at least one evaluation information of the object to be predicted, and extract the semantics of each evaluation information based on the classification model to obtain the semantic expression features of each evaluation information;

[0043] The objects to be predicted are those with whom users conduct online transactions based on the internet platform provided by the backend server 10. User feedback information regarding these objects can be obtained through the internet platform, such as user complaints and opinions. After obtaining this feedback information, it can be input into a classification model to categorize the merchants.

[0044] In one possible implementation, the classification network is a neural network model used to predict merchant types. Its input is the rating information of the object to be predicted (e.g., a merchant on an internet platform), and its output is the classification information of the object. Optionally, the classification model can obtain the predicted type of an object, the predicted attribute of the object under that predicted type, and the predicted rating type of the object's rating information. During the training process of the classification model, iterative training can be performed based on the losses between the predicted attributes, the predicted rating types of the rating information, and the corresponding true labels, allowing the model to gradually learn the ability to predict the type information of an object based on its rating information.

[0045] In one possible implementation, the classification model includes a semantic extraction network. After the evaluation information of the object to be predicted is input into the classification model, the semantic extraction network first extracts semantic features from the evaluation information so that type prediction can be made based on the extracted semantic expression features.

[0046] Specifically, different networks can be used for semantic extraction to obtain semantic expression features for evaluation information of different data types. The data type of the evaluation information of the object to be predicted can be image or text. For example, when the evaluation information contains an image, the convolutional network in the semantic extraction network can perform convolution processing on the image input to the classification model to extract image features. The semantic extraction network can also perform semantic extraction based on image features to obtain semantic expression features. Alternatively, the text recognition network in the semantic extraction network can perform semantic recognition on the text input to the model. For example, the evaluation information can be segmented into words, and the segmentation results can be recognized to obtain the semantic expression features of the evaluation information. The aforementioned text recognition network can be the first six layers of a BERT (bidirectional encoder representation from transformers) model, where the evaluation information is processed by the first six layers of the BERT model to obtain a 768-dimensional CLS vector.

[0047] It should be noted that after semantic extraction of evaluation information of different data types using different networks, the semantic expression features of different data types are fused. For example, if an evaluation message includes an image and text, the semantic expression features extracted based on the image feature map are fused with the semantic expression features extracted based on the text to obtain the semantic expression features of the evaluation message.

[0048] 202. Based on the classification model, weights are assigned to semantic expression features to obtain intermediate features for each evaluation information.

[0049] It should be noted that the classification model also includes a weight allocation network. After the semantic extraction network outputs the semantic expression features of the evaluation information, the semantic expression features can be input into the weight allocation network for weight allocation processing to obtain the intermediate features output by the weight allocation network.

[0050] In one possible implementation, weight allocation involves determining intermediate features based on semantic expression features and the weight coefficients of the corresponding evaluation information, or by performing calculations on the semantic expression features based on the weight coefficients. The weight coefficients of the evaluation information are positively correlated with the importance of the evaluation information; the more important the evaluation information, the higher its weight coefficient.

[0051] 203. Based on the classification model, all intermediate features corresponding to at least one evaluation information are fused, and the classification information of the object to be predicted is obtained according to the fusion result.

[0052] The classification information of the object to be predicted includes the prediction type and the prediction attributes of the object under that prediction type. The prediction type characterizes whether the object is abnormal; abnormality can indicate the presence of risk. For example, the business activities of the object based on an internet platform may be risky or abnormal. For instance, the prediction type could be "normal," "high risk," "low risk," "abnormal," or "risk exists." The prediction attributes characterize the specific type of abnormality of the object under the prediction type. For example, when the prediction type of the object is "abnormal," the classification model can also output the specific type of abnormality, such as transaction disputes, no service, impersonation fraud, fraudulent order placement, lottery scams, false information fraud, pornography, gambling, and illegal fundraising.

[0053] In one possible implementation, the classification model further includes a classification network. After the weight allocation network outputs intermediate features of the evaluation information, these intermediate features can be input into the classification network. The classification network can fuse the intermediate features of all the evaluation information of the object to be predicted and perform type prediction based on the fusion result, outputting the classification information of the object to be predicted. The classification network includes a binary classification function and a multi-class classification function. Inputting the fusion result into the binary classification function yields the predicted classification of the object to be predicted, and inputting the fusion result into the multi-class classification function yields the predicted attribute of the object to be predicted. The binary classification function can be a sigmoid function, and the multi-class classification function can be a softmax function.

[0054] For example, the fusion processing of intermediate features can be an addition process. For instance, if the object to be predicted i corresponds to J evaluation information, and the J evaluation information corresponds to J intermediate features, the J intermediate features can be added together to obtain the fusion result.

[0055] In this application embodiment, the prediction type and specific implementation method of the prediction attribute of the object to be predicted are also provided:

[0056] First, the semantic expression features of each evaluation information corresponding to the object to be predicted are aggregated, and the aggregated features are input into the binary classification function of the classification network to obtain the prediction type of the object to be predicted.

[0057] Specifically, for each semantic expression feature of the object to be predicted, intermediate features are determined based on the weight coefficient of the semantic expression feature and the evaluation information corresponding to the semantic expression feature; all intermediate features corresponding to the object to be predicted can also be fused and the fusion result can be scalarized to obtain a first vector; finally, binary classification prediction is performed on the first vector based on the binary classification function to obtain the prediction type of the object to be predicted.

[0058] For example, the object to be predicted, i, corresponds to J evaluation information, namely X i1X i2 …X iJ Evaluation Information C ij (j is greater than or equal to 1 and less than or equal to J) After inputting into the feature semantic extraction network, the semantic expression characteristics of BERT(C) are obtained. ij ), BERT(C ij () can be a 768-dimensional vector. Furthermore, the semantic representation characteristics of the J evaluation information can be determined based on the weight coefficients of each evaluation information. BERT(C) ij Aggregation is performed, and the aggregated features are transformed into a scalar. This can be achieved using an MLP (Multi-Level Processing). m Perform scalar transformation. Finally, input the scalar into the sigmoid function to obtain the prediction type Y of the object to be predicted, i. i The following formula (1) gives the prediction type Y of the object to be predicted. i The expression for ':

[0059] Y i = sigmoid(MLP) m (∑ j attention ij *BERT(C ij ))) (1)

[0060] Among them, attention ij C is the j-th evaluation information of the object to be predicted, i. ij The weight coefficients of attention ij *BERT(C ij ) is the evaluation information C ij The corresponding intermediate features, after aggregating the intermediate features of the J evaluation information, are ∑ j attention ij *BERT(C ij The first vector corresponding to the object to be predicted, i, is an MLP. m (∑ j attention ij *BERT(C ij )).

[0061] The following describes the output of the sigmoid function, a binary classification function. Assume the predicted evaluation types can be represented by 0 and 1, where "0" represents "normal" and "1" represents "abnormal". The output of the sigmoid function can be a 2x1 vector, where each element corresponds to a type. The value of each element (score) represents the probability that the object to be predicted belongs to that type, and the sum of all scores equals 1. A higher score indicates a higher probability that the object to be predicted belongs to that type; therefore, the type corresponding to the highest score can be chosen as the predicted type for the object.

[0062] Suppose that the first vector of the object to be predicted is input into the sigmoid function, and the output is [0.33, 0.67]. This means that the probability of the predicted type of the object being "0" is 0.33, and the probability of the predicted type being "1" is 0.67. In this case, "0" can be taken as the prediction result of the binary classification function, that is, the predicted type of the object is "0", which means that the object is normal.

[0063] Second, the semantic expression features of each evaluation information corresponding to the object to be predicted are aggregated, and the aggregated features are input into the multi-classification function of the classification network to obtain the predicted attributes of the object to be predicted.

[0064] Specifically, for each semantic expression feature of the object to be predicted, intermediate features are determined based on the semantic expression features and the weight coefficients of the corresponding evaluation information; all intermediate features corresponding to the object to be predicted are fused, and the fusion result is subjected to dimensionality transformation to obtain a second vector; multi-classification prediction is performed on the second vector based on a multi-classification function to obtain the predicted attributes of the object to be predicted. The number of dimensions of the second vector is the same as the number of preset merchant types, which is the number of merchant types that can be predicted using the method (or classification model) provided in this application embodiment.

[0065] For example, the object to be predicted, i, corresponds to J evaluation information, namely X i1 X i2 …X iJ Evaluation Information C ij (j is greater than or equal to 1 and less than or equal to J) After inputting into the feature semantic extraction network, the semantic expression characteristics of BERT(C) are obtained. ij ), BERT(C ij () can be a 768-dimensional vector. Furthermore, the semantic representation characteristics of the J evaluation information can be determined based on the weight coefficients of each evaluation information. BERT(C) ij The aggregated features are then subjected to dimensionality transformation. This can be achieved using MLP (Multi-Level Processing). mtPerform scalar transformation. Finally, input the scalar into the softmax function to obtain the predicted attribute Z of the object to be predicted, i. i The following formula (2) gives the predicted attribute Z of the object to be predicted. i The expression for ':

[0066] Z i = softmax(MLP) mt (∑ j attention ij *BERT(C ij (2)

[0067] Among them, attention ij C is the j-th evaluation information of the object to be predicted, i. ij The weight coefficients of attention ij *BERT(C ij ) is the evaluation information C ij The corresponding intermediate features, after aggregating the intermediate features of the J evaluation information, are ∑ j attention ij *BERT(C ij The second vector corresponding to the object to be predicted, i, is an MLP. mt (∑ j attention ij *BERT(C ij )).

[0068] The following uses 3-class classification as an example to introduce the output of a multi-class classification function (softmax). The attributes that a multi-class classification function can predict are represented by 1, 2, and 3. Its output can be a 3*1 vector, where each element corresponds to an attribute, and the value of the element (score) represents the probability that the input sample belongs to the corresponding attribute. A higher score indicates a higher probability that the attribute corresponding to that score is present in the object to be predicted. Therefore, the attribute corresponding to the highest score can be selected as the predicted attribute for the object to be predicted.

[0069] Suppose that the second vector of the object to be predicted is input into the softmax function, and the output is [0.11, 0.34, 0.55]. This means that the probability of the predicted attribute of the object being classified as class 1 is 0.11, the probability of the predicted attribute being class 2 is 0.34, and the probability of the predicted attribute being class 3 is 0.55. In this case, the highest score, "class 3," can be taken as the prediction result of the multi-class function, meaning that the predicted attribute of the object is "class 3."

[0070] The method provided in this application embodiment can directly predict the merchant type based on the merchant's evaluation information. Compared with the existing technology that relies on limited keywords to determine the type of evaluation information and then determines the merchant type based on the type of evaluation information, it is not limited by the hit rate of evaluation information on keywords and can determine the merchant's classification information based on the evaluation information, thus improving the accuracy of merchant classification.

[0071] In another embodiment of this application, weights can be assigned to the corresponding semantic expression features based on the weight coefficients of the evaluation information. For example, the aforementioned weight assignment of semantic expression features based on a classification model to obtain intermediate features for each evaluation information includes:

[0072] For each piece of evaluation information, the weight coefficient of the evaluation information is determined based on the quantified value of the semantic expression features of the evaluation information; the quantified value of the semantic expression features is positively correlated with the weight coefficient of the evaluation information; and the intermediate features of the evaluation information are obtained by multiplying the semantic expression features and the weight coefficient of the evaluation information.

[0073] In another embodiment of this application, weighting coefficients for evaluation information (e.g., the attention mentioned above) are also provided. ij The specific implementation method of ). In one possible implementation method, for the J evaluation information corresponding to object i, the evaluation information C ij The weighting coefficient is related to the magnitude of the quantified value of the evaluation information; the larger the quantified value of the evaluation information, the larger the weighting coefficient. Among them, the evaluation information C... ij The quantized value is obtained after quantifying the semantic representation features of the evaluation information. For example, the evaluation information is input into BERT, MLP... at Evaluation information C was subsequently obtained. ij Semantic representation features of MLP at BERT(C ij ), and its corresponding weight coefficient attention ij The following formula (3) must be satisfied:

[0074]

[0075] The exp operation is used to determine the quantification value of the evaluation information, that is, to obtain the quantification value of the evaluation information by performing an exponential operation on the semantic expression features of the evaluation information.

[0076] This represents the sum of the quantified values ​​of all evaluation information for object i. In other words, the weight coefficient of a certain evaluation information for object i is related to the ratio of its quantified value to the sum of its quantified values.

[0077] This application also introduces an attention mechanism, which assigns different weight coefficients to different evaluation information, which helps to obtain more accurate classification results.

[0078] In another embodiment of this application, a training method for the above-described classification model is also provided.

[0079] For example, refer to Figure 3 The method includes the following steps:

[0080] 301. Obtain historical data from multiple reference objects;

[0081] The historical data of the reference object includes the reference object's type label, the reference object's attribute label, multiple evaluation information for each reference object, and the evaluation type label for each evaluation information.

[0082] In this embodiment, after a merchant registers on the internet platform, the backend server can obtain historical evaluations based on the internet platform, i.e., the historical data described in this embodiment. The server can also train a model based on this historical data to obtain a model capable of classifying merchants based on their evaluation information, i.e., the classification model described in this embodiment.

[0083] It should be noted that the reference objects are those with whom users conduct online transactions based on the internet platform provided by the backend server 10. User feedback information, such as evaluations and complaints, regarding merchants can be obtained through the internet platform. For example, the reference objects can be merchants registered on the internet platform. After users evaluate merchants through the internet platform, the platform can store the user's evaluation data. In step 301, historical data of multiple reference objects stored on the internet platform can be retrieved.

[0084] Additionally, the type label of a reference object is used to characterize whether the reference object is abnormal. For example, it could be a "high-risk" label, a "low-risk" label, or a "risk exists" label, an "abnormal" label, or a "normal" label. In specific implementations, the type label of a reference object can be related to the historical handling results of the reference object. For example, by obtaining the handling results of the reference object (merchant), filtering out the more severe handling results (such as restricting funding, closing payment permissions, etc.), the type label of the corresponding merchant is set to "high-risk," "risk exists," or "abnormal"; the type label of the remaining merchants is set to "low-risk" or "normal."

[0085] The attribute tags of a reference object are used to characterize the specific type of anomaly, such as transaction disputes, lack of service, impersonation fraud, fraudulent order placement, lottery scams, false information fraud, pornography, gambling, and illegal fundraising. The attribute tags of a reference object can be the result of manual review. For example, risk control experts may identify merchant ratings based on existing review standards and derive a review result, which can then serve as the merchant's attribute tag.

[0086] The evaluation information of a reference object can be the content entered by a user when evaluating a reference object (e.g., a merchant) on an internet platform. The evaluation information can be text or image elements. When the user enters text, the backend server can store the text as the merchant's historical data. When the user enters image elements, the backend server can store the image element itself, or it can store the evaluation information represented by the image element. For example, image element 1 represents "transaction dispute exists," and image element 2 represents "fraudulent use exists." For instance, if a user selects image element 2 to evaluate a merchant, the merchant's historical data can include the evaluation information "fraudulent use exists," and / or image element 2.

[0087] The rating type label of the rating information is used to characterize the specific type of the rating information. For example, the rating type label of the rating information can be the rating type selected by the user when entering the rating information through the front-end entry of the Internet platform.

[0088] In one possible implementation, the evaluation type labels of the evaluation information can also be filtered to remove labels that do not match the corresponding reference object. For example, if the evaluation type label of the evaluation information does not belong to the type label of the corresponding reference object, then the evaluation type label is filtered out.

[0089] 302. Based on the initial model, perform type prediction on historical data to obtain the first prediction loss, the second prediction loss, and the third prediction loss for each reference object;

[0090] The first prediction loss is the loss between the type label of the reference object and the predicted type; the second prediction loss is the loss between the attribute label of the reference object and the predicted attribute of the reference object; and the third prediction loss is the loss between the evaluation type label of the evaluation information of the reference object and the predicted evaluation type of the evaluation information. The predicted type is the type of the reference object output by the model, the predicted attribute is the attribute of the reference object output by the model, and the predicted evaluation type is the evaluation type of the evaluation information output by the model.

[0091] In one possible implementation, the specific implementation of type prediction based on the initial model for historical data includes: semantic extraction of each historical data to obtain the semantic expression features of each evaluation information, and further type prediction based on the semantic expression features to obtain the predicted type, predicted attribute, and predicted evaluation type of each reference object.

[0092] For example, refer to Figure 4 The initial model can include a semantic extraction network and a classification network. Historical data can be input into the semantic extraction network for semantic extraction to obtain semantic expression features of the evaluation information. Alternatively, the semantic expression features can be input into the classification network, where the classification function operates on the semantic expression features to obtain the prediction results for the reference object and the prediction results for the evaluation information of the reference object. The prediction results for the reference object include its classification information, such as the predicted type and predicted attribute. The predicted type characterizes whether the reference object is abnormal, and can be "risky," "normal," "high-risk," "low-risk," etc., or it can be the probability that the reference object belongs to a certain type. The predicted attribute of the reference object can be the specific type of abnormality, such as "fraudulent," "transaction dispute," etc. The prediction results for the evaluation information can include the predicted evaluation type, that is, the specific type of evaluation information predicted by the model.

[0093] In one possible implementation, different networks can be used for semantic extraction to obtain semantic expression features for different data types of historical evaluations. Type prediction is then performed on the semantic expression features of each data type. The data type of the historical evaluation can be image or text. For example, when historical data contains images (e.g., image elements as described above), the convolutional network in the semantic extraction network can perform convolution processing on the historical data input to the model to extract image features. The semantic extraction network can also perform semantic extraction based on image features to obtain semantic expression features. Further, the semantic expression features are processed by the classification function of the input classification network to obtain the corresponding prediction results. Alternatively, the text recognition network in the semantic extraction network can perform semantic recognition on the text (e.g., evaluation information) input to the model. For example, the evaluation information can be segmented into words, and the segmentation results can be recognized to obtain the semantic expression features of the evaluation information. The aforementioned text recognition network can be the first six layers of a BERT (bidirectional encoder representation from transformers) model, where the evaluation information is processed by the first six layers of the BERT model to obtain a 768-dimensional CLS vector.

[0094] Another possible implementation involves using different networks to extract semantics from historical data of different data types, then fusing the semantic features of these different data types. Type prediction is then performed based on the fused semantic features. For example, refer to... Figure 5 When historical data is input into the initial model (or classification model), the data in the historical data can first be split and processed, and the text can be input into the text recognition network. The text recognition network can extract the semantics of the text and obtain the first type of semantic expression features.

[0095] Furthermore, the aforementioned splitting process allows images from historical data to be input into an image recognition network. This network can then use convolutional networks to extract features from the input images, obtaining feature maps. Semantic extraction can also be performed on these feature maps to obtain second-type semantic expression features.

[0096] Finally, the first and second types of semantic expression features can be input into the classification network. The classification network can fuse the first and second types of semantic expression features, and then input the fused semantic expression features into the classification function of the classification network for calculation to obtain the prediction results for historical data.

[0097] Alternatively, optical character recognition (OCR) can be performed on the images in the historical data before splitting the data, and the recognized text can be input into a text recognition network to obtain the first type of semantic expression features.

[0098] In one possible implementation, the classification function can be a sigmoid function for binary classification prediction, or a softmax function for multi-class classification. The binary classification function can output the predicted type of the reference object; the multi-class classification function can output the predicted type of the reference object, or the predicted evaluation type of the evaluation information.

[0099] In one possible implementation, this application embodiment also provides specific implementation methods for three prediction results of the reference object (i.e., prediction type of the reference object, prediction attribute, and prediction evaluation type of evaluation information):

[0100] First, for a single reference object, the semantic expression features of each evaluation information corresponding to the reference object can be aggregated, and the aggregated features can be input into the binary classification function of the classification network to obtain the predicted type of the reference object.

[0101] Specifically, the type prediction based on semantic expression features, as described above, to obtain the predicted classification for each reference object includes:

[0102] For each semantic expression feature of the reference object, intermediate features are determined based on the weight coefficient of the semantic expression feature and the evaluation information corresponding to the semantic expression feature. Alternatively, all intermediate features corresponding to the reference object can be fused, and the fusion result can be scalarized to obtain a first vector. Finally, binary classification prediction is performed on the first vector based on a binary classification function to obtain the predicted type of the reference object.

[0103] For example, reference object i corresponds to J evaluation information, namely X i1 X i2 …X iJ Evaluation Information C ij (j is greater than or equal to 1 and less than or equal to J) After inputting into the feature semantic extraction network, the semantic expression characteristics of BERT(C) are obtained. ij ), BERT(C ij () can be a 768-dimensional vector. Furthermore, the semantic representation characteristics of the J evaluation information can be determined based on the weight coefficients of each evaluation information. BERT(C) ij Aggregation is performed, and the aggregated features are transformed into a scalar. This can be achieved using an MLP (Multi-Level Processing). m Perform scalar transformation. Finally, input the scalar into the sigmoid function to obtain the prediction type Y of the reference object i. i Prediction type Y i The expression for ' is given in formula (1) mentioned above and will not be repeated here.

[0104] The following describes the predicted type of the reference object based on the output of the binary classification function (sigmoid). The types predicted by the binary classification function are represented by 0 and 1, where "0" indicates a predicted type of "normal" and "1" indicates a predicted type of "abnormal". The output of the sigmoid function can be a 2*1 vector, where each element corresponds to a type. The value of each element (score) represents the probability that the reference object belongs to the type corresponding to that element, and the sum of all scores equals 1. A higher score indicates a higher probability that the reference object matches the type corresponding to that score; therefore, the type corresponding to the highest score can be selected as the predicted type of the reference object.

[0105] Suppose that after inputting the first vector of the reference object into the sigmoid function, the output is [0.33, 0.67]. This means that the probability of the reference object's predicted type being "0" is 0.33, and the probability of the reference object's predicted type being "1" is 0.67. In this case, "0" can be taken as the prediction result of the binary classification function, that is, the predicted type of the reference object is "0", which means the reference object is normal.

[0106] The second approach, for a single reference object, is to aggregate the semantic expression features of each evaluation information corresponding to the reference object, and then input the aggregated features into the multi-classification function of the classification network to obtain the predicted attributes of the reference object.

[0107] Specifically, the type prediction based on semantic expression features, as described above, to obtain the predicted attributes of each reference object, includes:

[0108] For each semantic expression feature of the reference object, intermediate features are determined based on the semantic expression features and the weight coefficients of the corresponding evaluation information. All intermediate features corresponding to the reference object are fused, and the fusion result is subjected to dimensionality transformation to obtain a second vector. Multi-class prediction is performed on the second vector based on a multi-classification function to obtain the predicted attributes of the reference object. The number of dimensions of the second vector is the same as the number of preset attributes, which is the number of attributes that can be predicted using the method (or classification model) provided in this application embodiment.

[0109] For example, reference object i corresponds to J evaluation information, namely X i1 X i2 …X iJ Evaluation Information C ij (j is greater than or equal to 1 and less than or equal to J) After inputting into the feature semantic extraction network, the semantic expression characteristics of BERT(C) are obtained. ij ), BERT(C ij () can be a 768-dimensional vector. Furthermore, the semantic representation characteristics of the J evaluation information can be determined based on the weight coefficients of each evaluation information. BERT(C) ij The aggregated features are then subjected to dimensionality transformation. This can be achieved using MLP (Multi-Level Processing). mt A scalar transformation is performed. Finally, the scalar is input into the softmax function to obtain the predicted attribute Z of the reference object i. i '. Predicted attribute Z of the reference object i The expression for ' is given in formula (2) mentioned above and will not be repeated here.

[0110] The following uses 3-class classification as an example to introduce the output of a multi-class classification function (softmax). The attributes that a multi-class classification function can predict are represented by 1, 2, and 3. Its output can be a 3*1 vector, where each element corresponds to an attribute, and the value (score) of the element represents the probability that the reference object possesses the attribute corresponding to that element. A higher score indicates a higher probability of the reference object possessing the attribute corresponding to that score; therefore, the attribute corresponding to the highest score can be chosen as the predicted attribute for the reference object.

[0111] Suppose that inputting the second vector of the reference object into the softmax function yields an output of [0.11, 0.34, 0.55]. This indicates that the probability of the reference object's predicted attribute being class 1 is 0.11, the probability of its predicted attribute being class 2 is 0.34, and the probability of its predicted attribute being class 3 is 0.55. In this case, the highest score, "class 3," can be taken as the prediction result of the multi-class classification function, meaning the predicted attribute of the reference object is "class 3."

[0112] The third method involves inputting the semantic representation features of a single evaluation message into the multi-classification function of a classification network to obtain the predicted evaluation type of the evaluation message.

[0113] Specifically, the aforementioned method of predicting the type of evaluation information based on semantic expression features to obtain the predicted evaluation type includes: performing dimensional transformation processing on the semantic expression features of the evaluation information to obtain a third vector; and performing multi-classification prediction on the third vector based on a multi-classification function to obtain the predicted evaluation type of the evaluation information. The number of dimensions of the third vector is the same as the number of preset evaluation types, which is the number of evaluation types that can be predicted using the method provided in this application embodiment.

[0114] For example, the evaluation information of reference object i can be denoted as C. ij Evaluation Information C ij After inputting into the feature semantic extraction network, the semantic representation features of BERT(C) are obtained. ij ), BERT(C ij () can be a 768-dimensional vector. It can also be adjusted based on the number of preset evaluation types for BERT(C). ij Perform dimensionality transformation to obtain the vector BERT'(C) ij ), BERT'(C ij The number of dimensions is the same as the number of preset evaluation types. Dimension transformation processing can be performed using MLP. C Processing. Further processing of BERT'(C ij The input is processed by the softmax function, and the output is the evaluation information C. ij The prediction type. The following formula (3) is the evaluation information C. ij Prediction type T' ij The expression:

[0115] T' ij =softmax(MLP) C (BERT(C ij )))(3)

[0116] The following uses 3-class classification as an example to introduce the predicted evaluation type output by a multi-class classification function (softmax). The evaluation types that a multi-class classification function can predict are represented by A, B, and C. Its output can be a 3*1 vector, where each element corresponds to an evaluation type, and the element's value (score) represents the probability that the evaluation information corresponds to that score. A higher score indicates a higher probability that the evaluation information matches the preset evaluation type corresponding to that score; therefore, the evaluation type corresponding to the highest score can be selected as the predicted evaluation type for the evaluation information.

[0117] Suppose that after inputting BERT'(Cij) into the softmax function, the output is [0.09, 0.24, 0.67]. This means that the probability of predicting the evaluation type of this evaluation information as class A is 0.09, the probability of predicting the evaluation type as class B is 0.24, and the probability of predicting the evaluation type as class C is 0.67. In this case, the highest score, "class C", can be taken as the prediction result of the multi-class function, that is, the predicted evaluation type of the evaluation information is "class C".

[0118] In this application embodiment, based on the three different implementation methods described above, a detailed description of the network structure of the classification model is provided. For example, refer to... Figure 6 In addition to the semantic extraction networks and classification networks mentioned earlier, classification models can also include MLP networks and weight allocation networks. During the training phase, the input to the classification model includes evaluation information from multiple reference objects, evaluation type labels for the evaluation information, type labels for the reference objects, and attribute labels for the reference objects. Within the classification model, the evaluation information is first input into the semantic extraction network (e.g., the first six layers of the BERT model) to obtain the semantic representation features of the evaluation information.

[0119] Furthermore, semantic representation features can be input into an MLP network for processing, and intermediate features output by the MLP network can be input into a multi-classification function of a classification network to obtain the predicted evaluation type of the evaluation information.

[0120] Semantic expression features can also be input into a weighted distribution network. Based on the weight coefficients of the evaluation information, multiple semantic expression features corresponding to the same reference object (i.e., semantic expression features of multiple evaluation information of the reference object) are weighted and aggregated. The aggregated features are then input into an MLP network for processing. Finally, the output of the MLP network is input into the binary classification function of the classification network to obtain the predicted type of the reference object. Alternatively, the output of the MLP network can be input into the multi-classification function of the classification network to obtain the predicted attribute of the reference object.

[0121] refer to Figure 7After the model training is completed, during the inference and prediction process based on the trained model, the evaluation information of the object to be predicted is input into the semantic extraction network (e.g., the first six layers of the BERT model) to obtain the semantic expression features of the evaluation information.

[0122] Semantic expression features can also be input into a weighted distribution network. Based on the weight coefficients of the evaluation information, multiple semantic expression features corresponding to the same reference object (i.e., semantic expression features of multiple evaluation information of the reference object) are weighted and aggregated. The aggregated features are then input into an MLP network for processing. Finally, the output of the MLP network is input into the binary classification function of the classification network to obtain the predicted type of the object to be predicted. Alternatively, the output of the MLP network can be input into the multi-class classification function of the classification network to obtain the predicted attribute of the object to be predicted.

[0123] In one possible implementation, after obtaining the model prediction results, the model's prediction loss can be determined based on the model prediction results and labels. For example, a first prediction loss is determined based on a first loss function to compare the predicted type and type label of the reference object; a second prediction loss is determined based on a second loss function to compare the predicted attributes and attribute labels of the reference object; and a third prediction loss is determined based on a third loss function to compare the evaluation type labels of historical evaluation information with the predicted evaluation types of historical evaluation information.

[0124] For example, the type label of reference object i is Y. i The corresponding prediction type is Y. i ', then the first prediction loss corresponding to reference object i. i According to the first loss function -Y i log Y i '-(1-Y i log(1-Y) i ') Obtain; the attribute label of the reference object i is Z i The corresponding predicted attribute is Z. i ', the second prediction loss corresponding to reference object i i According to the second loss function -Z i log Z i '-(1-Z i log(1-Z) i ') Obtain; the j-th evaluation information C of reference object i ij The rating type label is T ij The corresponding prediction and evaluation type is T' ij Evaluation information C ij The corresponding third prediction loss ij It is based on the third loss function crossentropy(T' ij T ijThe value is obtained from crossentropy, where crossentropy is the cross-entropy operation.

[0125] 303. Iteratively train the initial model using the first prediction loss, the second prediction loss, and the third prediction loss to obtain a classification model.

[0126] It should be noted that before the model is trained, the accuracy of the initial model's prediction results for merchants is insufficient. It is necessary to iteratively train the model based on the loss between the initial model's prediction results for training samples and the sample labels. When the loss between the prediction results and the sample labels converges, the model gradually learns the ability to predict merchant types based on evaluation information.

[0127] For example, training samples can be historical data of a reference object. The labels of the training samples include the type label of the reference object, the attribute label of the reference object, and the evaluation type label of the evaluation information. By inputting the training samples into the initial model, three prediction results can be obtained from the model output: the predicted type of the reference object, the predicted attribute of the reference object, and the predicted evaluation type of the evaluation information.

[0128] Furthermore, the model loss function can be determined based on the first prediction loss, the second prediction loss, and the third prediction loss, and the initial model can be iteratively trained based on this loss function.

[0129] In one possible implementation, the first prediction losses corresponding to multiple reference objects can be aggregated to obtain the first component of the model loss function. For example, the type label of reference object i is Y. i The corresponding prediction type is Y. i ', the first prediction loss of reference object i i =-Y i log Y i '-(1-Y i log(1-Y) i '), sum the first predicted losses of all merchants to obtain Where I is the total number of reference objects, and riskloss is the first component mentioned above.

[0130] The second prediction losses corresponding to multiple reference objects can be aggregated to obtain the second component of the model loss function. The attribute label of reference object i is Z. i The corresponding predicted attribute is Z. i ', The second prediction loss of reference object i and the first prediction loss of reference object i i =-Z i log Z i '-(1-Z i log(1-Z)i '), sum the first predicted losses of all merchants to obtain Where I is the total number of reference objects, and riskTypeloss is the second component mentioned above.

[0131] Furthermore, all third-prediction losses corresponding to multiple reference objects can be aggregated to obtain the third component of the model loss function. For example, the j-th evaluation information C of reference object i... ij The rating type label is T ij The corresponding prediction type is T' ij Evaluation information C ij The third prediction loss ij =crossentropy(T' ij T ij The loss can be calculated from the J evaluation information of reference object i. ij Aggregation is performed to obtain the evaluation information loss of the reference object. Alternatively, the evaluation information losses of I reference objects can be aggregated to obtain the aforementioned third component. That is, the third component...

[0132] Understandably, the model loss function for training a classification model is LOSS = α * complaint loss + β * risk loss + γ * riskType loss. Here, α, β, and γ are hyperparameters of the model and can be dynamically adjusted based on training progress. Initially, α, β, and γ can be arbitrarily set to large values. During training, the hyperparameters of the loss components can be adjusted based on their convergence. For example, assuming complaint loss converges first, α can be reduced, and β and γ can be set to larger values. Subsequent training can focus on training the other two tasks until all loss components converge.

[0133] The method provided in this application embodiment can use historical data of multiple reference objects as training samples, and the type labels, attribute labels, and evaluation type labels of the evaluation information of the reference objects as labels for the training samples. Semantic extraction can also be performed on the historical data to obtain the semantic expression features of each evaluation information. Based on the semantic expression features of the evaluation information, type prediction is performed to obtain the predicted type, predicted attribute, and predicted evaluation type of the evaluation information of the reference object. Furthermore, the initial model can be iteratively trained based on the loss between the prediction results and the corresponding labels, so that the model gradually acquires the ability to predict types based on evaluation information, thus obtaining a classification model. This classification model can determine the type information of a merchant based on its evaluation information, such as the predicted type and predicted attribute of the merchant, achieving direct prediction of the merchant type. Compared with existing technologies that rely on limited keywords to predict evaluation information to determine the merchant type, this improves the accuracy of merchant classification prediction.

[0134] Furthermore, the lack of labels for most merchants can also affect model performance to some extent. However, this application's embodiment utilizes not only the prediction loss of the merchant dimension (i.e., the loss between the predicted type of the merchant and the type label of the merchant, and the loss between the predicted attribute of the merchant and the attribute label) during training, but also the prediction loss of the evaluation information dimension (i.e., the loss between the evaluation type label and the predicted type). Since there is a correlation between the evaluation type label of the merchant evaluation information and the merchant's own type label, this application can improve the model's ability to predict merchant types by leveraging the model's prediction loss of evaluation information, and can also improve the accuracy of merchant classification prediction to some extent.

[0135] In another embodiment of this application, semi-supervised training can be used to predict the classification model. That is, the training samples include both labeled and unlabeled samples. The labeled samples can be the reference objects described in the embodiments of this application, which have type labels. Alternatively, objects lacking type labels can be referred to as target objects. The embodiments of this application can also perform label completion processing on objects lacking type labels to determine type labels for target objects, thus constructing labeled training samples.

[0136] In one possible implementation, the target object can be padded with a first label based on its probability quantification value to obtain its type label. The probability quantification value represents the likelihood that the target object matches a preset type. The preset type is the merchant type that the classification model can predict, such as "abnormal" or "normal".

[0137] In practice, after inputting the evaluation information of the target object into the classification model, multi-class prediction can be performed based on the semantic expression features of the evaluation information to obtain the predicted type of the target object. Alternatively, the probability of the target object hitting a preset type before normalization can be used as the probability quantification value of the target object.

[0138] For example, if the probability quantification value of the target object is greater than the probability threshold corresponding to the preset type, then the preset type is used as the type label of the target object. Figure 8 For example, suppose the classification model provided in this application can predict preset types (e.g., merchant types) as y1, y2, and y3. Semantic extraction is performed on each evaluation information of the target object i, and the semantic expression features are aggregated to obtain the MLP. mt (∑ j attention ij *BERT(C ij MLP can be used mt (∑ j attention ij *BERT(C ij Input the softmax function to obtain a three-dimensional vector [p1, p2, p3]. Here, p1 is the probability that target object i hits type y1, p2 is the probability that target object i hits type y2, and p3 is the probability that target object i hits type y3.

[0139] Furthermore, determine the value of p1 before normalization, logit. mt1 If logit mt1 If the probability threshold F1 corresponding to type y1 is greater than the target object, then the target object is labeled according to type y1, that is, the type label of the target object includes "type 1".

[0140] Similarly, determine the value of p2 before normalization, logit. mt2 If logit mt2 If the probability threshold F2 corresponding to type y2 is greater than the target object, then the target object is labeled according to type y2, that is, the type label of the target object includes "type y2".

[0141] It is also possible to determine the value of p3 before normalization, logit. mt3 If logit mt3 If the probability threshold F3 corresponding to type y3 is greater than the target object, then the target object is tagged according to type y3, that is, the type label of the target object includes "type y3".

[0142] It should be noted that if a target object is tagged with two class targets after the first tag completion process described above, the predefined type corresponding to the larger probability quantification value of the target object will be taken as the type tag of the target object. For example, if the target object satisfies logit... mt2 >T2 and logit mt3 >T3, if logit mt2 logit mt3 If so, then type y2 is selected as the type label for the target object.

[0143] The method provided in this application embodiment can perform semi-supervised training. When there are insufficient merchant labels, the missing labels can be filled in to generate labeled training samples, thereby greatly enriching the training samples of the model. Training based on a richer model can also improve the model performance.

[0144] In another embodiment of this application, the target object may still lack a type label after the first label completion process described above. For example, if the probability quantification value of the target object does not exceed the probability threshold corresponding to any preset type, then the first label completion process cannot label the target object. Optionally, a second labeling process can be performed on the target object based on the probability quantification value of the evaluation information to obtain the type label of the target object. The probability quantification value of the evaluation information is used to characterize the probability that the evaluation information matches a preset evaluation type.

[0145] For example, embodiments of this application perform multi-class prediction based on the semantic expression features of the evaluation information to obtain the predicted evaluation type of the evaluation information. Specifically, the probability of the target object hitting a preset evaluation type before normalization can be used as the probability quantification value of the evaluation information.

[0146] In the specific implementation, the maximum quantification value among the multiple possible values ​​corresponding to each evaluation information is determined, and the maximum quantification value of each evaluation information is aggregated based on the weight of each evaluation information to obtain the aggregated quantification value.

[0147] If the aggregated quantization value is greater than the probability threshold corresponding to the preset type, then the preset type will be used as the type label of the target object.

[0148] For example, refer to Figure 9 Assuming that the classification model provided in this application can predict preset types (i.e., object types, such as merchant types) as y1, y2 and y3, and can predict preset evaluation types as t1, t2 and t3.

[0149] Assume that target object i corresponds to J evaluation information, and the j-th evaluation information C ij Semantic extraction is performed to obtain semantic representation features (MLP). C(BERT(C ij It can also use MLP to evaluate J pieces of information. C (BERT(C ij Input each vector into the softmax function to obtain J three-dimensional vectors [f] 11 f 12 f 13 ]、[f 21 f 22 f 23 ]…[f j1 f j2 f j3 ]. Where, f j1 f is the probability that evaluation information j matches evaluation type t1. j2 f is the probability that evaluation information j matches evaluation type t2. j3 It represents the probability that evaluation information j matches evaluation type t3. The value of each probability in the three-dimensional vector before normalization is called the probability quantification value.

[0150] Furthermore, for each of the J evaluation pieces of information, the maximum value W among the probability quantification values ​​corresponding to each three-dimensional vector can be determined. ij To obtain J maximum values ​​W ij Furthermore, the J maximum values ​​W can be determined based on the weighting coefficients of each evaluation information. ij Perform aggregation to obtain the aggregated quantized value R = ∑ j attention ij *W ij attention ij It is the weight coefficient of the evaluation information j of the target object i.

[0151] If the aggregated quantization value R is greater than the probability threshold F1 corresponding to type y1, then the target object is tagged according to type y1, that is, the type label of the target object includes "type y1".

[0152] Similarly, if the aggregated quantization value R is greater than the probability threshold F2 corresponding to type y2, then the target object is tagged according to type y2, that is, the type label of the target object includes "type y2".

[0153] If the aggregated quantization value R is greater than the probability threshold F3 corresponding to type y3, then the target object is tagged according to type y3, that is, the type label of the target object includes "type y3".

[0154] This application also provides a tag completion method that can complete the tags for merchants with missing tags as much as possible, greatly improving the tag completion efficiency. This results in a richer training sample pool and improved model performance.

[0155] In another embodiment of this application, a method for determining the probability threshold described above is also provided. An initial threshold can be determined based on the probability quantification value of a reference object, and the initial threshold can be verified and adjusted based on the actual prediction results of the reference object using a classification model, thereby improving the accuracy of the probability threshold in representing the "probability of hitting a preset type".

[0156] For example, for any preset type K that the classification model can predict, obtain P probability quantization values ​​for P reference objects. Here, the probability quantization values ​​are used to characterize the likelihood that a reference object hits the preset type K. Further, determine Q probability quantization values ​​among the P probability quantization values ​​that are greater than an initial threshold, and determine whether Q / P is greater than a preset ratio value.

[0157] If so, the initial threshold is determined to be the probability threshold corresponding to the preset type; otherwise, the initial threshold is adjusted until Q / P is greater than the preset ratio value.

[0158] refer to Figure 10 This provides a process for determining the probability threshold corresponding to a preset type. Assume that the classification model provided in this embodiment can predict H preset types (e.g., merchant types), y1, y2…y3… k …y H The following uses the default type y. k As an example, the process of determining the probability threshold of a preset type is introduced:

[0159] Semantic extraction is performed on each evaluation information of reference object i, and the semantic expression features are aggregated to obtain the MLP. mt (∑ j attention ij *BERT(C ij MLP can be used mt (∑ j attention ij *BERT(C ij Inputting the softmax function yields an H-dimensional vector [p1, p2, ..., p...]. k …p H ]. Where p k The reference object i hits the type y. k The probability of y. k The value before normalization, which determines the preset type y. k The probability quantization value used when determining the probability threshold.

[0160] Obtain the probability y corresponding to P reference objects k The corresponding P possible quantized values ​​can be used to set the maximum value among the P possible quantized values ​​as the initial threshold T. kDetermine which of the P possible quantified values ​​is greater than T. k The number of quantized values ​​Q is determined. If Q / P is less than a preset ratio (e.g., 80%), then T is reduced. k Continue until Q / P is greater than or equal to the preset ratio.

[0161] It should be noted that in the second label completion process described above, the probability threshold F corresponding to the preset type can be determined based on the aggregated quantization value of each reference object. ck To obtain the P aggregated quantization values ​​corresponding to P reference objects, the maximum value among the P aggregated quantization values ​​can be set as the initial threshold T. ck Determine which of the P aggregated quantization values ​​is greater than T. ck The number of quantized values ​​Q is determined. If Q / P is less than a preset ratio (e.g., 80%), then T is reduced. ck Continue until Q / P is greater than or equal to the preset ratio.

[0162] In another embodiment of this application, if the preset type is a high-priority type (e.g., a high-risk type), then the probability threshold of the preset type is determined based on reference objects with a high hit rate. That is, if the preset type is a high-risk type, then the P reference objects are those objects that hit the preset type and whose hit probability is higher than the probability threshold.

[0163] Among them, the "high-risk type" is a pre-set type, which can be one or more high-risk types determined based on prior knowledge or merchant classification rules in the field of network risk control. For example, "impersonation fraud" or "pornography".

[0164] The hit probability is the probability of belonging to the preset type, and can be the output of a multi-class classification function. For example, assume that the classification model provided in this embodiment can predict preset types y1, y2, and y3. Semantic extraction is performed on each evaluation information of the target object i, and the semantic expression features are aggregated to obtain the MLP. mt (∑ j attention ij *BERT(C ij MLP can be used mt (∑ j attention ij *BERT(C ij Input the softmax function to obtain a three-dimensional vector [p1, p2, p3]. Here, p1 is the probability that target object i hits type y1, p2 is the probability that target object i hits type y2, and p3 is the probability that target object i hits type y3.

[0165] If type y1 is a high type, then reference objects with p1 greater than the probability threshold are selected from multiple reference objects. The probability threshold corresponding to type y1 is determined based on the probability quantification value of the selected reference objects.

[0166] In another embodiment of this application, after labeling the target object with missing type labels based on the label completion processing described above, labeled training samples can be obtained. The training sample set is updated, and the model can be updated and trained based on the updated training sample set.

[0167] For example, historical data of the target object and historical data of the reference object are used as training samples to input into the classification model to obtain the predicted type, predicted attributes, and predicted evaluation type of the evaluation information of the training samples.

[0168] The classification model is iteratively trained based on the loss between the predicted type and the type label of the training sample, the loss between the predicted attribute and the attribute label of the training sample, and the loss between the predicted evaluation type and the evaluation type label of the evaluation information.

[0169] In one possible implementation, retraining of the classification model is only initiated when the number of objects with labels changes significantly after label completion. For example, the number of target objects with type labels is counted after the first and second label completion processes. If the ratio of this number to the number of reference objects before label completion is greater than a preset value (e.g., 5%), the model is retrained.

[0170] The reference object is the object with a type label before the label completion process.

[0171] In another embodiment of this application, to ensure that the device's (e.g., the backend server's) video memory does not overflow during training, and to limit the amount of historical data loaded each time, the data can be filtered. For example, the initially obtained data can be sampled during model training.

[0172] For example, the acquisition of historical data from multiple reference objects mentioned above includes:

[0173] Obtain the initial data of the reference object. If the number of evaluation information in the initial data is greater than the preset threshold M, then select N evaluation information from the initial data in descending order of priority of the evaluation information.

[0174] Randomly sample the remaining evaluation information from the initial data to obtain MN evaluation information;

[0175] Historical data is determined based on N evaluation information entries, evaluation type labels for N evaluation information entries, MN evaluation information entries, evaluation type labels for MN evaluation information entries, type labels for reference objects, and attribute labels for reference objects. Here, M and N can be preset data.

[0176] In other words, when the amount of evaluation information is too large, the evaluation information can be sampled to reduce the memory usage of model training data. The priority of the evaluation information can be related to the time of its occurrence or the size of the data. For example, the more recent the occurrence of the evaluation information, the higher its priority. The larger the amount of evaluation information, the higher its priority. This application does not limit the priority of the evaluation information; any scheme that can represent priority falls within the protection scope of this application.

[0177] refer to Figure 11 If the number of merchant reviews in the initial data is greater than 50, then the 10 most recent reviews of the merchants are selected from the initial data, and then 40 reviews are randomly sampled from the remaining reviews. Finally, the 50 selected reviews are used to build training samples for model training.

[0178] This application also provides a classification information determination device, such as... Figure 12 As shown, the device includes a semantic extraction unit 1201, a weight allocation unit 1202, and a determination unit 1203.

[0179] The semantic extraction unit 1201 is used to obtain at least one evaluation information of the object to be predicted, and to perform semantic extraction on each evaluation information based on a classification model to obtain the semantic expression features of each evaluation information.

[0180] The weight allocation unit 1202 is used to allocate weights to the semantic expression features based on the classification model to obtain intermediate features for each evaluation information.

[0181] The determining unit 1203 is used to perform fusion processing on all intermediate features corresponding to the at least one evaluation information based on the classification model, and determine the classification information of the object to be predicted based on the fusion result; the classification information includes at least one of prediction type and prediction attribute under the prediction type.

[0182] In one embodiment, the weight allocation unit 1202 assigns weights to the semantic expression features based on the classification model to obtain intermediate features for each evaluation information, including:

[0183] For each piece of evaluation information, a weight coefficient for the evaluation information is determined based on the quantized value of the semantic expression feature of the evaluation information; the quantized value of the semantic expression feature is positively correlated with the weight coefficient of the evaluation information.

[0184] The intermediate features of the evaluation information are obtained by multiplying the semantic expression features and the weight coefficients of the evaluation information.

[0185] In one embodiment, such as Figure 13 As shown, the device also includes a model training unit 1204. The model training unit 1204 is used to acquire historical data of multiple reference objects, perform type prediction on the historical data based on an initial model, and obtain a first prediction loss, a second prediction loss, and a third prediction loss for each reference object.

[0186] The initial model is iteratively trained based on the first prediction loss, the second prediction loss, and the third prediction loss to obtain the classification model;

[0187] Wherein, the first prediction loss is the loss between the type label of the reference object and the predicted type, the second prediction loss is the loss between the attribute label of the reference object and the predicted attribute of the reference object, and the third prediction loss is the loss between the evaluation type label of the evaluation information of the reference object and the predicted evaluation type of the evaluation information.

[0188] In one embodiment, the historical data includes the type label of the reference object, the attribute label of the reference object, the historical evaluation information of the reference object, and the evaluation type label of each piece of historical evaluation information;

[0189] The model training unit 1204 performs type prediction on the historical data to obtain a first prediction loss, a second prediction loss, and a third prediction loss for each reference object, including:

[0190] The historical data is input into the initial model, and semantic extraction is performed on each piece of historical evaluation information to obtain the semantic expression features of each piece of historical evaluation information.

[0191] Based on the semantic expression features, type prediction is performed to obtain the predicted type, predicted attribute, and predicted evaluation type of each historical evaluation information for each reference object.

[0192] A first prediction loss is determined based on a first loss function to determine the predicted type and type label of the reference object; a second prediction loss is determined based on a second loss function to determine the predicted attributes and attribute labels of the reference object; and a third prediction loss is determined based on a third loss function to determine the relationship between the evaluation type label of the historical evaluation information and the predicted evaluation type of the historical evaluation information.

[0193] In one embodiment, the model training unit 1204 performs type prediction based on the semantic expression features to obtain the predicted classification for each reference object, including:

[0194] For each semantic expression feature of the reference object, an intermediate feature is determined based on the semantic expression feature and the weight coefficient of the evaluation information corresponding to the semantic expression feature;

[0195] All intermediate features corresponding to the reference object are fused, and the fusion result is scalarized to obtain the first vector;

[0196] The first vector is subjected to binary classification prediction based on a binary classification function to obtain the prediction type of the reference object; the prediction type is used to characterize whether the reference object is abnormal.

[0197] In one embodiment, the model training unit 1204 performs type prediction based on the semantic expression features to obtain the predicted attributes of each reference object, including:

[0198] For each semantic expression feature of the reference object, an intermediate feature is determined based on the semantic expression feature and the weight coefficient of the evaluation information corresponding to the semantic expression feature;

[0199] All intermediate features corresponding to the reference object are fused, and the fusion result is transformed to obtain a second vector.

[0200] Based on the multi-classification function, the second vector is used to perform multi-classification prediction to obtain the predicted attributes of the reference object.

[0201] In one embodiment, the model training unit 1204 performs type prediction based on the semantic expression features to obtain the predicted evaluation type of the historical evaluation information, including:

[0202] The semantic expression features are subjected to dimensionality transformation processing to obtain a third vector;

[0203] Based on the multi-classification function, the third vector is used to perform multi-classification prediction to obtain the predicted evaluation type of the historical evaluation information.

[0204] In one embodiment, the model training unit 1204 is further configured to perform a first label completion process on the target object based on the probability quantization value of the target object to obtain the type label of the target object; the target object is an object that is missing a type label, and the probability quantization value of the target object is used to characterize the probability that the target object hits a preset type.

[0205] In one embodiment, the model training unit 1204 performs first label completion processing on the target object based on the probability quantization value of the target object to obtain the type label of the target object, including:

[0206] If the probability quantification value of the target object is greater than the probability threshold corresponding to the preset type, then the preset type is used as the type label of the target object.

[0207] In one embodiment, the model training unit 1204 is further configured to obtain P probability quantization values ​​for P reference objects; the probability quantization values ​​of the reference objects are used to characterize the probability that the reference objects hit the preset type; the labels of the P reference objects are the preset type;

[0208] Determine Q possible quantization values ​​that are greater than the initial threshold from among the P possible quantization values, and determine whether Q / P is greater than a preset ratio value;

[0209] If so, the initial threshold is determined to be the probability threshold corresponding to the preset type; otherwise, the initial threshold is adjusted until Q / P is greater than the preset ratio value.

[0210] In one embodiment, the P reference objects are objects whose hit probability is higher than a probability threshold among objects of the preset type.

[0211] In one embodiment, the model training unit 1204 is further configured to, if the target object still lacks a type label after the first label completion processing, perform a second label completion processing on the target object based on the probability quantification value of the evaluation information of the target object to obtain the type label of the target object; the probability quantification value of the evaluation information is used to characterize the probability that the evaluation information hits a preset evaluation type.

[0212] In one embodiment, the model training unit 1204 performs second label completion processing on the target object based on the probability quantification value of the evaluation information of the target object to obtain the type label of the target object, including:

[0213] Determine the maximum quantization value among the multiple probability values ​​corresponding to each evaluation information, and aggregate the maximum quantization values ​​of each evaluation information based on the weight of each evaluation information to obtain an aggregated quantization value.

[0214] If the aggregated quantization value is greater than the probability threshold corresponding to the preset type, then the preset type is used as the type label of the target object.

[0215] In one embodiment, the model training unit 1204 is further configured to input the historical data of the target object and the historical data of the reference object as training samples into the classification model to obtain the prediction type, prediction attribute, and prediction evaluation type of the evaluation information in each training sample.

[0216] The classification model is iteratively trained based on the loss between the predicted type and the type label of the training sample, the loss between the predicted attribute and the attribute label of the training sample, and the loss between the predicted evaluation type and the evaluation type label of the evaluation information of the training sample.

[0217] In one embodiment, the semantic extraction unit 1201 acquires historical data of multiple reference objects, including:

[0218] Obtain the initial data of the reference object. If the number of evaluation information in the initial data is greater than a preset threshold M, then select N evaluation information from the initial data in descending order of priority of the evaluation information.

[0219] Randomly sample the remaining evaluation information from the initial data to obtain MN evaluation information;

[0220] The historical data of the reference object is determined based on the N evaluation information, the evaluation type labels of the N evaluation information, the MN evaluation information, the evaluation type labels of the MN evaluation information, the type label of the reference object, and the attribute label of the reference object.

[0221] It should be understood that the units described in the above-mentioned device are related to the reference. Figure 2 The steps in the described method correspond to each other. Therefore, the operations and features described above for the method are also applicable to the classification information determination device and its constituent units, and will not be repeated here. This device can be pre-implemented in a computer device's browser or other security applications, or it can be loaded into a computer device's browser or its security applications through download or other means. The corresponding units in this device can cooperate with the units in the computer device to implement the solutions of the embodiments of this application.

[0222] The division of modules or units mentioned in the detailed description above is not mandatory. In fact, according to the embodiments of this disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0223] It should be noted that for details not disclosed in the apparatus of this application embodiments, please refer to the details disclosed in the above embodiments of this application, which will not be repeated here.

[0224] The following is for reference. Figure 14 , Figure 14 A schematic diagram of a computer device suitable for implementing embodiments of this application is shown, such as... Figure 14 As shown, the computer system 1400 includes a central processing unit (CPU) 1401, which can perform various appropriate actions and processes based on programs stored in read-only memory (ROM) 1402 or programs loaded from storage section 1408 into random access memory (RAM) 1403. The RAM 1403 also stores various programs and data required for the system's operating instructions. The CPU 1401, ROM 1402, and RAM 1403 are interconnected via a bus 1404. An input / output (I / O) interface 1405 is also connected to the bus 1404.

[0225] The following components are connected to I / O interface 1405: an input section 1406 including a keyboard, mouse, etc.; an output section 1407 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 1408 including a hard disk, etc.; and a communication section 1409 including a network interface card such as a LAN card, modem, etc. The communication section 1409 performs communication processing via a network such as the Internet. A drive 1410 is also connected to I / O interface 1405 as needed. Removable media 1411, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 1410 as needed so that computer programs read from them can be installed into storage section 1408 as needed.

[0226] Specifically, according to embodiments of this application, the flowchart above refers to... Figure 2The described process can be implemented as a computer software program. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowchart. In such an embodiment, the computer program contains program code for performing the methods shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via communication section 1409, and / or installed from removable medium 1411. When the computer program is executed by central processing unit (CPU) 1401, it performs the functions defined in the system of this application.

[0227] It should be noted that the computer-readable medium shown in this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media can also be any computer-readable medium other than computer-readable storage media, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.

[0228] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operational instructions of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two connected blocks may actually be executed substantially in parallel, or they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified functions or operational instructions, or using a combination of dedicated hardware and computer instructions.

[0229] The units or modules described in the embodiments of this application can be implemented in software or hardware. The described units or modules can also be housed in a processor; for example, a processor may be described as including a semantic extraction unit, a weight allocation unit, and a determination unit. The names of these units or modules do not necessarily constitute a limitation on the unit or module itself.

[0230] On the other hand, this application also provides a computer-readable storage medium, which may be included in the computer device described in the above embodiments, or may exist independently and not assembled into the computer device. The aforementioned computer-readable storage medium stores one or more programs that, when used by one or more processors, execute the methods described in this application. For example, it may execute... Figure 2 or Figure 11 The steps of the method shown.

[0231] This application provides a computer program product including instructions that, when executed, cause the method described in this application to be performed. For example, it can execute... Figure 2 or Figure 11 The steps of the method shown.

[0232] The above description is merely a preferred embodiment of this application and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of disclosure in this application is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the foregoing disclosed concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions disclosed in this application.

Claims

1. A method for determining classification information, characterized in that, include: Obtain at least one evaluation information of a merchant on an internet platform, and perform semantic extraction on each evaluation information based on a classification model to obtain the semantic expression features of each evaluation information; Based on the classification model, the semantic expression features are weighted to obtain the intermediate features of each evaluation information. Based on the classification model, all intermediate features corresponding to the at least one evaluation information are fused, and the classification information of the merchant is determined according to the fusion result; The classification information includes at least one of the prediction type and the prediction attribute under the prediction type; The training process of the classification model includes: Acquire historical data for multiple reference objects; the historical data includes type labels of the reference objects, attribute labels of the reference objects, historical evaluation information of the reference objects, and evaluation type labels for each piece of historical evaluation information. Input the historical data into the initial model, perform semantic extraction on each piece of historical evaluation information, and obtain the semantic expression features of each piece of historical evaluation information. Type prediction is performed based on the semantic expression features of each historical evaluation information to obtain the predicted type, predicted attribute, and predicted evaluation type of each historical evaluation information for each reference object. A first prediction loss is determined based on a first loss function to determine the predicted type and type label of the reference object; a second prediction loss is determined based on a second loss function to determine the predicted attributes and attribute labels of the reference object; and a third prediction loss is determined based on a third loss function to determine the relationship between the evaluation type label of the historical evaluation information and the predicted evaluation type of the historical evaluation information. The initial model is iteratively trained based on the first prediction loss, the second prediction loss, and the third prediction loss to obtain the classification model.

2. The method according to claim 1, characterized in that, The step of assigning weights to the semantic expression features based on the classification model to obtain intermediate features for each evaluation information includes: For each piece of evaluation information, a weight coefficient for the evaluation information is determined based on the quantified value of the semantic expression feature of the evaluation information; the quantified value of the semantic expression feature is positively correlated with the weight coefficient of the evaluation information. The intermediate features of the evaluation information are obtained by multiplying the semantic expression features and the weight coefficients of the evaluation information.

3. The method according to claim 1, characterized in that, The step of predicting the type based on the semantic expression features to obtain the predicted classification of each reference object includes: For each semantic expression feature of the reference object, an intermediate feature is determined based on the semantic expression feature and the weight coefficient of the evaluation information corresponding to the semantic expression feature; All intermediate features corresponding to the reference object are fused, and the fusion result is scalar transformed to obtain the first vector; The first vector is subjected to binary classification prediction based on a binary classification function to obtain the prediction type of the reference object; the prediction type is used to characterize whether the reference object is abnormal.

4. The method according to claim 1, characterized in that, The step of performing type prediction based on the semantic expression features to obtain the predicted attributes of each reference object includes: For each semantic expression feature of the reference object, an intermediate feature is determined based on the semantic expression feature and the weight coefficient of the evaluation information corresponding to the semantic expression feature; All intermediate features corresponding to the reference object are fused, and the fusion result is transformed to obtain a second vector. Based on the multi-classification function, the second vector is used to perform multi-classification prediction to obtain the predicted attributes of the reference object.

5. The method according to claim 1, characterized in that, The step of predicting the type based on the semantic expression features to obtain the predicted evaluation type of the historical evaluation information includes: The semantic expression features are subjected to dimensionality transformation processing to obtain a third vector; Based on the multi-classification function, the third vector is used to perform multi-classification prediction to obtain the predicted evaluation type of the historical evaluation information.

6. The method according to claim 1, characterized in that, The method further includes: The target object is subjected to first label completion processing based on the probability quantification value of the target object to obtain the type label of the target object; the target object is an object that is missing a type label, and the probability quantification value of the target object is used to characterize the probability that the target object hits a preset type.

7. The method according to claim 6, characterized in that, The step of performing a first label completion process on the target object based on the probability quantification value of the target object to obtain the type label of the target object includes: If the probability quantification value of the target object is greater than the probability threshold corresponding to the preset type, then the preset type is used as the type label of the target object.

8. The method according to claim 7, characterized in that, The method further includes: Obtain P probability quantization values ​​for P reference objects; the probability quantization values ​​of the reference objects are used to characterize the probability that the reference objects hit the preset type; the labels of the P reference objects are the preset type; Determine Q possible quantization values ​​that are greater than the initial threshold from among the P possible quantization values, and determine whether Q / P is greater than a preset ratio value; If so, the initial threshold is determined to be the probability threshold corresponding to the preset type; otherwise, the initial threshold is adjusted until Q / P is greater than the preset ratio value.

9. The method according to claim 8, characterized in that, The P reference objects are those objects of the preset type whose hit probability is higher than a probability threshold.

10. The method according to any one of claims 6-9, characterized in that, The method further includes: If the target object still lacks a type label after the first label completion process, then the target object is subjected to a second label completion process based on the probability quantification value of the evaluation information of the target object to obtain the type label of the target object; the probability quantification value of the evaluation information is used to characterize the probability that the evaluation information hits a preset evaluation type.

11. The method according to claim 10, characterized in that, The process of performing second label completion on the target object based on the probability quantification value of the evaluation information of the target object to obtain the type label of the target object includes: Determine the maximum quantization value among the multiple probability values ​​corresponding to each evaluation information, and aggregate the maximum quantization values ​​of each evaluation information based on the weight of each evaluation information to obtain an aggregated quantization value. If the aggregated quantization value is greater than the probability threshold corresponding to the preset type, then the preset type is used as the type label of the target object.

12. The method according to any one of claims 6-9, characterized in that, The method further includes: The historical data of the target object and the historical data of the reference object are used as training samples and input into the classification model to obtain the predicted type, predicted attributes, and predicted evaluation type of the evaluation information in each training sample. The classification model is iteratively trained based on the loss between the predicted type and the type label of the training sample, the loss between the predicted attribute and the attribute label of the training sample, and the loss between the predicted evaluation type and the evaluation type label of the evaluation information of the training sample.

13. The method according to claim 1, characterized in that, The acquisition of historical data from multiple reference objects includes: Obtain the initial data of the reference object. If the number of evaluation information in the initial data is greater than a preset threshold M, then select N evaluation information from the initial data in descending order of priority of the evaluation information. Randomly sample the remaining evaluation information from the initial data to obtain MN evaluation information; The historical data of the reference object is determined based on the N evaluation information, the evaluation type labels of the N evaluation information, the MN evaluation information, the evaluation type labels of the MN evaluation information, the type label of the reference object, and the attribute label of the reference object.

14. A classification information determination device, characterized in that, include: The semantic extraction unit is used to obtain at least one evaluation information of merchants in the Internet platform, and to perform semantic extraction on each evaluation information based on the classification model to obtain the semantic expression features of each evaluation information. The weight allocation unit is used to allocate weights to the semantic expression features based on the classification model to obtain intermediate features for each evaluation information. The determining unit is configured to perform fusion processing on all intermediate features corresponding to the at least one evaluation information based on the classification model, and determine the classification information of the merchant based on the fusion result; the classification information includes at least one of prediction type and prediction attribute under the prediction type; A model training unit is used to acquire historical data of multiple reference objects, including type labels, attribute labels, historical evaluation information, and evaluation type labels for each historical evaluation information. The historical data is input into an initial model to perform semantic extraction on each historical evaluation information, obtaining semantic expression features for each historical evaluation information. Type prediction is performed based on the semantic expression features of each historical evaluation information to obtain the predicted type, predicted attribute, and predicted evaluation type for each reference object. A first prediction loss is determined based on a first loss function to compare the predicted type and type label of the reference object; a second prediction loss is determined based on a second loss function to compare the predicted attribute and attribute label of the reference object; and a third prediction loss is determined based on a third loss function to compare the evaluation type label and the predicted evaluation type of the historical evaluation information. The initial model is iteratively trained based on the first prediction loss, the second prediction loss, and the third prediction loss to obtain the classification model.

15. The apparatus according to claim 14, characterized in that, The weight allocation unit is further configured to determine the weight coefficient of each evaluation information based on the quantized value of the semantic expression feature of the evaluation information; the quantized value of the semantic expression feature is positively correlated with the weight coefficient of the evaluation information. The intermediate features of the evaluation information are obtained by multiplying the semantic expression features and the weight coefficients of the evaluation information.

16. The apparatus according to claim 14, characterized in that, The model training unit is also used to determine intermediate features for each semantic expression feature of the reference object based on the semantic expression feature and the weight coefficient of the evaluation information corresponding to the semantic expression feature; All intermediate features corresponding to the reference object are fused, and the fusion result is scalarized to obtain the first vector; The first vector is subjected to binary classification prediction based on a binary classification function to obtain the prediction type of the reference object; the prediction type is used to characterize whether the reference object is abnormal.

17. The apparatus according to claim 14, characterized in that, The model training unit is also used to determine intermediate features for each semantic expression feature of the reference object based on the semantic expression feature and the weight coefficient of the evaluation information corresponding to the semantic expression feature; All intermediate features corresponding to the reference object are fused, and the fusion result is transformed to obtain a second vector. Based on the multi-classification function, the second vector is used to perform multi-classification prediction to obtain the predicted attributes of the reference object.

18. The apparatus according to claim 14, characterized in that, The model training unit is also used to perform dimensionality transformation processing on the semantic expression features to obtain a third vector; Based on the multi-classification function, the third vector is used to perform multi-classification prediction to obtain the predicted evaluation type of the historical evaluation information.

19. The apparatus according to claim 14, characterized in that, The model training unit is also used to perform first label completion processing on the target object based on the probability quantization value of the target object to obtain the type label of the target object; the target object is an object that is missing a type label, and the probability quantization value of the target object is used to characterize the probability that the target object hits a preset type.

20. The apparatus according to claim 19, characterized in that, The model training unit is further configured to use the preset type as the type label of the target object if the probability quantification value of the target object is greater than the probability threshold corresponding to the preset type.

21. The apparatus according to claim 20, characterized in that, The model training unit is further configured to obtain P probability quantization values ​​for P reference objects; the probability quantization values ​​of the reference objects are used to characterize the probability that the reference objects hit the preset type; the labels of the P reference objects are the preset type; Determine Q possible quantization values ​​that are greater than the initial threshold from among the P possible quantization values, and determine whether Q / P is greater than a preset ratio value; If so, the initial threshold is determined to be the probability threshold corresponding to the preset type; otherwise, the initial threshold is adjusted until Q / P is greater than the preset ratio value.

22. The apparatus according to claim 21, characterized in that, The P reference objects are those objects of the preset type whose hit probability is higher than a probability threshold.

23. The apparatus according to any one of claims 19-22, characterized in that, The model training unit is further configured to perform a second label completion process on the target object based on the probability quantification value of the evaluation information of the target object if the target object still lacks a type label after the first label completion process, thereby obtaining the type label of the target object; the probability quantification value of the evaluation information is used to characterize the probability that the evaluation information hits a preset evaluation type.

24. The apparatus according to claim 23, characterized in that, The model training unit is further configured to determine the maximum quantization value among the multiple probability values ​​corresponding to each evaluation information, and to aggregate the maximum quantization value of each evaluation information based on the weight of each evaluation information to obtain an aggregated quantization value. If the aggregated quantization value is greater than the probability threshold corresponding to the preset type, then the preset type is used as the type label of the target object.

25. The apparatus according to any one of claims 19-22, characterized in that, The model training unit is also used to input the historical data of the target object and the historical data of the reference object as training samples into the classification model to obtain the prediction type, prediction attribute and prediction evaluation type of the evaluation information in each training sample. The classification model is iteratively trained based on the loss between the predicted type and the type label of the training sample, the loss between the predicted attribute and the attribute label of the training sample, and the loss between the predicted evaluation type and the evaluation type label of the evaluation information of the training sample.

26. The apparatus according to claim 14, characterized in that, The semantic extraction unit is also used to obtain the initial data of the reference object. If the number of evaluation information in the initial data is greater than a preset threshold M, then N evaluation information are selected from the initial data in descending order of priority of the evaluation information. Randomly sample the remaining evaluation information from the initial data to obtain MN evaluation information; The historical data of the reference object is determined based on the N evaluation information, the evaluation type labels of the N evaluation information, the MN evaluation information, the evaluation type labels of the MN evaluation information, the type label of the reference object, and the attribute label of the reference object.

27. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method for determining classification information as described in any one of claims 1-13.

28. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the method for determining classification information as described in any one of claims 1-13.

29. A computer program product, characterized in that, The computer program product includes instructions that, when executed, cause the method as described in any one of claims 1 to 13 to be performed.

Citation Information

Patent Citations

  • Method and device for training a multi-label classification model

    CN109840531A

  • Emotion classification method and system for commodity user comment texts

    CN111666410A

  • Commodity recommendation method based on neighbor users and comment information

    CN112884551A