APP classification method and system based on virtual class and electronic equipment
Through the APP classification method based on virtual classes, the BERT model and the virtual class generation model are used to generate virtual class features, and the original class features are combined for classification prediction, which solves the problem of low classification accuracy in traditional APP classification methods, and achieves higher classification accuracy and adaptability.
Patent Information
- Application Number
- CN202510037055.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-09
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing APP classification methods rely on manual annotation and rule-based classification algorithms, which have problems such as strong subjectivity, inefficiency, and difficulty in formulating rules, resulting in inaccurate classification.
A virtual class-based APP classification method is adopted, and the APP title and text description data are obtained for preprocessing, and a word chunk vector sequence is generated using the BERT model, and virtual class features are generated in the virtual class generation model, and classification prediction is performed based on the original class features.
It improves the classification accuracy of the APP, expands the category boundaries, explores potential category relationships, provides more discriminant information for the classification model, reduces the risk of classification label missing, and improves the overall classification performance and adaptability.
Smart Images

Figure CN119961752A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of natural language processing and machine learning, and in particular to a method, system and electronic device for classifying APPs based on virtual categories. Background Art
[0002] At present, the number and types of APPs are showing explosive growth. In order to facilitate users to use and developers to understand market trends, APP classification technology has come into being. However, traditional APP classification methods mainly rely on manual labeling and rule-based classification algorithms. These methods have problems such as strong subjectivity, low efficiency, and difficulty in formulating rules, resulting in inaccurate classification.
[0003] Therefore, a new APP classification method is needed to achieve accurate classification of APPs. Summary of the invention
[0004] The purpose of this application is to provide a virtual class-based APP classification method, system and electronic device to solve the problem of low classification accuracy in existing APP classification methods in related technologies.
[0005] In order to achieve the above objectives, this application provides the following technical solutions:
[0006] In the first aspect, the present application provides a virtual class-based APP classification method, including:
[0007] Obtain the title and text description data of the APP, and preprocess the title and text description data to obtain a word block sequence corresponding to the title and text description data;
[0008] Encode the word chunk sequence through the BERT model to generate a word chunk vector sequence;
[0009] Inputting the chunk vector sequence into a preset virtual class generation model to generate at least one virtual class feature, and forming a virtual class feature set from the virtual class features;
[0010] According to the virtual class feature set and the preset original class feature set, the APP is classified and predicted, and the classification result of the APP is obtained.
[0011] Furthermore, the preprocessing of the title and text description data to obtain a word block sequence corresponding to the title and text description data includes:
[0012] Performing data cleaning on the title and text description data to obtain cleaned title and text description data;
[0013] The cleaned title and text description data are segmented to obtain word block sequences corresponding to the title and text description data.
[0014] Furthermore, encoding the word chunk sequence through a BERT model to generate a word chunk vector sequence includes:
[0015] Each word block in the word block set is passed through the BERT model to generate a corresponding word block vector; all word block vectors are aggregated to form a word block vector sequence.
[0016] Furthermore, the word block vector sequence is input into a preset virtual class generation model to generate at least one virtual class feature, and the virtual class features are formed into a virtual class feature set, using the following calculation formula:
[0017] Q=W Q h T ,K=W K h T ,V=W V h T
[0018]
[0019] Among them, Q is the query vector space, K is the key vector space, V is the value vector space, and W Q , W K and W V are weight matrices, Attention is the attention score, and are the output results of the first and second layers of the perceptron, ReLU is the activation function, h T is the input word chunk vector sequence, and W G3 The weight matrix, and b G3 is the bias vector, v j It is a virtual class feature.
[0020] Further, the classifying the APP according to the virtual class feature set and the preset original class feature set, performing classification prediction on the APP, and obtaining the classification result of the APP includes:
[0021] Merging the virtual class feature set with a preset original class feature set to obtain an extended feature set;
[0022] Each feature in the extended feature set is converted into a corresponding discriminant embedding vector through a preset embedding layer E;
[0023] According to the discriminant embedding vector, the word block vector sequence is classified and predicted using a fully connected layer F to obtain the classification result of the APP.
[0024] Furthermore, the embedding layer E is calculated using the following formula:
[0025]
[0026] Among them, e exit is the discriminant embedding vector, ReLU is the activation function, W E is the weight matrix of the embedding layer E, c exit is the feature in the extended feature set, b E is the bias vector.
[0027] Furthermore, according to the discriminant embedding vector, the word block vector sequence is classified and predicted using a fully connected layer F, thereby obtaining the classification result of the APP:
[0028]
[0029] Among them, z exit is the vector calculated by the fully connected layer F, W F is the weight matrix of the fully connected layer F, b F is the bias vector, p(c k ∣h T ) is the discriminant embedding vector belonging to category c k probability.
[0030] Furthermore, the following joint loss function is used between the virtual class features and the original class features:
[0031]
[0032] L=L ce +λL vc
[0033] Among them, L ce is the cross entropy loss function, y k is the word vector sequence of the original class feature, L vc is the virtual class constraint loss function, consine is the cosine distance formula, and L is the joint loss function.
[0034] In a second aspect, the present application also provides an APP classification system based on virtual categories, including:
[0035] A preprocessing module, used to obtain the title and text description data of the APP, and preprocess the title and text description data to obtain a word block sequence corresponding to the title and text description data;
[0036] An encoding module, used to encode the word chunk sequence through a BERT model to generate a word chunk vector sequence;
[0037] A feature generation module, used for inputting the word chunk vector sequence into a preset virtual class generation model, generating at least one virtual class feature, and forming the virtual class features into a virtual class feature set;
[0038] The classification module is used to perform classification prediction on the APP according to the virtual class feature set and the preset original class feature set, and obtain the classification result of the APP.
[0039] In a third aspect, the present application also provides a computer electronic device, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of any one of the above-mentioned virtual class-based APP classification methods are implemented.
[0040] The present application provides a virtual class-based APP classification method, system, and electronic device. The method obtains the title and text description data of the APP, and pre-processes the title and text description data to obtain a word block sequence corresponding to the title and text description data; encodes the word block sequence through the BERT model to generate a word block vector sequence; inputs the word block vector sequence into a preset virtual class generation model to generate at least one virtual class feature, and forms a virtual class feature set with the virtual class features; classifies and predicts the APP based on the virtual class feature set and the preset original class feature set, and obtains the classification result of the APP. The present application innovatively uses a generative network with a specific structure to generate virtual class features from APP text features, expands the category boundaries, mines potential category relationships, and provides more discriminant information for the classification model. It can effectively deal with complex APP text classification, reduce the risk of missing classification labels, and improve overall classification performance and adaptability. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 It is a flowchart of a virtual class-based APP classification method according to an embodiment of the present application;
[0042] Figure 2 It is a structural diagram of an APP classification system based on virtual classes according to an embodiment of the present application;
[0043] Figure 3 It is a structural schematic diagram of a computer electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0044] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0045] It should be noted that when an element is referred to as being "fixed to" another element, it may be directly on the other element or there may be a central element. When an element is considered to be "connected to" another element, it may be directly connected to the other element or there may be a central element at the same time. In contrast, when an element is referred to as being "directly on" another element, there is no intermediate element. The terms "vertical", "horizontal", "left", "right" and similar expressions used herein are for illustrative purposes only.
[0046] In this application, unless otherwise clearly specified and limited, the terms "installed", "connected", "connected", "fixed" and the like should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, it can be the internal connection of two elements or the interaction relationship between two elements. For ordinary technicians in this field, the specific meanings of the above terms in this application can be understood according to specific circumstances.
[0047] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of this application, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.
[0048] The terms used in one or more embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to limit one or more embodiments of the present application. The singular forms "a", "said" and "the" used in one or more embodiments of the present application are also intended to include plural forms, unless the context clearly indicates other meanings.
[0049] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein in the specification of the template are only for the purpose of describing specific embodiments and are not intended to limit this application. The term "and / or" used herein includes any and all combinations of one or more of the associated listed items.
[0050] It should be understood that, although the terms first, second, etc. may be used to describe various information in one or more embodiments of the present application, these information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of the present application, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when the time of".
[0051] At present, the existing APP classification methods have the following problems: Most traditional methods only rely on a predefined fixed category system for feature learning and classification. They usually extract features directly from the text data of the APP and use simple classification models (such as Naive Bayes, support vector machines, or basic neural networks) to map the APP to existing category labels. These methods do not take into account the potential relationship and boundary ambiguity between categories. The feature learning process is relatively simple and cannot fully mine the deep semantic information in the text. As a result, it is difficult to accurately determine the category of some emerging or complex APPs, and it is easy to have inaccurate or missing labels.
[0052] The technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems are described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes are not repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.
[0053] Please refer to Figure 1 The embodiment of the present application provides an APP classification method based on virtual classes, which is applied to an APP classification system and includes at least the following steps:
[0054] S10. Obtain the title and text description data of the APP, and pre-process the title and text description data to obtain a word block sequence corresponding to the title and text description data.
[0055] Specifically, in this embodiment, it is first necessary to obtain the title content and corresponding text description data of the APP to be classified, that is, a brief description of the functions of the APP.
[0056] After the text data is obtained, it is necessary to pre-process the text data to facilitate the subsequent use of the text data.
[0057] In one embodiment of the present application, the preprocessing of the title and text description data to obtain a word block sequence corresponding to the title and text description data includes:
[0058] S101. Clean the title and text description data to obtain the cleaned title and text description data.
[0059] S102. Perform word segmentation on the cleaned title and text description data to obtain the word block sequence corresponding to the title and text description data.
[0060] Specifically, data cleaning can include removing stop words and punctuation. Exemplarily, for example, there is an APP with the title "Intelligent Healthy Diet Assistant and Exercise Plan Customization" and the description "This APP combines advanced artificial intelligence algorithms and can formulate personalized healthy diet plans according to the user's physical condition, dietary preferences, and fitness goals, and provide professional exercise training plans. It also has functions such as diet check-in, exercise record, and social sharing to help users easily manage a healthy life."
[0061] First, clean this text data to remove stop words (such as "of", "this", "and", etc.) and punctuation, and then perform word segmentation to obtain the word sequence: ["Intelligent", "Healthy", "Diet", "Assistant", "Exercise", "Plan", "Customization", "Combine", "Advanced", "Artificial", "Intelligence", "Algorithm", "According", "To", "User", "Physical", "Condition", "Preference", "Fitness", "Goal", "Formulate", "Personalized", "Plan", "Have", "Check-in", "Record", "Social", "Sharing", "Function", "Help", "Easily", "Manage", "Life"].
[0062] S20. Encode the word block sequence through the BERT model to generate a word block vector sequence.
[0063] It should be noted that BERT (Bidirectional Encoder Representations from Transformers) is a pre-trained language model launched by Google in 2018. Its emergence is to solve the limitations of unidirectional language models in natural language processing (NLP). Before BERT, language models such as early recurrent neural networks (RNNs) and long short-term memory networks (LSTMs) were mainly unidirectional and could only utilize the information of the previous or subsequent words in the text, while the innovation of BERT is that it can utilize the information of both the previous and subsequent words of a word, thus providing a more comprehensive language understanding ability.
[0064] In a certain embodiment of the present application, step S20. Encode the word block sequence through the BERT model to generate a word block vector sequence, including:
[0065] S201, passing each word block in the word block set through the BERT model to generate a corresponding word block vector;
[0066] S202: Aggregate all chunk vectors to form a chunk vector sequence.
[0067] Specifically, the pre-trained BERT model is used to map each token into a 768-dimensional vector representation. For each token in the APP text, if it exists in the BERT vocabulary, the corresponding 768-dimensional vector is directly obtained; if it does not exist, the unknown token ([UNK]) vector representation of BERT is used.
[0068] S30, inputting the word block vector sequence into a preset virtual class generation model to generate at least one virtual class feature, and forming a virtual class feature set from the virtual class features.
[0069] Specifically, a generation network is used to generate virtual class features. For the convenience of subsequent description, in this embodiment and subsequent embodiments, the number of virtual class features is uniformly set to 20, as follows:
[0070] 1. First, h is transformed through a linear transformation layer T (Assume that the APP text feature vector after BERT encoding has a dimension of 768) is mapped to the query vector space Q, key vector space K and value vector space V, and the weight matrices are W Q ∈R 384×768 ,W K ∈R 384×768 and W V ∈R 384×768
[0071] Q=W Q h T ,K=W K h T ,V=W V h T
[0072] 2. Then calculate the attention score:
[0073]
[0074] 3. The feature vector processed by the attention mechanism enters the multi-layer perceptron (MLP). The weight matrix of the first hidden layer of MLP is The bias vector is The activation function uses Relu, and the output is:
[0075]
[0076] 4. The weight matrix of the second hidden layer is The bias vector is The output is
[0077]
[0078] 5. Finally, a linear layer is used to map it into a virtual class feature space, and the weight matrix is W G3 ∈R 768×768 , the bias vector is b G3 ∈R 768 Generate m = 20 virtual class features V = {v1, v2, ..., v 20},in:
[0079]
[0080] Among them, Q is the query vector space, K is the key vector space, V is the value vector space, and W Q , W K and W V are weight matrices, Attention is the attention score, and are the output results of the first and second layers of the perceptron, ReLU is the activation function, h T is the input word chunk vector sequence, and W G3 The weight matrix, and b G3 is the bias vector, v j It is a virtual class feature.
[0081] S40: According to the virtual class feature set and the preset original class feature set, classify and predict the APP, and obtain the classification result of the APP.
[0082] Specifically, in one embodiment of the present application, step S40 includes:
[0083] S401: Merge the virtual class feature set and the preset original class feature set to obtain an extended feature set.
[0084] S402: Convert each feature in the extended feature set into a corresponding discriminant embedding vector through a preset embedding layer E.
[0085] S403: Based on the discriminant embedding vector, a fully connected layer F is used to perform classification prediction on the word block vector sequence, so as to obtain a classification result of the APP.
[0086] It can be understood that the preset original class feature set is a pre-set original category feature label, such as "health management category", "fitness and exercise category", "life service category", etc.
[0087] It should be noted that the embedding layer E (Embedding Layer) is a layer in the neural network that is used to convert discrete categorical data (such as words, category labels, etc.) into a low-dimensional continuous vector space representation. Its essence is a mapping function that maps discrete symbols to a vector space of fixed dimension. The fully connected layer (Fully-ConnectedLayer), also known as the dense layer (Dense Layer), is a neural network layer. In this layer, each neuron is connected to all neurons in the previous layer. This means that if the previous layer has n neurons and the current fully connected layer has m neurons, then there are n*m connection weights from the previous layer to the current layer.
[0088] In this example, the original class feature label is 100 and the virtual class feature is 20. The specific process is as follows:
[0089] 1. The original three-level industry label set C = {c1, c2, ..., c 100} (assuming a total of 100 original categories, such as "health management", "fitness", "life service", etc.) and the generated virtual class feature V are merged into an extended feature set C ext =C∪V.
[0090] 2. For each feature c exti ∈C ext (i=1,2,…,120), which is mapped into a 384-dimensional discriminant embedding vector through the embedding layer E The weight matrix of the embedding layer E is W E ∈R 384×768 , the bias vector is b E ∈R 384 ,Right now:
[0091]
[0092] 3. Then use the fully connected layer F for classification prediction. The weight matrix of the fully connected layer is W F ∈R 100x38 The bias vector is b F ∈R 100 For the input APP text feature h T (i.e., word vector sequence), which belongs to category c k The probability p(c k ∣h T ) is calculated as follows:
[0093]
[0094] Among them, e exitis the discriminant embedding vector, ReLU is the activation function, W E is the weight matrix of the embedding layer E, c exit is the feature in the extended feature set, b E is the bias vector, z exit is the vector calculated by the fully connected layer F, W F is the weight matrix of the fully connected layer F, b F is the bias vector, p(c k ∣h T ) is the discriminant embedding vector belonging to category c k probability.
[0095] It should be noted that the virtual category is only used for auxiliary classification and is not used as the final classification result. That is, the real classification label of the APP still uses the label in the original category.
[0096] In one example of the present application, the following joint loss function is used between the virtual class features and the original class features:
[0097]
[0098] L=L ce +λL vc
[0099] Among them, L ce is the cross entropy loss function, y k is the word vector sequence of the original class feature, L vc is the virtual class constraint loss function, consine is the cosine distance formula, and L is the joint loss function.
[0100] Specifically, in order to enable the model to effectively learn the discriminative relationship between virtual classes and original classes, a joint loss function L is designed that combines the cross-extraction loss and the virtual class constraint loss. ce It is used to measure the difference between the predicted label and the true label. Assume that the one-hot encoding of the true label is y = [y1, y2, ..., y 100 ],but:
[0101]
[0102] Virtual class constraint loss L vc The distribution of virtual classes is constrained based on the cosine distance between the virtual class and the original class. For the original class c j The discriminant embedding vector of For virtual class v i The discriminant embedding vector of is, and the cosine distance formula is The virtual classes are constrained by maximizing the following loss:
[0103]
[0104] The joint loss function is:
[0105] L=L ce +λL vc
[0106] Among them, λ is the balance coefficient, which is used to adjust the weight of the virtual class constraint loss in the total loss. In this example, the value of λ is 0.5.
[0107] The present application provides a virtual class-based APP classification method, which obtains the title and text description data of the APP, and pre-processes the title and text description data to obtain the word block sequence corresponding to the title and text description data; encodes the word block sequence through the BERT model to generate a word block vector sequence; inputs the word block vector sequence into a preset virtual class generation model to generate at least one virtual class feature, and forms a virtual class feature set with the virtual class features; classifies and predicts the APP according to the virtual class feature set and the preset original class feature set, and obtains the classification result of the APP. The present application innovatively uses a generative network with a specific structure to generate virtual class features from APP text features, expands the category boundaries, explores potential category relationships, and provides more discriminant information for the classification model. It can effectively deal with complex APP text classification, reduce the risk of missing classification labels, and improve overall classification performance and adaptability.
[0108] See also Figure 2 The present application also provides an APP classification system 200 based on virtual classes, including:
[0109] A preprocessing module 201 is used to obtain the title and text description data of the APP, and preprocess the title and text description data to obtain a word block sequence corresponding to the title and text description data;
[0110] An encoding module 202, configured to encode the word chunk sequence through a BERT model to generate a word chunk vector sequence;
[0111] A feature generation module 203, configured to input the chunk vector sequence into a preset virtual class generation model, generate at least one virtual class feature, and form a virtual class feature set from the virtual class features;
[0112] The classification module 204 is used to perform classification prediction on the APP according to the virtual class feature set and the preset original class feature set, and obtain the classification result of the APP.
[0113] See also Figure 3The embodiment of the present application further provides a computer electronic device 300, including a memory 303 and a processor 302, wherein the memory 303 stores a computer program, and when the processor executes the computer program, the steps of any of the above-mentioned virtual class-based APP classification methods are implemented.
[0114] Specifically, the electronic device 300 includes: a transceiver 301, a bus interface and a processor 302, the processor 302 is used to obtain the title and text description data of the APP, and pre-process the title and text description data to obtain a word block sequence corresponding to the title and text description data; encode the word block sequence through a BERT model to generate a word block vector sequence; input the word block vector sequence into a preset virtual class generation model to generate at least one virtual class feature, and form the virtual class features into a virtual class feature set; classify and predict the APP according to the virtual class feature set and the preset original class feature set, and obtain the classification result of the APP.
[0115] In the embodiment of the present application, the electronic device 300 further includes: a memory 303. Figure 3 In the embodiment, the bus architecture may include any number of interconnected buses and bridges, specifically linking various circuits of one or more processors represented by processor 302 and memory represented by memory 303. The bus architecture may also link various other circuits such as peripherals, voltage regulators, and power management circuits, which are well known in the art and are therefore not further described herein. The bus interface provides an interface. The transceiver 301 may be a plurality of components, i.e., including a transmitter and a receiver, providing a unit for communicating with various other devices on a transmission medium. The processor 302 is responsible for managing the bus architecture and general processing, and the memory 303 may store data used by the processor 302 when performing operations.
[0116] In this embodiment, the computer readable storage medium may be a non-volatile storage medium or a volatile storage medium. For example, the computer storage medium may include but is not limited to: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and other media that can store program codes.
[0117] In all examples shown and described herein, any specific values should be interpreted as merely exemplary and not as limiting, and thus other examples of the exemplary embodiments may have different values.
[0118] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, further definition and explanation thereof is not required in subsequent drawings.
[0119] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely schematic. For example, the flowcharts and structure diagrams in the accompanying drawings show the possible architecture, functions and operations of the devices, methods and computer program products according to multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of a code, and the module, a program segment or a part of a code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in an alternative implementation, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the structure diagram and / or the flow diagram, and the combination of boxes in the structure diagram and / or the flow diagram, can be implemented with a dedicated hardware-based system that performs a specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.
[0120] In addition, the functional modules or units in the various embodiments of the present application may be integrated together to form an independent part, or each module may exist separately, or two or more modules may be integrated to form an independent part.
[0121] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a terminal device (which can be a smart phone, a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in the various embodiments of the present application.
[0122] The above description is only a specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application.
Claims
1. A virtual class-based APP classification method, characterized in that: include: Obtain the title and text description data of the APP, and preprocess the title and text description data to obtain a word block sequence corresponding to the title and text description data; Encode the word chunk sequence through the BERT model to generate a word chunk vector sequence; Inputting the chunk vector sequence into a preset virtual class generation model to generate at least one virtual class feature, and forming a virtual class feature set from the virtual class features; According to the virtual class feature set and the preset original class feature set, the APP is classified and predicted, and the classification result of the APP is obtained.
2. The APP classification method based on virtual classes according to claim 1 is characterized in that: The preprocessing of the title and text description data to obtain a word block sequence corresponding to the title and text description data includes: Performing data cleaning on the title and text description data to obtain cleaned title and text description data; The cleaned title and text description data are segmented to obtain word block sequences corresponding to the title and text description data.
3. The APP classification method based on virtual classes according to claim 1 is characterized in that: The step of encoding the word chunk sequence through a BERT model to generate a word chunk vector sequence includes: Pass each word block in the word block set through the BERT model to generate a corresponding word block vector; All chunk vectors are aggregated to form a chunk vector sequence.
4. The APP classification method based on virtual classes according to claim 1 is characterized in that: The word block vector sequence is input into a preset virtual class generation model to generate at least one virtual class feature, and the virtual class features are formed into a virtual class feature set, using the following calculation formula: Q=W Q h T ,K=W K h T ,V=W V h T Among them, Q is the query vector space, K is the key vector space, V is the value vector space, and W Q , W K and W V are weight matrices, Attention is the attention score, and are the output results of the first and second layers of the perceptron, ReLU is the activation function, h T is the input word chunk vector sequence, and W G3 The weight matrix, and b G3 is the bias vector, v j It is a virtual class feature.
5. The APP classification method based on virtual classes according to claim 1 is characterized in that: The classifying the APP according to the virtual class feature set and the preset original class feature set, performing classification prediction on the APP, and obtaining the classification result of the APP includes: Merging the virtual class feature set with a preset original class feature set to obtain an extended feature set; Each feature in the extended feature set is converted into a corresponding discriminant embedding vector through a preset embedding layer E; According to the discriminant embedding vector, the word block vector sequence is classified and predicted using a fully connected layer F to obtain the classification result of the APP.
6. The method for classifying virtual apps according to claim 5, characterized in that: The embedding layer E is calculated using the following formula: Among them, e exit is the discriminant embedding vector, ReLU is the activation function, W E is the weight matrix of the embedding layer E, c exit is the feature in the extended feature set, b E is the bias vector.
7. The APP classification method based on virtual classes according to claim 5 is characterized in that: According to the discriminant embedding vector, the block vector sequence is classified and predicted using a fully connected layer F, thereby obtaining the classification result of the APP: Among them, z exit is the vector calculated by the fully connected layer F, W F is the weight matrix of the fully connected layer F, b F is the bias vector, p(c k ∣h T ) is the discriminant embedding vector belonging to category c k probability.
8. The APP classification method based on virtual classes according to claim 1 is characterized in that: The following joint loss function is used between the virtual class features and the original class features: L=L ce +λL vc Among them, L ce is the cross entropy loss function, y k is the word vector sequence of the original class feature, L vc is the virtual class constraint loss function, consine is the cosine distance formula, and L is the joint loss function.
9. An APP classification system based on virtual classes, characterized in that: include: A preprocessing module, used to obtain the title and text description data of the APP, and preprocess the title and text description data to obtain a word block sequence corresponding to the title and text description data; An encoding module, used to encode the word chunk sequence through a BERT model to generate a word chunk vector sequence; A feature generation module, used for inputting the word chunk vector sequence into a preset virtual class generation model, generating at least one virtual class feature, and forming the virtual class features into a virtual class feature set; The classification module is used to perform classification prediction on the APP according to the virtual class feature set and the preset original class feature set, and obtain the classification result of the APP.
10. A computer electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the virtual class-based APP classification method described in any one of claims 1 to 8 when executing the computer program.
Citation Information
Patent Citations
Application program classification method and device and terminal equipment
CN111797239A
Application classification method, device and apparatus
CN113553434A
Program classification model training method and device and program category identification method and device
CN117056836A